Sign in

Simone Scardapane

@sscardapane.bsky.social
526 followers 35 following 42 posts

I fall in love with a new #machinelearning topic every month 🙄 Ass. Prof. Sapienza (Rome) | Author: Alice in a differentiable wonderland (www.sscardapane.it/alice-book)

PostsRepliesMedia
Reposted by Simone Scardapane
Nathan Godey @nthngdy.bsky.social · 06/03/2025
Thanks a lot to all my amazing co-authors @alessiodevoto.bsky.social @sscardapane.bsky.social @yuzhaouoe.bsky.social @neuralnoise.com Eric de la Clergerie @bensagot.bsky.social And a special thanks to @edoardo-ponti.bsky.social for the academic visit that made this work possible!
121
Reposted by Simone Scardapane
Donato Crisostomi ✈️ NeurIPS @crisostomi.bsky.social · 11/03/2025
Will present this at #CVPR ✈️ See you in Nashville 🇺🇸! Kudos to the team 👏 Antonio A. Gargiulo, @mariasofiab.bsky.social, @sscardapane.bsky.social, Fabrizio Silvestri, Emanuele Rodolà.
052
Reposted by Simone Scardapane
Pasquale Minervini @neuralnoise.com · 13/03/2025
Please share it within your circles! edin.ac/3DDQK1o
0139
Reposted by Simone Scardapane
Nathan Godey @nthngdy.bsky.social · 06/03/2025
🚀 New Paper Alert! 🚀 We introduce Q-Filters, a training-free method for efficient KV Cache compression! It is compatible with FlashAttention and can compress along generation which is particularly useful for reasoning models ⚡ TLDR: we make Streaming-LLM smarter using the geometry of attention
1207
Reposted by Simone Scardapane
Nathan Godey @nthngdy.bsky.social · 06/03/2025
Q-Filters is very efficient which allows streaming compression at virtually no latency cost, just like Streaming-LLM... ...but it is also much better at retaining relevant KV pairs compared to fast alternatives (and can even beat slower algorithms such as SnapKV)
111
Simone Scardapane @sscardapane.bsky.social · 27/02/2025
*Compositionality and Ambiguity: Latent Co-occurrence and Interpretable Subspaces* by @maclarke.bsky.social et al. Studies co-occurence of SAE features and how they can be understood as composite / ambiguous concepts. www.lesswrong.com/posts/WNoqEi...
lesswrong.com
Compositionality and Ambiguity:  Latent Co-occurrence and Interpretable Subspaces — LessWrong
Matthew A. Clarke, Hardik Bhatnagar and Joseph Bloom
030
Simone Scardapane @sscardapane.bsky.social · 18/02/2025
*Weighted Skip Connections are Not Harmful for Deep Nets* by @rupspace.bsky.social Cool blog post "in defense" of weighted variants of ResNets (aka HighwayNets) - as a follow up to a previous post by @giffmana.ai. rupeshks.cc/blog/skip.html
rupeshks.cc
Weighted Skip Connections are Not Harmful for Deep Nets
Give Gates a Chance
091
Simone Scardapane @sscardapane.bsky.social · 17/02/2025
*CAT: Content-Adaptive Image Tokenization* by @junhongshen1.bsky.social @lukezettlemoyer.bsky.social et al. They use an LLM to predict a "complexity score" for each image token, which in turns decides the size of its VAE latent representation. arxiv.org/abs/2501.03120
000
Simone Scardapane @sscardapane.bsky.social · 14/02/2025
*Accurate predictions on small data with a tabular foundation model* by Noah Hollmann et al. A transformer for tabular data that takes an entire training set as input and provides predictions - trained on millions of synthetic datasets. www.nature.com/articles/s41...
011
Simone Scardapane @sscardapane.bsky.social · 13/02/2025
*Insights on Galaxy Evolution from Interpretable Sparse Feature Networks* by @jwuphysics.bsky.social Integrates a sparse dictionary step on the last layer of a CNN to obtain a set of interpretable features on multiple astronomical prediction tasks. arxiv.org/abs/2501.00089
030
Simone Scardapane @sscardapane.bsky.social · 10/02/2025
*Round and Round We Go! What makes Rotary Positional Encodings useful?* by @petar-v.bsky.social et al. They show RoPE has distinct behavior for different rotation angles - high freq for position, low freq for semantics. arxiv.org/abs/2410.06205
061
Simone Scardapane @sscardapane.bsky.social · 03/02/2025
*Cautious Optimizers: Improving Training with One Line of Code* by Liang et al. Adding a simple masking operation to momentum-based optimizers can significantly boost their speed. arxiv.org/abs/2411.16085
021
Simone Scardapane @sscardapane.bsky.social · 31/01/2025
*Byte Latent Transformer: Patches Scale Better Than Tokens* by @artidoro.bsky.social et al. Trains a small encoder to dynamically aggregate bytes into tokens, which are input to a standard autoregressive model. Nice direction! arxiv.org/abs/2412.09871
040
Simone Scardapane @sscardapane.bsky.social · 28/01/2025
*Understanding Gradient Descent through the Training Jacobian* by @norabelrose.bsky.social @eleutherai.bsky.social Analyzes training through the spectrum of the "training Jacobian" (∇ of trained weights wrt initial weights), identifying a large inactive subspace. arxiv.org/abs/2412.07003
050
Simone Scardapane @sscardapane.bsky.social · 27/01/2025
*Mixture of A Million Experts* by Xu Owen He Scales a MoE architecture up to millions of experts by implementing a fast retrieval method in the router, inspired by recent MoE scaling laws. arxiv.org/abs/2407.04153
020
Simone Scardapane @sscardapane.bsky.social · 23/01/2025
*Restructuring Vector Quantization with the Rotation Trick* by Fifty et al. Replaces the "closest codebook" operation in vector quantization with a rotation and rescaling operations to improve the back-propagation of gradients. arxiv.org/abs/2410.06424
261
Simone Scardapane @sscardapane.bsky.social · 21/01/2025
*On the Surprising Effectiveness of Attention Transfer for Vision Transformers* by Li et al. Shows that distilling attention patterns in ViTs is competitive with standard fine-tuning. arxiv.org/abs/2411.09702
090
Simone Scardapane @sscardapane.bsky.social · 17/01/2025
*The Super Weight in Large Language Models* by Yu et al. Identifies single weights in LLMs that destroy inference when deactivated. Tracks their mechanisms through the LLM and proposes quantization-specific techniques. arxiv.org/abs/2411.07191
020
Simone Scardapane @sscardapane.bsky.social · 16/01/2025
*The Surprising Effectiveness of Test-Time Training for Abstract Reasoning* by @ekinakyurek.bsky.social et al. Shows that test-time training (fine-tuning at inference time) strongly improves performance on the ARC dataset. arxiv.org/abs/2411.07279
030
Reposted by Simone Scardapane
Fabio Montello @zioictus.bsky.social · 15/01/2025
Our paper “A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion" is out as preprint! By myself, @sscardapane.bsky.social, @rgring.bsky.social and @lanalpa.bsky.social 📄 arxiv.org/abs/2501.07451
arxiv.org
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
Model compression is essential in the deployment of large Computer Vision models on embedded devices. However, static optimization techniques (e.g. pruning, quantization, etc.) neglect the fact that d...
175
Simone Scardapane @sscardapane.bsky.social · 15/01/2025
*Large Concept Models* by Barrault et al. Builds an autoregressive model in a "concept" space by wrapping the LLM in a pre-trained sentence embedder (also works with diffusion models). arxiv.org/abs/2412.08821
060
Reposted by Simone Scardapane
Bruno Neri @neribr.bsky.social · 15/01/2025
"Task Singular Vectors: Reducing Task Interference in Model Merging" by Antonio Andrea Gargiulo, @crisostomi.bsky.social , @mariasofiab.bsky.social , @sscardapane.bsky.social, Fabrizio Silvestri, Emanuele Rodolà Paper: arxiv.org/abs/2412.00081 Code: github.com/AntoAndGar/t... #machinelearning
042
Simone Scardapane @sscardapane.bsky.social · 14/01/2025
*Adaptive Length Image Tokenization via Recurrent Allocation* by @phillipisola.bsky.social et al. An encoder to compress an image into a sequence of 1D tokens whose length can dynamically vary depending on the specific image. arxiv.org/abs/2411.02393
020
Simone Scardapane @sscardapane.bsky.social · 14/01/2025
*Deep Learning Through A Telescoping Lens* by @alanjeffares.bsky.social @aliciacurth.bsky.social Shows that tracking 1st-order approximations to the training dynamics provides insights into many phenomena (e.g., double descent, grokking). arxiv.org/abs/2411.00247
0101
Simone Scardapane @sscardapane.bsky.social · 10/01/2025
*MoE Graph Transformers for Interpretable Particle Collision Detection* by @alessiodevoto.bsky.social @sgiagu.bsky.social et al. We propose a MoE graph transformer for particle collision analysis, with many nice interpretability insights (e.g., expert specialization). arxiv.org/abs/2501.03432
0124
Simone Scardapane @sscardapane.bsky.social · 09/01/2025
*A Meticulous Guide to Advances in Deep Learning Efficiency over the Years* by Alex Zhang Part deep learning history, part overview on the vast landscape of "efficiency" in DL (hardware, compilers, architecture, ...). Fantastic post! alexzhang13.github.io/blog/2024/ef...
1102
Reposted by Simone Scardapane
Fabio Montello @zioictus.bsky.social · 08/01/2025
First little project of the year: an awesome collection of papers on Dynamic Neural Networks for Computer Vision and Sensor Fusion! Each paper comes with a brief summary and code link. 👉 github.com/DTU-PAS/awes...
github.com
GitHub - DTU-PAS/awesome-dynn-for-cv: Awesome collection of DyNN papers for Computer Vision and Sensor Fusion applications :sparkles:
Awesome collection of DyNN papers for Computer Vision and Sensor Fusion applications :sparkles: - DTU-PAS/awesome-dynn-for-cv
282
Reposted by Simone Scardapane
Donato Crisostomi ✈️ NeurIPS @crisostomi.bsky.social · 08/01/2025
Don’t miss out on these insights and more — check out the paper! 📄 Preprint → arxiv.org/abs/2412.00081 💻 Code → github.com/AntoAndGar/t... Joint work w/ Antonio A. Gargiulo, @mariasofiab.bsky.social, @sscardapane.bsky.social, Fabrizio Silvestri, Emanuele Rodolà. (6/6)
arxiv.org
Task Singular Vectors: Reducing Task Interference in Model Merging
Task Arithmetic has emerged as a simple yet effective method to merge models without additional training. However, by treating entire networks as flat parameter vectors, it overlooks key structural in...
022
Simone Scardapane @sscardapane.bsky.social · 03/01/2025
*Modular Duality in Deep Learning* Develops a theory of "modular duality" for designing principled optimizers that respect the "type semantics" of each layer. arxiv.org/abs/2410.21265
010
Simone Scardapane @sscardapane.bsky.social · 28/12/2024
*Understanding Visual Feature Reliance through the Lens of Complexity* by @thomasfel.bsky.social @louisbethune.bsky.social @lampinen.bsky.social Wonderful work! They rank features' complexity with a variant of mutual information, before analyzing their dynamics. arxiv.org/abs/2407.06076
0212
Reposted by Simone Scardapane
Alessio Devoto @alessiodevoto.bsky.social · 19/12/2024
In Vision & Audio transformers, not all tokens need the same compute resources! We propose “modular learners” to control compute at token-level granularity (MHA & MLP): hard tokens get more, easy ones get less! w/ @sscardapane.bsky.social @neuralnoise.com @bartoszWojcik Soon #AAAI25 Link 👇
1193
Simone Scardapane @sscardapane.bsky.social · 19/12/2024
*Pooling in graph neural networks* My friend FM Bianchi made an awesome introduction to GNNs and pooling techniques over graphs, full of nice visuals and details! 🔥 gnn-pooling.notion.site/1-3-pooling-...
0104
Reposted by Simone Scardapane
Timothy O'Leary @timothyoleary.bsky.social · 17/12/2024
Apropos of never ending discussions about whether ANNs are "good" models of the nervous system, here is a slide I present to masters students showing a network that is found in motor control circuits *across phyla* (that's pretty ubiquitous!) I ask them to guess what it does...
616663
Simone Scardapane @sscardapane.bsky.social · 18/12/2024
*Relaxed Recursive Transformers* by @talschuster.bsky.social et al. Converts pre-trained transformers to a more efficient version by turning blocks of layers into a single layer which is iterated. Lots of interesting tricks! arxiv.org/abs/2410.20672
152
Simone Scardapane @sscardapane.bsky.social · 13/12/2024
*Relaxed Equivariance via Multitask Learning* by @tkrusch.bsky.social @mmbronstein.bsky.social They propose a regularization approach for exploiting symmetries over data (penalizing variable predictions over augmented data). arxiv.org/abs/2410.17878
1353
Simone Scardapane @sscardapane.bsky.social · 13/12/2024
*Rethinking Softmax: Self-Attention with Polynomial Activations* They show that using softmax in the attention computation upper-bounds the Frobenius norm of the attention matrix, and similar results can be obtained with a polynomial normalization. arxiv.org/abs/2410.18613
071
Simone Scardapane @sscardapane.bsky.social · 11/12/2024
*Adaptive Computation Modules: Granular Conditional Computation For Efficient Inference* with @alessiodevoto.bsky.social @neuralnoise.com Happy to share our work on distilling efficient transformers with dynamic modules' activation was accepted at #AAAI2025. 🔥 arxiv.org/abs/2312.10193
0102
Simone Scardapane @sscardapane.bsky.social · 09/12/2024
*Linear Algebra via Exterior Products* by Sergei Winitzki By far one of the best coordinate-free linear algebra books out there! Full of intuitions, great notation, and the PDF is free. 🙃 github.com/winitzki/lin...
github.com
GitHub - winitzki/linear-algebra-book: The full source code and hyperlinked PDF of the book "Linear Algebra via Exterior Products" (2010)
The full source code and hyperlinked PDF of the book "Linear Algebra via Exterior Products" (2010) - winitzki/linear-algebra-book
0134
Reposted by Simone Scardapane
Valerio Marsocci @valeriomarsocci.bsky.social · 06/12/2024
🚀🚀🌏 Are geospatial foundation models really impactful? Check it in our new pre-print! Welcome to **PANGAEA: a global and inclusive benchmark for GFMs** arxiv.org/abs/2412.04204 Check also the public GitHub repo (other news/updates soon): github.com/VMarsocci/pa... a short thread 🧵
2104
Simone Scardapane @sscardapane.bsky.social · 06/12/2024
*Sparse Crosscoders for Cross-Layer Features and Model Diffing* by @colah.bsky.social @anthropic.com Investigates stability & dynamics of "interpretable features" with cross-layers SAEs. Can also be used to investigate differences in fine-tuned models. transformer-circuits.pub/2024/crossco...
084
Reposted by Simone Scardapane
Donato Crisostomi ✈️ NeurIPS @crisostomi.bsky.social · 05/12/2024
First blue post (still have to figure out how tweets are called here) 💡idea: we consider task vectors at the layer level and reduce task interference by decorrelating the task-specific singular vectors of any matrix-structured layer 🔬results: large-margin improvements across all vision benchmarks
072
Simone Scardapane @sscardapane.bsky.social · 04/12/2024
*Task Singular Vectors: Reducing Task Interference in Model Merging* We show that task vectors are inherently low-rank, and we propose a merging method that significantly improves SOTA. arxiv.org/abs/2412.00081
0172
Simone Scardapane @sscardapane.bsky.social · 29/11/2024
*A Journey into PyTorch, the Ecosystem, and Deep Learning Compilers* Great seminar by @lantiga.bsky.social on the compiler's ecosystem and the recent Thunder framework by LightningAI! I have uploaded slides, video, & notebooks publicly. 🙂 www.sscardapane.it/seminars/pyt...
0386
Simone Scardapane @sscardapane.bsky.social · 28/11/2024
*Duo-LLM: A Framework for Studying Adaptive Computation in LLMs* Learns for each token & layer whether to use the original block or a smaller one, based on an overall compute budget. Interesting analysis based on an "oracle routing" on the relative importance of layers. arxiv.org/abs/2410.10846
021
Simone Scardapane @sscardapane.bsky.social · 27/11/2024
*Automatically Interpreting Millions of Features in LLMs* by @norabelrose.bsky.social et al. An open-source pipeline for finding interpretable features in LLMs with sparse autoencoders and automated explainability methods from @eleutherai.bsky.social. arxiv.org/abs/2410.13928
0276
Simone Scardapane @sscardapane.bsky.social · 25/11/2024
*Text classification with 1D CNNs in Equinox* 2nd part of our Equinox lab! We implement a simple 1D convolutional network with a pre-trained tokenizer and trainable embeddings - I am quite happy about the outcome! 🙂 colab.research.google.com/drive/1ik7cT...
050
Simone Scardapane @sscardapane.bsky.social · 22/11/2024
*Building neural networks in Equinox* Third lab for the course, after JAX and Keras we see Equinox (from @patrickkidger.bsky.social) - including callable pytrees, stateful modules, and filtering. Stay tuned for part b with a text classification model. 😎 colab.research.google.com/drive/19NHkb...
0192
Reposted by Simone Scardapane
nathan @nds.bsky.social · 21/11/2024
here’s a start! (no pun intended) go.bsky.app/2E66DfJ
021
Simone Scardapane @sscardapane.bsky.social · 20/11/2024
I fell on the keyboard and bought a bunch of physics textbooks (and one intruder). 😅 With great suggestions from @sgiagu.bsky.social @wellingmax.bsky.social
292
Simone Scardapane @sscardapane.bsky.social · 20/11/2024
*JAX - Why is Everyone So Excited About This Framework* by Yashovardhan Srivastava Detailed overview of a toy reimplementation of the core JAX engine (similar to autodidax) - a bit rough at times but truly interesting if you want to understand the framework better! yash-sri.xyz/blogs
0141