Sign in

Abhinav Moudgil

@amoudgl.bsky.social
181 followers 85 following 25 posts

PhD student, Mila abhinavmoudgil.com

PostsRepliesMedia
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Finally, I'd like to acknowledge the Google TRC program for the generous TPU support that made this research possible!
000
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Work done with Boris Knyazev and @ebelilov.bsky.social. If you are interested in this line of research or related topics, our lab is hiring: eugenium.github.io/Projects/pos...
eugenium.github.io
Research Openings
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
All paper links Celo2: arxiv.org/abs/2602.19142 Celo (its predecessor): arxiv.org/abs/2501.12670 I'll also be presenting these two papers at #ICLR26 main track in 3 weeks at Rio de Janeiro, Brazil 🇧🇷 feel free to reach out if you'd like to chat more in-person!
arxiv.org
Celo2: Towards Learned Optimization Free Lunch
Learned optimizers are powerful alternatives to hand-designed update rules like Adam, yet they have seen limited practical adoption since they often fail to meta-generalize beyond their training distr...
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
We release a single-file self-contained Optax implementation of Celo2 with meta-training support to facilitate future research: github.com/amoudgl/celo2
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Comparison with Muon: Our Celo2 meta-trained on small classification tasks performs slightly better but tbh doesn't justify Muon's replacement yet due to higher memory cost (3 extra per-param + sublinear adafactor accumulators). That said, Celo2 can improve with learning, so stay tuned!
comparing celo2 with muon
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Atari RL:
celo2 atari reinforcement learning results
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
ImageNet classification:
celo2 imagenet classification results
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Now on to results: Celo2 shows strong out-of-distribution performance and stable optimization behaviour in all the domains we tested on: language modeling (shown above), image classification and reinforcement learning. Recall that update rule is learned on just image classification tasks.
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Above are the core bits of our recipe. However, we do a lot of ablations to make it work well (task augmentation, MLP architecture, normalization forms, etc) from a generalization-first perspective. Check out our paper below for all the details!
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
(4) Modern optimization harness: we find that Celo2 is highly compatible with recent techniques like orthogonalization (Newton-Schulz), distinct 1D/2D update rules (Muon), decoupled weight decay, suggesting that future improvements in optimization may translate to learned optimizers as well!
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
This RMS step normalization also nicely ties our work with recent work such as Kimi-Muon that does the same RMS normalization of the update but with a different objective: to match empirical RMS of Adam and reduce tuning effort by having a shared learning rate across Adam and Muon.
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
(3) RMS normalized update: prior work normalizes MLP inputs but uses raw MLP outputs as the step update, which is prone to overfitting on meta-training tasks. We normalize the outputs too, forcing the MLP to learn a task-invariant update that can work on any task at test-time.
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Celo2 architecture is a small 2-layer MLP (<200 params) that works as a drop-in replacement for Adam but it can effectively absorb and utilize a lot more information from tensor + param levels, thus allowing us to add more compute/FLOPs in optimizer to improve performance.
celo2 vs adam
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
(2) Simple architecture that scales: We start with hierarchical architecture (tensor-level LSTMs with param-level MLPs) from VeLO and massively simplify it in our work (Celo, Celo2) to achieve high generalization performance.
celo2 architecture comparison with prior work
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
(1) Tunable step size: unlike prior work, our approach solely focuses on learning the update rule, thus decoupling step size tuning from update rule learning. This does result in extra tunable knob (learning rate) but we find that the learned rule generalizes surprisingly well to large-scale tasks.
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
We obtain a compute-efficient learned optimizer (or Celo) recipe by iterating over learned optimizer architecture and meta-training, pushing the generalization frontier while also maintaining performance. Our recipe is simple and consists of the following core components:
celo2 algorithm
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Note that this setting is quite ambitious! It is testing generalization from meta-training across all axes in terms of task diversity (dataset, model, objective), unroll length, and scale.
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Specifically, we meta-train on small image MLP classification tasks from Celo (our prior work) and stress test it on standard unseen tasks such as GPT-2/3 pretraining, ViT image classification, and Atari RL.
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
In our work, we treat generalization as a first class citizen and develop a recipe to improve generalization performance on unseen / out-of-distribution tasks.
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
A seminal work by Google brain team in this direction, VeLO, scaled meta-training to 4000 TPU months (!) but it fails to generalize to out-of-distribution tasks such as ones larger than 600M parameters, RL, finetuning, etc. arxiv.org/abs/2211.09760
arxiv.org
VeLO: Training Versatile Learned Optimizers by Scaling Up
While deep learning models have replaced hand-designed features across many domains, these models are still trained with hand-designed optimizers. In this work, we leverage the same scaling approach b...
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Quick recap: learned optimizers are a compelling alternative to hand-designed update rules since they can be improved by learning, compute and data but their practical applicability has been hindered by their inability to generalize beyond tasks in their meta-training distribution.
100
Abhinav Moudgil @amoudgl.bsky.social · 31/03/2026
Introducing Celo2: Towards Learned Optimization Free Lunch We show that learned optimizers can generalize to practical tasks like GPT-3 1.3B pretraining and several out-of-distribution vision/RL tasks from limited meta-training (~4.5 GPU hours)! 🧵
celo2 language model pretraining results
121
Abhinav Moudgil @amoudgl.bsky.social · 15/06/2025
This tool is especially useful in cases when evaluations are expensive (e.g. LM harness eval) and you want to track model performance during training.
010
Abhinav Moudgil @amoudgl.bsky.social · 15/06/2025
github: github.com/amoudgl/assa...
github.com
GitHub - amoudgl/assayer: Python RQ watchdog to automatically evaluate ML model checkpoints offline during training
Python RQ watchdog to automatically evaluate ML model checkpoints offline during training - amoudgl/assayer
120
Abhinav Moudgil @amoudgl.bsky.social · 15/06/2025
New side project! assayer: A simple Python-RQ based tool to automatically monitor and evaluate ML model checkpoints offline during training.
141
Reposted by Abhinav Moudgil
Aditi Mavalankar @aditimavalankar.bsky.social · 17/03/2025
Excited to share our recent work, AuPair, an inference-time technique that builds on the premise of in-context learning to improve LLM coding performance! arxiv.org/abs/2502.18487 🧵
arxiv.org
AuPair: Golden Example Pairs for Code Repair
Scaling up inference-time compute has proven to be a valuable strategy in improving the performance of Large Language Models (LLMs) without fine-tuning. An important task that can benefit from additio...
1124