Sign in

Chandar Research Lab

@chandar-lab.bsky.social
10 followers 37 following 59 posts

Sarath Chandar's research group at @polymtl , @UMontreal and @Mila_Quebec focusing on Machine Learning!

PostsRepliesMedia
Chandar Research Lab @chandar-lab.bsky.social · 21/04/2026
Sugar-shack with the lab!! 🍁 Saying goodbye to a very long winter, and welcoming sunnier days 🤩🍃
000
Chandar Research Lab @chandar-lab.bsky.social · 17/03/2026
🗣️ Shoutout to the authors: Pranshu Malviya, Balaraman Ravindran and @sarath-chandar.bsky.social!!! (published at CoLLAs 2022). 🔗 Learn more at: lnkd.in/eSSd9m56
lnkd.in
LinkedIn
This link will take you to a page that’s not on LinkedIn
001
Chandar Research Lab @chandar-lab.bsky.social · 17/03/2026
This was the first work to show that you can successfully use adaptive gradient optimizers for lifelong learning and still beat Stochastic Gradient Descent (i.e., RMSProp < SGD < TAG-RMSProp!). 🔥 Across benchmarks, TAG was shown to improve final accuracy over baselines like ER and A-GEM.
100
Chandar Research Lab @chandar-lab.bsky.social · 17/03/2026
🔥 TAG tracks gradient traces from previously learned tasks and estimates task similarity: ⬇️ Lower α → related tasks → transfer is encouraged ⬆️ Higher α → conflicting tasks → reduce interference
100
Chandar Research Lab @chandar-lab.bsky.social · 17/03/2026
In lifelong learning, acquiring new tasks can cause ML models to forget previously learned knowledge. For this, our lab introduced TAG (Task-based Accumulated Gradients), a general wrapper on top of adaptive gradient optimizers. 📈
100
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
As we hope for women to strive in research every year more than the last, we encourage all of them to apply to our lab for internships, Master’s or PhD degrees with @sarath-chandar.bsky.social !
000
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
The Chandar Research Lab remains committed to supporting women and other underrepresented communities @mila-quebec.bsky.social and in ML with initiatives such as the graduate application assistance program or a Computer Science summer school for high school students goint to undergrad.👩‍🔬
100
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
This is only a sneak peek into the actual work they did last year, as much of their research is still under submission. Stay tuned for more interesting papers spanning ML for Biology, model merging, continual learning, etc...
100
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
Generalization Can Emerge in Tabular Foundation Models From a Single Table by Nour Shaheen at the AI for Tabular Data workshop @euripsconf.bsky.social 2025! arxiv.org/abs/2511.09665
arxiv.org
Generalization Can Emerge in Tabular Foundation Models From a Single Table
Deep tabular modelling increasingly relies on in-context learning where, during inference, a model receives a set of $(x,y)$ pairs as context and predicts labels for new inputs without weight updates....
100
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
The Expressive Limits of Diagonal SSMs for State-Tracking by Behnoush Khavari @iclr-conf.bsky.social 2026. iclr.cc/virtual/2026...
iclr.cc
ICLR Poster The Expressive Limits of Diagonal SSMs for State-TrackingICLR 2026
100
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
NeoBERT: A Next Generation BERT by @lola-le-breton.bsky.social published @tmlr-pub.bsky.social and @iclr-conf.bsky.social in Rio this year. arxiv.org/abs/2502.19587
arxiv.org
NeoBERT: A Next-Generation BERT
Recent innovations in architecture, pre-training, and fine-tuning have led to the remarkable in-context learning and reasoning abilities of large auto-regressive language models such as LLaMA and Deep...
100
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models by Istabrak Abbes @collasconf.bsky.social arxiv.org/abs/2508.01908
arxiv.org
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
Training large language models (LLMs) typically involves pre-training on massive corpora, only to restart the process entirely when new data becomes available. A more efficient and resource-conserving...
110
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
📜 Small Encoders Can Rival Large Decoders in Detecting Groundedness by Istabrak Abbes published @aclmeeting.bsky.social 2025. aclanthology.org/2025.finding...
aclanthology.org
Small Encoders Can Rival Large Decoders in Detecting Groundedness
Istabrak Abbes, Gabriele Prato, Quentin Fournier, Fernando Rodriguez, Alaa Boukhary, Adam Elwood, Sarath Chandar. Findings of the Association for Computational Linguistics: ACL 2025. 2025.
100
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
Maryam Hashemzadeh, @lola-le-breton.bsky.social, Istabrak Abbes, Nour Shaheen, Behnoush Khavari, Anabel Tan and @katelobacheva.bsky.social. Give them a follow and look at this list of their publications with our lab in the past year!⬇️
100
Chandar Research Lab @chandar-lab.bsky.social · 10/03/2026
This week, as we celebrated International Women’s Right Day for the 115th time on Sunday, the Chandar Lab wanted to pay tribute to all the amazing women doing research👩‍🎓, and to highlight the cutting-edge work they do at our lab everyday...🧵
122
Chandar Research Lab @chandar-lab.bsky.social · 24/02/2026
Work done by @nilaksh404.bsky.social, Antoine Clavaud, @mreymond.bsky.social, Francois Rivest, and @sarath-chandar.bsky.social Checkout the paper at : arxiv.org/abs/2602.09396 Code : github.com/chandar-lab/...
github.com
GitHub - chandar-lab/stream-rep-rl: Streaming setup with representation learning for RL
Streaming setup with representation learning for RL - chandar-lab/stream-rep-rl
010
Chandar Research Lab @chandar-lab.bsky.social · 24/02/2026
Look at the latents! t-SNE analysis shows that our method (top) learns structured, temporally coherent representations faster than standard streaming RL
110
Chandar Research Lab @chandar-lab.bsky.social · 24/02/2026
Our method systematically outperforms existing baselines across Atari, MinAtar, and Octax. The best part? It remains efficient enough to train on just a few CPU cores.
100
Chandar Research Lab @chandar-lab.bsky.social · 24/02/2026
Streaming data is highly correlated, which usually causes poor training. To fix this, we introduced Orthogonal Gradient Updates. By projecting gradients onto a subspace orthogonal to their history, we keep learning stable and effective.
110
Chandar Research Lab @chandar-lab.bsky.social · 24/02/2026
We bring Self-Predictive Representations (SPR) to the streaming pipeline. By predicting future latent states, we force the encoder to learn much richer features from every observed frame: without needing a massive memory footprint of a replay buffer.
110
Chandar Research Lab @chandar-lab.bsky.social · 24/02/2026
Without a replay buffer, streaming agents struggle to build meaningful representations. Traditional value-based losses alone can’t exploit the full informational content of transient data before it's gone.
110
Chandar Research Lab @chandar-lab.bsky.social · 24/02/2026
Streaming Reinforcement Learning (RL) is a huge challenge: transitions are used once and discarded immediately. This makes agents extremely sample-inefficient. But what if we could "squeeze" more information out of every single frame? Check out our latest paper!
133
Chandar Research Lab @chandar-lab.bsky.social · 17/02/2026
Shoutout to the authors: Kamran Chitsaz, Milad Aghajohari, @a-kazemnejad.bsky.social. Supervised by: @sarath-chandar@bsky.social, @murefil.bsky.social, AaronCourville and @sivareddyg.bsky.social 🔗 Learn more at: arxiv.org/abs/2510.06557 🔗Build with: github.com/McGill-NLP/the-markovian-thinker
arxiv.org
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Reinforcement learning (RL) has recently become a strong recipe for training reasoning LLMs that produce long chains of thought (LongCoT). Yet the standard RL "thinking environment", where the state i...
000
Chandar Research Lab @chandar-lab.bsky.social · 17/02/2026
🧩 Even state-of-the-art models show Markovian Thinking at zero-shot: both GPT-oss-120B and Qwen3-30B-A3B recover/track LongCoT with no special prompting/training required, and lots of in-distribution positives on initialization, so RL with Delethink is primed to scale!!
100
Chandar Research Lab @chandar-lab.bsky.social · 17/02/2026
🔥 Further, we scaled DeepSeek R1-1.5B to a thinking budget of 96K in 150 RL steps. Accuracy jumped, with mean trace lengths at around 40K tokens.
100
Chandar Research Lab @chandar-lab.bsky.social · 17/02/2026
Markovian Thinking is instantiated by Delethink, an RL enviroment. With it, we trained DeepSeek R1-1.5B and demonstrated: 1️⃣ The same scaling as LongCoT-RL, but at lower costs, 2️⃣ Better test-time scaling, improving past 24K tokens, while LongCoT-RL plateaus. 3️⃣ All this while keeping linear costs!!
100
Chandar Research Lab @chandar-lab.bsky.social · 17/02/2026
Markovian Thinking works by: 1️⃣ Making LLMs reason in 8K chunks. 2️⃣ At each boundary, context is reset and a small textual state from the last chunk is carried over. 🔃 Continues from that state. ✅ This decouples thinking length from context size, achieving linear compute and constant memory!
100
Chandar Research Lab @chandar-lab.bsky.social · 17/02/2026
‘The Markovian Thinker’, developed by our lab, has been accepted at @iclr-conf.bsky.social 

This work achieved long reasoning without the quadratic attention tax by making LLMs reason in chunks with a bounded state, achieving linear compute, constant memory and scaling beyond its training limits! 🔥
110
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
📝 openreview.net/forum?id=5bg... Joint work of Mehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh and @sarath-chandar.bsky.social @mila-quebec.bsky.social .
openreview.net
The Expressive Limits of Diagonal SSMs for State-Tracking
State-Space Models (SSMs) have recently been shown to achieve strong empirical performance on a variety of long-range sequence modeling tasks while remaining efficient and highly-parallelizable....
021
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
Takeaways for architecture design: - Diagonal structure imposes a precise group-theoretic ceiling on expressivity - Depth helps in a principled way (one layer per Abelian factor) - But training algorithms need to catch up — expressivity alone isn't enough
100
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
Interestingly, initializing near the analytical solution does help the model learns and generalizes. This suggests solutions sit in a basin of attraction, but training can't find it from random init. A very different failure mode from what's been observed for Transformers on similar tasks.
100
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
But there is a catch: expressivity ≠ learnability. In our experiments, multi-layer diagonal SSMs consistently fail to learn S₃ and A₄ with gradient-based optimization, even though solutions provably exist in the hypothesis class!
100
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
We give an explicit 2-layer diagonal SSM construction for S₃: the first layer tracks a C₂ parity automaton, the second tracks a C₃ rotation conditioned on the first layer's state — mirroring the semi-direct product decomposition.
100
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
What this means concretely: - Parity (C₂), modular counting (Cₙ): 1 layer suffices - Permutations of 3 elements (S₃): exactly 2 layers needed - S₄: 3 layers - A₅ (non-solvable): no number of diagonal layers will ever work - Rubik’s cube: same
100
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
Theorem 2: A k-layer Complex Diagonal SSM can track a group G ⟺ G has a subnormal series of length ≤ k with Abelian factor groups. This characterizes the expressivity of SSMs: depth lets you “peel off” one Abelian layer at a time.
110
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
Theorem 1: A single-layer Complex Diagonal SSM (input-dependent, complex-valued) can track a group if and only if that group is Abelian. Even with complex eigenvalues and a powerful decoder, diagonality forces commutativity. The bias term b(x) doesn't help either.
110
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
State-tracking is fundamental: tracking variables in code, game states, parse trees. It boils down to simulating group operations over sequences. Diagonal SSMs are popular for their efficiency, but how expressive are they really?
110
Chandar Research Lab @chandar-lab.bsky.social · 10/02/2026
New work, just accepted @ICLR: "The Expressive Limits of Diagonal SSMs for State-Tracking" We give a complete characterization of what diagonal SSMs can and cannot compute on state-tracking tasks and the answer is deeply connected to group theory. 🧵👇
122
Chandar Research Lab @chandar-lab.bsky.social · 03/02/2026
NeoBERT: A Next-Generation BERT (TMLR Journal-to-Conference Track) We modernized BERT (RoPE, SwiGLU, 4k context). At just 250M params, it outperforms RoBERTa and ModernBERT on the MTEB benchmark. 📄 arxiv.org/abs/2502.19587
arxiv.org
NeoBERT: A Next-Generation BERT
Recent innovations in architecture, pre-training, and fine-tuning have led to the remarkable in-context learning and reasoning abilities of large auto-regressive language models such as LLaMA and Deep...
031
Chandar Research Lab @chandar-lab.bsky.social · 03/02/2026
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning We achieve linear-complexity reasoning. Our "Delethink" decouples thought length from context, matching LongCoT performance with ≈25% of the compute. 📄 arxiv.org/abs/2510.06557
arxiv.org
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Reinforcement learning (RL) has recently become a strong recipe for training reasoning LLMs that produce long chains of thought (LongCoT). Yet the standard RL "thinking environment", where the state i...
111
Chandar Research Lab @chandar-lab.bsky.social · 03/02/2026
The Expressive Limits of Diagonal SSMs for State-Tracking We prove a tight bound: Diagonal SSMs are theoretically incapable of tracking non-Abelian groups. A critical look at where efficient models fail vs. where they succeed. 📄 openreview.net/forum?id=5bg...
openreview.net
The Expressive Limits of Diagonal SSMs for State-Tracking
State-Space Models (SSMs) have recently been shown to achieve strong empirical performance on a variety of long-range sequence modeling tasks while remaining efficient and highly-parallelizable....
111
Chandar Research Lab @chandar-lab.bsky.social · 03/02/2026
Excited to share that we have 3 papers accepted at #ICLR2026! 🇧🇷 Our work this year focuses on efficiency and expressivity: deriving theoretical limits for SSMs, achieving linear scaling for reasoning, and modernizing encoder architectures. A summary of our work 👇 🧵
121
Chandar Research Lab @chandar-lab.bsky.social · 27/01/2026
We believe this is a critical missing piece for the future of interactive agents. Explore the project: 🌐 Website: chandar-lab.github.io/hangman-webs... 💻 Code: github.com/chandar-lab/... 📄 Paper: arxiv.org/html/2601.06...
chandar-lab.github.io
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
000
Chandar Research Lab @chandar-lab.bsky.social · 27/01/2026
The Solution? A Private Working Memory. We propose an architecture with a "scratchpad" that: ✅ Persists across turns ✅ Is readable/writable by the agent ✅ Is NEVER shown to the user This Generative-Retention Loop allows agents to finally play Hangman correctly.
100
Chandar Research Lab @chandar-lab.bsky.social · 27/01/2026
Without being able to generate and retain intermediate reasoning traces, models face a dangerous Secrecy vs. Helpfulness trade-off. We found aligned models often leak the secret, as "helpfulness" overrides the game's hidden constraints.
100
Chandar Research Lab @chandar-lab.bsky.social · 27/01/2026
Semantic retrieval (like RAG) isn't enough because it is designed to find external or public knowledge. And while reasoning models generate intermediate thoughts, these traces are lost instantly.
100
Chandar Research Lab @chandar-lab.bsky.social · 27/01/2026
...Just like in Hangman, the model is structurally incapable of this. It will invent symptoms on the fly rather than adhering to a consistent, hidden truth.
110
Chandar Research Lab @chandar-lab.bsky.social · 27/01/2026
Why does this matter? It has real implications. Imagine being a med student trying to test your diagnostic capabilities. You ask your favourite LLM agent to simulate a patient with a specific medical condition, but to keep that condition hidden so you can diagnose it through questioning...
110
Chandar Research Lab @chandar-lab.bsky.social · 27/01/2026
We prove an Impossibility Theorem: Standard chat agents face a fundamental trade-off between Secrecy and Consistency. 1️⃣ If the secret is in the context, then it’s leaked (violates Secrecy). 2️⃣ If it’s NOT in the context, then the model forgets it (violates Consistency).
120
Chandar Research Lab @chandar-lab.bsky.social · 27/01/2026
Nowadays, LLMs can solve Math Olympiad problems and one-shot complex code. But they fail at "trivial" games like Hangman. Why? The standard "chat interface" is fundamentally broken for tasks that require keeping a secret (what we call Private State Interactive Tasks).
110