Reposted by Saurav JhaChandar Research Lab @chandar-lab.bsky.social · 24/02/2026Streaming Reinforcement Learning (RL) is a huge challenge: transitions are used once and discarded immediately. This makes agents extremely sample-inefficient. But what if we could "squeeze" more information out of every single frame? Check out our latest paper! 133
Reposted by Saurav JhaChandar Research Lab @chandar-lab.bsky.social · 10/02/2026New work, just accepted @ICLR: "The Expressive Limits of Diagonal SSMs for State-Tracking" We give a complete characterization of what diagonal SSMs can and cannot compute on state-tracking tasks and the answer is deeply connected to group theory. 🧵👇 122
Reposted by Saurav JhaChandar Research Lab @chandar-lab.bsky.social · 27/01/2026Can LLMs play Hangman? Spoiler alert: Not yet. Check out “LLMs Can’t Play Hangman: On the Necessity of a Private Working Memory for Language Agents”, led by Davide Baldelli, Ali Parviz, AmalZouaq and Sarath Chandar. 111
Reposted by Saurav JhaChandar Research Lab @chandar-lab.bsky.social · 20/01/2026Can LLMs become CAD designers? Check out “CADmium: Fine-Tuning Code Language Models for Text-Driven Sequential CAD Design”, which is now published in Transactions on Machine Learning Research (TMLR)! 131
Saurav Jha @saurav-jha.bsky.social · 03/10/2025Life update - last month I moved to #montreal 🇨🇦 from #Sydney 🇦🇺 to kick off my @ivado.bsky.social postdoc fellowship at @mila-quebec.bsky.social. Must say I am constantly amused by: 1. How walkable the city is. 2. How easy is it to reach out to diverse research communities within #mila ! 😀 030
Saurav Jha @saurav-jha.bsky.social · 22/01/2025🎉 Happy to share that our paper “Mining your own secrets: Diffusion Classifier scores for Continual Personalization of Text-to-Image Diffusion Models” has been accepted to #ICLR2025! 👉 The work results from my #Sony summer internship in the stunning #Tokyo🗼 city Preprint: arxiv.org/pdf/2410.00700 010
Saurav Jha @saurav-jha.bsky.social · 25/12/2024I ran across a busy Sander at a #neurips party with a similar question - he was still patient enough to explain stuff. This talk further clarifies a good amount of my doubts. Recommend watching if you're working on diffusion / LLMs for generation! 071