Sign in

Vaishnavh Nagarajan

@vaishnavh.bsky.social
3.4K followers 387 following 237 posts

Foundations of AI. I like simple and minimal examples and creative ideas. I also like thinking about the next token 🧮🧸 Google | PhD, CMU | arxiv.org/abs/2504.15266 | arxiv.org/abs/2403.06963 vaishnavh.github.io

PostsRepliesMedia
Vaishnavh Nagarajan @vaishnavh.bsky.social · 25/08/2026
Loved this analogy between how we perceive a scientific discovery that takes its full form from nascent glimpses and little elements of that idea and how an orchestra is set up and perceived. From Loren Eiseley's The Firmament of Time.
screenshot of paragraph from book
020
Vaishnavh Nagarajan @vaishnavh.bsky.social · 25/06/2026
In new post, I write about how we misunderstand the famous lottery ticket hypothesis. We tend to think that a network needs to be exponentially large for there to be a subnetwork to win the lottery. This is incorrect---a small network suffices! I use a dart throwing analogy to make sense of this.
1152
Vaishnavh Nagarajan @vaishnavh.bsky.social · 19/06/2026
In my next blogpost, I write about how I view technical communication: it's like trying to communicate an escape route to someone without a map but with a catch: you're not with them. You only have a walkie-talkie. Also, they're in panic.
130
Vaishnavh Nagarajan @vaishnavh.bsky.social · 14/04/2026
Consolidated my armchair thoughts about "how may an LLM (not) differ from a human who thinks in images/text?". I split this as 2 qns: - is text sufficient to be correct about say, a circle? - does correctness imply sharing the same "understanding" as humans? vaishnavh.github.io/blog/what-ll...
120
Vaishnavh Nagarajan @vaishnavh.bsky.social · 19/03/2026
In practice, scientists do not seem to behave like "neutral, rational agents" but rather behave like "zealous advocates" for an idea that has "hired them". I wrote about how I think this "courtroom" view of science works and what I learned from it! vaishnavh.github.io/blog/emotion...
1252
Vaishnavh Nagarajan @vaishnavh.bsky.social · 18/03/2026
Also, what's the catch with punishing bad reviews by preventing future submissions? Say: if your reviews are egregiously bad as flagged by multiple ACs across at least two conferences, you won't be able to submit papers to the next N conferences. (Possible that I'm missing something here.)
100
Vaishnavh Nagarajan @vaishnavh.bsky.social · 18/03/2026
Curious why conferences don't have a system where the authors of every paper together guarantee N reviews per paper (and they can distribute the load amongst themselves). This way wouldn't we tax authors in proportion to the number of papers they burden the system with?
110
Vaishnavh Nagarajan @vaishnavh.bsky.social · 03/03/2026
A recent paper (arxiv.org/abs/2602.18671) made me question something basic: do the logits of a language model model the next-token or the full sequence distribution? It really messed with my brain (in a fun way!). I wrote about the paper to clarify my thinking. vaishnavh.github.io/blog/joint-o...
vaishnavh.github.io
What does a language model model? - Vaishnavh Nagarajan
TL;DR: Does the next-token logit track the conditional or the joint probability of the whole sequence?I had an invisi...
1133
Vaishnavh Nagarajan @vaishnavh.bsky.social · 19/02/2026
Really liked this paper which ties up two observations that are equally mindboggling (low-rank logits & subliminal/weird generalization effects) and presents one other such observation arxiv.org/abs/2602.04863
arxiv.org
Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
Training modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop techniques to understa...
2216
Reposted by Vaishnavh Nagarajan
Declan Campbell @thisisadax.bsky.social · 05/02/2026
The visual world is composed of objects, and those objects are composed of features. But do VLMs exploit this compositional structure when processing multi-object scenes? In our 🆒🆕 #ICLR2026 paper, we find they do – via emergent symbolic mechanisms for visual binding. 🧵👇
18326
Vaishnavh Nagarajan @vaishnavh.bsky.social · 13/01/2026
Currently reading "a mathematician's apology" by GH Hardy. This is excerpt the foreword by CP Snow describing Hardy's personality and his work:
1131
Reposted by Vaishnavh Nagarajan
Vaishnavh Nagarajan @vaishnavh.bsky.social · 08/01/2026
in associative memory, the latent space doesn't really encode any interesting distance. imagine you're trying to store which countries share borders. you could simply write down a list of adjacent countries OR you could visualize the world map in your head. this is "associative" vs "geometric".
111
Vaishnavh Nagarajan @vaishnavh.bsky.social · 12/01/2026
fascinating!
010
Vaishnavh Nagarajan @vaishnavh.bsky.social · 09/01/2026
Rare to see such long term efforts these days 🫡
0141
Reposted by Vaishnavh Nagarajan
Andrew Gordon Wilson @andrewgwils.bsky.social · 07/01/2026
We introduce epiplexity, a new measure of information that provides a foundation for how to select, generate, or transform data for learning systems. We have been working on this for almost 2 years, and I cannot contain my excitement! arxiv.org/abs/2601.03220 1/7
814234
Reposted by Vaishnavh Nagarajan
Jeff Dean @jeffdean.bsky.social · 07/01/2026
Please welcome Google's Open Source efforts to Blue Sky at @opensource.google!
724740
Vaishnavh Nagarajan @vaishnavh.bsky.social · 08/01/2026
1/ We found that deep sequence models memorize atomic facts "geometrically" -- not as an associative lookup table as often imagined. This opens up practical questions on reasoning/memory/discovery, and also poses a theoretical "memorization puzzle."
312833
Vaishnavh Nagarajan @vaishnavh.bsky.social · 25/12/2025
If X, Y, Z are iid high-dim Gaussian N(0, I), what's the angle between X-Y and Z-Y? A. Concentrates at 90 deg B. Concentrates, NOT at 90° C. Doesn't concentrate anywhere. My (and most people's) instincts got this wrong! vaishnavh.github.io/blog/high-di...
vaishnavh.github.io
Angles between high-dimensional vectors - Vaishnavh Nagarajan
Switch off your brain and answer this:Given three points $\mathbf{X}, \mathbf{Y}, \mathbf{Z}$ sampled from a high-dim...
160
Reposted by Vaishnavh Nagarajan
CMU Computer Science Department @csdatcmu.bsky.social · 17/07/2025
Congratulations to CSD faculty Aditi Raghunathan and her research collaborators on receiving an ICML Outstanding Paper award for Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction (icml.cc/virtual/2025...). Paper: arxiv.org/abs/2504.15266
icml.cc
ICML 2025 AwardsICML 2025
061
Reposted by Vaishnavh Nagarajan
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 04/07/2025
Reading the dedications of a PhD thesis is often a cure for a bad day. There’s so much affection in them
4431
Reposted by Vaishnavh Nagarajan
Ahmad Beirami @abeirami.bsky.social · 02/07/2025
As NeurIPS review deadline is around the corner, please remember that you cannot use any non-local LLM like chatgpt/gemini for understanding the paper and drafting/revising your review as that breaks the confidentiality agreement. NeurIPS 2025 Official LLM Policy: neurips.cc/Conferences/...
neurips.cc
LLM Policy
061
Reposted by Vaishnavh Nagarajan
Giovanni Toffetti @gtof.eurosky.social · 22/06/2025
I really enjoyed "When We Cease to Understand the World", although it's more fiction than history of science
021
Reposted by Vaishnavh Nagarajan
Javier Burroni @jburroni.bsky.social · 23/06/2025
“Science in history” by Bernal is my first recommendation. The work of Ian Hacking is a good recommendation for Probability
011
Reposted by Vaishnavh Nagarajan
Alexandra Proca @aproca.bsky.social · 20/06/2025
How do task dynamics impact learning in networks with internal dynamics? Excited to share our ICML Oral paper on learning dynamics in linear RNNs! with @clementinedomine.bsky.social @mpshanahan.bsky.social and Pedro Mediano openreview.net/forum?id=KGO...
openreview.net
Learning dynamics in linear recurrent neural networks
Recurrent neural networks (RNNs) are powerful models used widely in both machine learning and neuroscience to learn tasks with temporal dependencies and to model neural dynamics. However, despite...
13411
Reposted by Vaishnavh Nagarajan
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 22/06/2025
When we are doing science, we are unknowingly executing our mythology, taken from movies and friends and textbooks, of what science is. History of science helps us ground that myth in reality
3536
Vaishnavh Nagarajan @vaishnavh.bsky.social · 13/06/2025
I finally wrote a full-fledged blog about this: reading the history of science is an **amazing** yet under-recognized way to develop (emotional) maturity as a researcher. If you have thoughts/recommendations, please share! vaishnavh.github.io/2025/04/29/h...
vaishnavh.github.io
Why PhD students should read the history of science - Vaishnavh Nagarajan
Here’s a secret that I accidentally discovered during my PhD: consuming history-of-science content is an efficient wa...
3357
Reposted by Vaishnavh Nagarajan
Tiago Pimentel @tpimentel.bsky.social · 04/06/2025
A string may get 17 times less probability if tokenised as two symbols (e.g., ⟨he, llo⟩) than as one (e.g., ⟨hello⟩)—by an LM trained from scratch in each situation! Our new ACL paper proposes an observational method to estimate this causal effect! Longer thread soon!
Title of paper "Causal Estimation of Tokenisation Bias" and schematic of how we define tokenisation bias, which is the causal effect we are interested in.
1518
Reposted by Vaishnavh Nagarajan
Andrew Saxe @saxelab.bsky.social · 04/06/2025
How does in-context learning emerge in attention models during gradient descent training? Sharing our new Spotlight paper @icmlconf.bsky.social: Training Dynamics of In-Context Learning in Linear Attention arxiv.org/abs/2501.16265 Led by Yedi Zhang with @aaditya6284.bsky.social and Peter Latham
15318
Reposted by Vaishnavh Nagarajan
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 02/06/2025
This paper is quite nice. It mixes some useful toy models of creativity with insights about how to induce more creativity in LLMs that are better than greedy sampling
1164
Vaishnavh Nagarajan @vaishnavh.bsky.social · 02/06/2025
📢 New #paper on creativity & multi-token prediction! We design minimal open-ended tasks to argue: → LLMs are limited in creativity as they learn to predict the next token → creativity can be improved via multi-token learning & injecting noise ("seed-conditioning" 🌱) 1/ #MLSky #AI #arxiv 🧵👇🏽
1272
Reposted by Vaishnavh Nagarajan
Nathan Lambert @natolambert.bsky.social · 27/05/2025
This isn't fake news. One of the craziest AI research papers I've been on in a while. Weird ablations on RLVR shows that the Qwen 2.5 models can learn with literally random rewards, likely due to some funkiness in mid-training and the GRPO setup.
4476
Vaishnavh Nagarajan @vaishnavh.bsky.social · 27/05/2025
can someone reconcile these two contradictory findings?! two papers find entropy *minimization*/confidence maximization helps performance, and the RL-on-one-sample finds entropy maximization/increasing exploration alone helps performance?!
572
Reposted by Vaishnavh Nagarajan
Core Francisco Parkg @corefpark.bsky.social · 21/05/2025
🚨 New Paper! A lot happens in the world every day—how can we update LLMs with belief-changing news? We introduce a new dataset "New News" and systematically study knowledge integration via System-2 Fine-Tuning (Sys2-FT). 1/n
161
Vaishnavh Nagarajan @vaishnavh.bsky.social · 21/05/2025
The contextual shadowing effect in this paper is interesting & reminds me of the infamous "cow-in-a-beach" spurious correlation examples. The model needs to learn a core feature (here. the store-in-memory feature) but it instead relies on a simpler "spurious" feature (the in-context feature).
111
Vaishnavh Nagarajan @vaishnavh.bsky.social · 18/05/2025
My #neurips bidding pool is REALLY bad and I see that I'm not alone. Was the pool restricted to 100 papers to break collusion? if so, seems like everyone else is going to pay for this
240
Vaishnavh Nagarajan @vaishnavh.bsky.social · 08/05/2025
looks like a fun workshop!
010
Reposted by Vaishnavh Nagarajan
Ahmad Beirami @abeirami.bsky.social · 23/04/2025
Excited that our paper "safety alignment should be made more than just a few tokens deep" was recognized as an #ICLR2025 Outstanding Paper! We identified a common root cause to many safety vulnerabilities and pointed out some paths forward to address it!
2323
Reposted by Vaishnavh Nagarajan
Nikhil Garg @nkgarg.bsky.social · 07/05/2025
At the moment, we just show you posts by people you follow. Following someone will make their paper posts appear! In the future, we'll expand this (slightly) to show you other paper posts that may be of interest
241
Reposted by Vaishnavh Nagarajan
William B. Fuckley @opinionhaver.bsky.social · 05/05/2025
Yeah this is actual unironic advice for any 1st years: you seriously need to go get beers or otherwise hangout socially with more senior grad students (no one else will really know) and hopefully get ensconced enough that someone will tell you “hey just fyi Dr. GoodPubs is low-key a total sociopath”
633333
Reposted by Vaishnavh Nagarajan
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 06/05/2025
@vaishnavh.bsky.social and crew do it again: - a benchmark for open-ended creativity - a demonstration of challenges of next-token prediction - a technique to improve transformer randomness through inputs not sampling arxiv.org/abs/2504.15266
arxiv.org
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
We design a suite of minimal algorithmic tasks that are a loose abstraction of open-ended real-world tasks. This allows us to cleanly and controllably quantify the creative limits of the present-day l...
1376
Vaishnavh Nagarajan @vaishnavh.bsky.social · 06/05/2025
@icmlconf.bsky.social (assuming this is the legit ICML account), the Canadian visa application requires providing the inviting contact to have a Canadian address. But the one in the visa letter is a US address. Could you look into this?
000
Reposted by Vaishnavh Nagarajan
Ahmad Beirami @abeirami.bsky.social · 03/05/2025
If you are at #AISTATS2025 and are interested in concept erasure, talk to @somnathbrc.bsky.social at Poster Session 1 on Saturday May 3.
0164
Reposted by Vaishnavh Nagarajan
NeurIPS Conference @neuripsconf.bsky.social · 02/05/2025
Responsible reviewing initiatives for NeurIPS 2025 - read more about changes to reviewing that that will safeguard reviewing quality and timeline in our blog post below: blog.neurips.cc/2025/05/02/r...
blog.neurips.cc
Responsible Reviewing Initiative for NeurIPS 2025 – NeurIPS Blog
Communications Chairs 2025 2021 Conference
1255
Reposted by Vaishnavh Nagarajan
Dimitris Papailiopoulos @dimitrisp.bsky.social · 02/02/2025
Self-improving Transformers can overcome easy-to-hard and length generalization challenges. Paper on arxiv coming on Monday. Link to a talk I gave on this below 👇 Super excited about this work! Talk : youtube.com/watch?v=szhE... slides: tinyurl.com/SelfImprovem...
0151
Reposted by Vaishnavh Nagarajan
Gautam Kamath @gautamkamath.com · 01/05/2025
I wrote a post on how to connect with people (i.e., make friends) at CS conferences. These events can be intimidating so here's some suggestions on how to navigate them I'm late for #ICLR2025 #NAACL2025, but in time for #AISTATS2025 #ICML2025! 1/3 kamathematics.wordpress.com/2025/05/01/t...
kamathematics.wordpress.com
Tips on How to Connect at Academic Conferences
I was a kinda awkward teenager. If you are a CS researcher reading this post, then chances are, you were too. How to navigate social situations and make friends is not always intuitive, and has to …
36618
Reposted by Vaishnavh Nagarajan
Danica Sutherland @djsutherland.ml · 23/04/2025
Thrilled to announce that Joshua’s paper won one of three Outstanding Paper awards at ICLR. Come to the poster on Friday afternoon (#376 in Hall 3) or the talk on Saturday (4:30 in Hall 1), and while you’re at it snag him for a postdoc!
1366
Reposted by Vaishnavh Nagarajan
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 14/04/2025
Been waiting for a work making this clear for a while: zero-shot, LLMs do not hedge and do not ask questions to solve under-specified problems arxiv.org/abs/2503.22674
arxiv.org
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
Recently, a large amount of work has focused on improving large language models' (LLMs') performance on reasoning benchmarks such as math and logic. However, past work has largely assumed that tasks a...
38122
Vaishnavh Nagarajan @vaishnavh.bsky.social · 03/04/2025
What tools do people use to manage AC/meta-reviewing responsibilities (like consolidating key points, flagging which papers/reviews need attention, etc., etc.,)?
000
Reposted by Vaishnavh Nagarajan
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
o3 can't multiply beyond a few digits... But I think multiplication, addition, maze solving and easy-to-hard generalization is actually solvable on standard transformers... with recursive self-improvement Below is the acc of a tiny model teaching itself how to add and multiply
2214
Reposted by Vaishnavh Nagarajan
Jacob Springer @jacobspringer.bsky.social · 26/03/2025
Training with more data = better LLMs, right? 🚨 False! Scaling language models by adding more pre-training data can decrease your performance after post-training! Introducing "catastrophic overtraining." 🥁🧵👇 arxiv.org/abs/2503.19206 1/10
13314