Vaishnavh Nagarajan @vaishnavh.bsky.social · 25/08/2026Loved this analogy between how we perceive a scientific discovery that takes its full form from nascent glimpses and little elements of that idea and how an orchestra is set up and perceived. From Loren Eiseley's The Firmament of Time. 020
Vaishnavh Nagarajan @vaishnavh.bsky.social · 25/06/2026In new post, I write about how we misunderstand the famous lottery ticket hypothesis. We tend to think that a network needs to be exponentially large for there to be a subnetwork to win the lottery. This is incorrect---a small network suffices! I use a dart throwing analogy to make sense of this. 1152
Vaishnavh Nagarajan @vaishnavh.bsky.social · 19/06/2026In my next blogpost, I write about how I view technical communication: it's like trying to communicate an escape route to someone without a map but with a catch: you're not with them. You only have a walkie-talkie. Also, they're in panic. 130
Vaishnavh Nagarajan @vaishnavh.bsky.social · 14/04/2026Consolidated my armchair thoughts about "how may an LLM (not) differ from a human who thinks in images/text?". I split this as 2 qns: - is text sufficient to be correct about say, a circle? - does correctness imply sharing the same "understanding" as humans? vaishnavh.github.io/blog/what-ll... 120
Vaishnavh Nagarajan @vaishnavh.bsky.social · 19/03/2026In practice, scientists do not seem to behave like "neutral, rational agents" but rather behave like "zealous advocates" for an idea that has "hired them". I wrote about how I think this "courtroom" view of science works and what I learned from it! vaishnavh.github.io/blog/emotion... 1252
Vaishnavh Nagarajan @vaishnavh.bsky.social · 18/03/2026Also, what's the catch with punishing bad reviews by preventing future submissions? Say: if your reviews are egregiously bad as flagged by multiple ACs across at least two conferences, you won't be able to submit papers to the next N conferences. (Possible that I'm missing something here.) 100
Vaishnavh Nagarajan @vaishnavh.bsky.social · 18/03/2026Curious why conferences don't have a system where the authors of every paper together guarantee N reviews per paper (and they can distribute the load amongst themselves). This way wouldn't we tax authors in proportion to the number of papers they burden the system with? 110
Vaishnavh Nagarajan @vaishnavh.bsky.social · 03/03/2026A recent paper (arxiv.org/abs/2602.18671) made me question something basic: do the logits of a language model model the next-token or the full sequence distribution? It really messed with my brain (in a fun way!). I wrote about the paper to clarify my thinking. vaishnavh.github.io/blog/joint-o...vaishnavh.github.ioWhat does a language model model? - Vaishnavh NagarajanTL;DR: Does the next-token logit track the conditional or the joint probability of the whole sequence?I had an invisi... 1133
Vaishnavh Nagarajan @vaishnavh.bsky.social · 19/02/2026Really liked this paper which ties up two observations that are equally mindboggling (low-rank logits & subliminal/weird generalization effects) and presents one other such observation arxiv.org/abs/2602.04863arxiv.orgSubliminal Effects in Your Data: A General Mechanism via Log-LinearityTraining modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop techniques to understa... 2216
Reposted by Vaishnavh NagarajanDeclan Campbell @thisisadax.bsky.social · 05/02/2026The visual world is composed of objects, and those objects are composed of features. But do VLMs exploit this compositional structure when processing multi-object scenes? In our 🆒🆕 #ICLR2026 paper, we find they do – via emergent symbolic mechanisms for visual binding. 🧵👇 18326
Vaishnavh Nagarajan @vaishnavh.bsky.social · 13/01/2026Currently reading "a mathematician's apology" by GH Hardy. This is excerpt the foreword by CP Snow describing Hardy's personality and his work: 1131
Reposted by Vaishnavh NagarajanVaishnavh Nagarajan @vaishnavh.bsky.social · 08/01/2026in associative memory, the latent space doesn't really encode any interesting distance. imagine you're trying to store which countries share borders. you could simply write down a list of adjacent countries OR you could visualize the world map in your head. this is "associative" vs "geometric". 111
Vaishnavh Nagarajan @vaishnavh.bsky.social · 09/01/2026Rare to see such long term efforts these days 🫡 0141
Reposted by Vaishnavh NagarajanAndrew Gordon Wilson @andrewgwils.bsky.social · 07/01/2026We introduce epiplexity, a new measure of information that provides a foundation for how to select, generate, or transform data for learning systems. We have been working on this for almost 2 years, and I cannot contain my excitement! arxiv.org/abs/2601.03220 1/7 814234
Reposted by Vaishnavh NagarajanJeff Dean @jeffdean.bsky.social · 07/01/2026Please welcome Google's Open Source efforts to Blue Sky at @opensource.google! 724740
Vaishnavh Nagarajan @vaishnavh.bsky.social · 08/01/20261/ We found that deep sequence models memorize atomic facts "geometrically" -- not as an associative lookup table as often imagined. This opens up practical questions on reasoning/memory/discovery, and also poses a theoretical "memorization puzzle." 312833
Vaishnavh Nagarajan @vaishnavh.bsky.social · 25/12/2025If X, Y, Z are iid high-dim Gaussian N(0, I), what's the angle between X-Y and Z-Y? A. Concentrates at 90 deg B. Concentrates, NOT at 90° C. Doesn't concentrate anywhere. My (and most people's) instincts got this wrong! vaishnavh.github.io/blog/high-di...vaishnavh.github.ioAngles between high-dimensional vectors - Vaishnavh NagarajanSwitch off your brain and answer this:Given three points $\mathbf{X}, \mathbf{Y}, \mathbf{Z}$ sampled from a high-dim... 160
Reposted by Vaishnavh NagarajanCMU Computer Science Department @csdatcmu.bsky.social · 17/07/2025Congratulations to CSD faculty Aditi Raghunathan and her research collaborators on receiving an ICML Outstanding Paper award for Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction (icml.cc/virtual/2025...). Paper: arxiv.org/abs/2504.15266icml.ccICML 2025 AwardsICML 2025 061
Reposted by Vaishnavh NagarajanEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 04/07/2025Reading the dedications of a PhD thesis is often a cure for a bad day. There’s so much affection in them 4431
Reposted by Vaishnavh NagarajanAhmad Beirami @abeirami.bsky.social · 02/07/2025As NeurIPS review deadline is around the corner, please remember that you cannot use any non-local LLM like chatgpt/gemini for understanding the paper and drafting/revising your review as that breaks the confidentiality agreement. NeurIPS 2025 Official LLM Policy: neurips.cc/Conferences/...neurips.ccLLM Policy 061
Reposted by Vaishnavh NagarajanGiovanni Toffetti @gtof.eurosky.social · 22/06/2025I really enjoyed "When We Cease to Understand the World", although it's more fiction than history of science 021
Reposted by Vaishnavh NagarajanJavier Burroni @jburroni.bsky.social · 23/06/2025“Science in history” by Bernal is my first recommendation. The work of Ian Hacking is a good recommendation for Probability 011
Reposted by Vaishnavh NagarajanAlexandra Proca @aproca.bsky.social · 20/06/2025How do task dynamics impact learning in networks with internal dynamics? Excited to share our ICML Oral paper on learning dynamics in linear RNNs! with @clementinedomine.bsky.social @mpshanahan.bsky.social and Pedro Mediano openreview.net/forum?id=KGO...openreview.netLearning dynamics in linear recurrent neural networksRecurrent neural networks (RNNs) are powerful models used widely in both machine learning and neuroscience to learn tasks with temporal dependencies and to model neural dynamics. However, despite... 13411
Reposted by Vaishnavh NagarajanEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 22/06/2025When we are doing science, we are unknowingly executing our mythology, taken from movies and friends and textbooks, of what science is. History of science helps us ground that myth in reality 3536
Vaishnavh Nagarajan @vaishnavh.bsky.social · 13/06/2025I finally wrote a full-fledged blog about this: reading the history of science is an **amazing** yet under-recognized way to develop (emotional) maturity as a researcher. If you have thoughts/recommendations, please share! vaishnavh.github.io/2025/04/29/h...vaishnavh.github.ioWhy PhD students should read the history of science - Vaishnavh NagarajanHere’s a secret that I accidentally discovered during my PhD: consuming history-of-science content is an efficient wa... 3357
Reposted by Vaishnavh NagarajanTiago Pimentel @tpimentel.bsky.social · 04/06/2025A string may get 17 times less probability if tokenised as two symbols (e.g., ⟨he, llo⟩) than as one (e.g., ⟨hello⟩)—by an LM trained from scratch in each situation! Our new ACL paper proposes an observational method to estimate this causal effect! Longer thread soon! 1518
Reposted by Vaishnavh NagarajanAndrew Saxe @saxelab.bsky.social · 04/06/2025How does in-context learning emerge in attention models during gradient descent training? Sharing our new Spotlight paper @icmlconf.bsky.social: Training Dynamics of In-Context Learning in Linear Attention arxiv.org/abs/2501.16265 Led by Yedi Zhang with @aaditya6284.bsky.social and Peter Latham 15318
Reposted by Vaishnavh NagarajanEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 02/06/2025This paper is quite nice. It mixes some useful toy models of creativity with insights about how to induce more creativity in LLMs that are better than greedy sampling 1164
Vaishnavh Nagarajan @vaishnavh.bsky.social · 02/06/2025📢 New #paper on creativity & multi-token prediction! We design minimal open-ended tasks to argue: → LLMs are limited in creativity as they learn to predict the next token → creativity can be improved via multi-token learning & injecting noise ("seed-conditioning" 🌱) 1/ #MLSky #AI #arxiv 🧵👇🏽 1272
Reposted by Vaishnavh NagarajanNathan Lambert @natolambert.bsky.social · 27/05/2025This isn't fake news. One of the craziest AI research papers I've been on in a while. Weird ablations on RLVR shows that the Qwen 2.5 models can learn with literally random rewards, likely due to some funkiness in mid-training and the GRPO setup. 4476
Vaishnavh Nagarajan @vaishnavh.bsky.social · 27/05/2025can someone reconcile these two contradictory findings?! two papers find entropy *minimization*/confidence maximization helps performance, and the RL-on-one-sample finds entropy maximization/increasing exploration alone helps performance?! 572
Reposted by Vaishnavh NagarajanCore Francisco Parkg @corefpark.bsky.social · 21/05/2025🚨 New Paper! A lot happens in the world every day—how can we update LLMs with belief-changing news? We introduce a new dataset "New News" and systematically study knowledge integration via System-2 Fine-Tuning (Sys2-FT). 1/n 161
Vaishnavh Nagarajan @vaishnavh.bsky.social · 21/05/2025The contextual shadowing effect in this paper is interesting & reminds me of the infamous "cow-in-a-beach" spurious correlation examples. The model needs to learn a core feature (here. the store-in-memory feature) but it instead relies on a simpler "spurious" feature (the in-context feature). 111
Vaishnavh Nagarajan @vaishnavh.bsky.social · 18/05/2025My #neurips bidding pool is REALLY bad and I see that I'm not alone. Was the pool restricted to 100 papers to break collusion? if so, seems like everyone else is going to pay for this 240
Reposted by Vaishnavh NagarajanAhmad Beirami @abeirami.bsky.social · 23/04/2025Excited that our paper "safety alignment should be made more than just a few tokens deep" was recognized as an #ICLR2025 Outstanding Paper! We identified a common root cause to many safety vulnerabilities and pointed out some paths forward to address it! 2323
Reposted by Vaishnavh NagarajanNikhil Garg @nkgarg.bsky.social · 07/05/2025At the moment, we just show you posts by people you follow. Following someone will make their paper posts appear! In the future, we'll expand this (slightly) to show you other paper posts that may be of interest 241
Reposted by Vaishnavh NagarajanWilliam B. Fuckley @opinionhaver.bsky.social · 05/05/2025Yeah this is actual unironic advice for any 1st years: you seriously need to go get beers or otherwise hangout socially with more senior grad students (no one else will really know) and hopefully get ensconced enough that someone will tell you “hey just fyi Dr. GoodPubs is low-key a total sociopath” 633333
Reposted by Vaishnavh NagarajanEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 06/05/2025@vaishnavh.bsky.social and crew do it again: - a benchmark for open-ended creativity - a demonstration of challenges of next-token prediction - a technique to improve transformer randomness through inputs not sampling arxiv.org/abs/2504.15266arxiv.orgRoll the dice & look before you leap: Going beyond the creative limits of next-token predictionWe design a suite of minimal algorithmic tasks that are a loose abstraction of open-ended real-world tasks. This allows us to cleanly and controllably quantify the creative limits of the present-day l... 1376
Vaishnavh Nagarajan @vaishnavh.bsky.social · 06/05/2025@icmlconf.bsky.social (assuming this is the legit ICML account), the Canadian visa application requires providing the inviting contact to have a Canadian address. But the one in the visa letter is a US address. Could you look into this? 000
Reposted by Vaishnavh NagarajanAhmad Beirami @abeirami.bsky.social · 03/05/2025If you are at #AISTATS2025 and are interested in concept erasure, talk to @somnathbrc.bsky.social at Poster Session 1 on Saturday May 3. 0164
Reposted by Vaishnavh NagarajanNeurIPS Conference @neuripsconf.bsky.social · 02/05/2025Responsible reviewing initiatives for NeurIPS 2025 - read more about changes to reviewing that that will safeguard reviewing quality and timeline in our blog post below: blog.neurips.cc/2025/05/02/r...blog.neurips.ccResponsible Reviewing Initiative for NeurIPS 2025 – NeurIPS BlogCommunications Chairs 2025 2021 Conference 1255
Reposted by Vaishnavh NagarajanDimitris Papailiopoulos @dimitrisp.bsky.social · 02/02/2025Self-improving Transformers can overcome easy-to-hard and length generalization challenges. Paper on arxiv coming on Monday. Link to a talk I gave on this below 👇 Super excited about this work! Talk : youtube.com/watch?v=szhE... slides: tinyurl.com/SelfImprovem... 0151
Reposted by Vaishnavh NagarajanGautam Kamath @gautamkamath.com · 01/05/2025I wrote a post on how to connect with people (i.e., make friends) at CS conferences. These events can be intimidating so here's some suggestions on how to navigate them I'm late for #ICLR2025 #NAACL2025, but in time for #AISTATS2025 #ICML2025! 1/3 kamathematics.wordpress.com/2025/05/01/t...kamathematics.wordpress.comTips on How to Connect at Academic ConferencesI was a kinda awkward teenager. If you are a CS researcher reading this post, then chances are, you were too. How to navigate social situations and make friends is not always intuitive, and has to … 36618
Reposted by Vaishnavh NagarajanDanica Sutherland @djsutherland.ml · 23/04/2025Thrilled to announce that Joshua’s paper won one of three Outstanding Paper awards at ICLR. Come to the poster on Friday afternoon (#376 in Hall 3) or the talk on Saturday (4:30 in Hall 1), and while you’re at it snag him for a postdoc! 1366
Reposted by Vaishnavh NagarajanEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 14/04/2025Been waiting for a work making this clear for a while: zero-shot, LLMs do not hedge and do not ask questions to solve under-specified problems arxiv.org/abs/2503.22674arxiv.orgQuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?Recently, a large amount of work has focused on improving large language models' (LLMs') performance on reasoning benchmarks such as math and logic. However, past work has largely assumed that tasks a... 38122
Vaishnavh Nagarajan @vaishnavh.bsky.social · 03/04/2025What tools do people use to manage AC/meta-reviewing responsibilities (like consolidating key points, flagging which papers/reviews need attention, etc., etc.,)? 000
Reposted by Vaishnavh NagarajanDimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025o3 can't multiply beyond a few digits... But I think multiplication, addition, maze solving and easy-to-hard generalization is actually solvable on standard transformers... with recursive self-improvement Below is the acc of a tiny model teaching itself how to add and multiply 2214
Reposted by Vaishnavh NagarajanJacob Springer @jacobspringer.bsky.social · 26/03/2025Training with more data = better LLMs, right? 🚨 False! Scaling language models by adding more pre-training data can decrease your performance after post-training! Introducing "catastrophic overtraining." 🥁🧵👇 arxiv.org/abs/2503.19206 1/10 13314