Sign in

Ekdeep Singh @ ICML

@ekdeepl.bsky.social
280 followers 380 following 48 posts

Postdoc at CBS, Harvard University (New around here)

PostsRepliesMedia
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 06/08/2025
Tubingen just got ultra-exciting :D
030
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/07/2025
Submit your latest and greatest papers to the hottest workshop on the block---on cognitive interpretability! 🔥
081
Reposted by Ekdeep Singh @ ICML
Jennifer Hu @jennhu.bsky.social · 16/07/2025
Excited to announce the first workshop on CogInterp: Interpreting Cognition in Deep Learning Models @ NeurIPS 2025! 📣 How can we interpret the algorithms and representations underlying complex behavior in deep learning models? 🌐 coginterp.github.io/neurips2025/ 1/4
coginterp.github.io
Home
First Workshop on Interpreting Cognition in Deep Learning Models (NeurIPS 2025)
15819
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 12/07/2025
I'll be at ICML beginning this Monday---hit me up if you'd like to chat!
030
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 08/07/2025
Our recent paper may be relevant (arxiv.org/abs/2506.17859)! We take a rational analysis lens to argue that beyond simplicity bias, we must model how well a hypothesis explains the data to yield a *predictive* account of behavior in neural nets! This helps explain learning of more complex functions.
arxiv.org
In-Context Learning Strategies Emerge Rationally
Recent work analyzing in-context learning (ICL) has identified a broad set of strategies that describe model behavior in different experimental conditions. We aim to unify these findings by asking why...
030
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 28/06/2025
I am definitely not biased :)
010
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 28/06/2025
Check out one of the most exciting papers of the year! :D
120
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 29/04/2025
I'll be attending NAACL at New Mexico beginning today---hit me up if you'd like to chat!
000
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 07/03/2025
Check out our new work on the duality between SAEs and how concepts are organized in model representations!
020
Reposted by Ekdeep Singh @ ICML
Sumedh Hindupur @sumedh-hindupur.bsky.social · 07/03/2025
New preprint alert! Do Sparse Autoencoders (SAEs) reveal all concepts a model relies on? Or do they impose hidden biases that shape what we can even detect? We uncover a fundamental duality between SAE architectures and concepts they can recover. Link: arxiv.org/abs/2503.01822
1142
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 17/02/2025
Oh god, I had no clue this happened. :/ Does bsky not do GIFs?
100
Reposted by Ekdeep Singh @ ICML
Andrew Lampinen @lampinen.bsky.social · 16/02/2025
Very nice paper; quite aligned with the ideas in our recent perspective on the broader spectrum of ICL. In large models, there's probably a complicated, context dependent mixture of strategies that get learned, not a single ability.
1172
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
Paper co-led with @corefpark, and in collaboration with the ever-awesome @PresItamar and @Hidenori8Tanaka! :) arXiv link: arxiv.org/abs/2412.01003
arxiv.org
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
In-Context Learning (ICL) has significantly expanded the general-purpose nature of large language models, allowing them to adapt to novel tasks using merely the inputted context. This has motivated a ...
030
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
Dynamics of attention maps is particularly striking in this competition: e.g., with high diversity, we see a bigram counter forming, but memorization eventually occurs and the pattern becomes uniform! This means models can remove learned components if they are not useful anymore!
130
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
Beyond corroborating our phase diagram, LIA confirms a persistent competition underlies ICL: once the induction head forms, with enough diversity bigram-based inference takes over; under low diversity, memorization occurs faster, yielding a bigram-based retrieval solution!
110
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
Given optimization and data diversity are independent axes, we then ask if these forces race against each other to yield our observed algorithmic phases. We propose a tool called LIA (linear interpolation of algorithms) for this analysis.
120
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
The tests check out! We see before/after a critical # of train steps are met (where induction head emerges), the model relies on unigram/bigram stats. With few chains (less diversity), there is retrieval behavior: we can literally reconstruct transition matrices from MLP neurons!
110
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
To test the above claim, we compute the effect of shuffling a sequence on next-token probs: this breaks bigram stats, but preserves unigrams. We check how “retrieval-like” or memorization-based model behavior is by comparing predicted transitions’ KL to a random set of chains.
120
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
We claim four algorithms explain the model’s behavior in diff. train/test settings. These algos compute uni-/bi- gram frequency statistics of an input to either *retrieve* a memorized chain or to in-context *infer* the chain used to define the input: latter performs better OOD!
120
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
We analyze models trained on a fairly simple task: learning to simulate a *finite mixture* of Markov chains. The sequence modeling nature of this task makes it a better abstraction for studying ICL abilities in LMs, (compared to abstractions of few-shot learning like linear reg.)
130
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
New paper–accepted as *spotlight* at #ICLR2025! 🧵👇 We show a competition dynamic between several algorithms splits a toy model’s ICL abilities into four broad phases of train/test settings! This means ICL is akin to a mixture of different algorithms, not a monolithic ability.
2325
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 12/02/2025
bsky.app/profile/ajyl...
110
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 12/02/2025
Some threads about recent works ;-) bsky.app/profile/ekde...
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 12/02/2025
Our group, funded by NTT Research, Inc., uniquely bridges industry and academia. We integrate approaches from physics, neuroscience, and psychology while grounding our work in empirical AI research. Apply from "Physics of AI Group Research Intern": careers.ntt-research.com
careers.ntt-research.com
CareerPortal
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 12/02/2025
Check out our recent work, including 5 ICLR 2025 papers all (co)led by amazing past and current interns: sites.google.com/view/htanaka...
sites.google.com
Hidenori Tanaka
Hidenori Tanaka Group Leader, Science of Intelligence for Alignment CBS-NTT Program in Physics of Intelligence, Harvard University Google Scholar email: hidenori_tanaka [at] fas.harvard.edu, twitter ...
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 12/02/2025
There's never been a more exciting time to explore the science of intelligence! 🧠 What can ideas and approaches from science tell us about how AI works? What might superhuman AI reveal about human cognition? Join us for an internship at Harvard to explore together! 1/
121
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 30/01/2025
Are those... bats?
110
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 23/01/2025
Now accepted at NAACL! This would be my first time presenting at an ACL conference---I've got almost first-year grad school level of excitement! :P
040
Reposted by Ekdeep Singh @ ICML
Core Francisco Parkg @corefpark.bsky.social · 05/01/2025
New paper! “In-Context Learning of Representations” What happens to an LLM’s internal representations in the large context limit? We find that LLMs form “in-context representations” to match the structure of the task given in context!
172
Reposted by Ekdeep Singh @ ICML
Andrew Lee @ajyl.bsky.social · 05/01/2025
New paper <3 Interested in inference-time scaling? In-context Learning? Mech Interp? LMs can solve novel in-context tasks, with sufficient examples (longer contexts). Why? Bc they dynamically form *in-context representations*! 1/N
25316
Reposted by Ekdeep Singh @ ICML
Naomi Saphra @nsaphra.bsky.social · 20/12/2024
Transformer LMs get pretty far by acting like ngram models, so why do they learn syntax? A new paper by sunnytqin.bsky.social, me, and @dmelis.bsky.social illuminates grammar learning in a whirlwind tour of generalization, grokking, training dynamics, memorization, and random variation. #mlsky #nlp
arxiv.org
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
Language models (LMs), like other neural networks, often favor shortcut heuristics based on surface-level patterns. Although LMs behave like n-gram models early in training, they must eventually learn...
514230
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
Overall, there's a long way to go, but I'm hopeful for the next year of SAE research and am looking forward to contributing to it! Paper link again: arxiv.org/abs/2410.11767 And note again that paper co-lead Abhinav Menon is awesome and looking for PhD positions! Recruit him!
000
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
2. We should build on ideas from disentanglement research to avoid reinventing the wheel! For ex., the community came up with really ingenious approaches to exploit "weak supervision" (kinda like self-supervision) to enable disentanglement! See paper linked in pic and our paper for a jab at this!
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
1. Can we show useful non-linear features exist? If they do, we can be certain our current SAE paradigm needs change! I don't think we as a community have tried sufficiently hard to find such features, and for that matter we have no clarity on what features end up linear!
110
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
So where do we go from here? Personally, I think SAEs are a really promising direction, but our current paradigm is bound for failure if model features turn out to be non-linear (which may not be the case!). To this end, two directions of future work seem promising to me.👇
110
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
Result 3: None of the latents turn out to be causally relevant to the model's computation! E.g., intervening on a specific part-of-speech feature induces no effects of distribution of other parts in model's generations! Again, we expect this based on results from disentanglement!
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
Result 2: The results above are however quite brittle. Small architectural changes yield huge deviations in fraction of variance explained of the original representations! This inability / limitation of SAEs is exactly what the disentanglement community's results suggested!
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
Result 1: We can find highly interpretable features in SAEs! We follow standard protocols to arrive at these results, yielding, e.g., features for parts-of-speech, counters, and depth trackers in different languages!
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
Taking inspiration from the disentanglement literature, we first define toy settings where we know what ground-truth latents to expect: specifically, we use formal grammars like Dyck-2, Expr, and English. We pretrain models on these grammars and then train SAEs on their features!
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
Do these results extend to SAEs? Are SAE features causally relevant to the model's computation? Is there sufficient consistency across training runs? We investigate these questions in our work!
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
Interestingly, the ICML best paper winner in 2019 established an impossibility theorem for disentanglement: if the DGP is nonlinear, there are infinite solutions to the problem. Hence, identified latents will depend on what AE pipeline is used, and may not be causally relevant!
210
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
In Disentanglement, one trains Autoencoders (possibly with some regularization like sparsity) to identify latents underlying the data-generating process (DGP); with SAEs, one is essentially trying to solve the same problem, but the data source is now model representations!
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
Paper co-led with Abhinav Menon (bsky.app/profile/bani...), who is awesome and is looking for PhD positions---recruit him! arXiv link: arxiv.org/abs/2410.11767
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 18/12/2024
Paper alert––*Awarded best paper* at NeurIPS workshop on Foundation Model Interventions! 🧵👇 We analyze the (in)abilities of SAEs by relating them to the field of disentangled rep. learning, where limitations of AE based interpretability protocols have been well established!🤯
2369
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 10/11/2024
Also highlighting our previous paper in this concept learning series that this work builds on top of, and a theoretical follow up with @YongyiYang7 & @weihu_ where we find we can prove the bulk of empirical results shown above! Previous: arxiv.org/abs/2310.09336 Follow-up: arxiv.org/abs/2410.08309
000
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 10/11/2024
Finally, shoutout to my amazing collaborators @corefpark, @MayaOkawa, @a_jy_l, and @Hidenori8Tanaka! This was a really fun collaboration where everybody played a very different, extremely synergistic role. :)
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 10/11/2024
The paper contains several more results! E.g., theoretical models that capture the learning dynamics, impact of underspecification (i.e., correlated concepts), and further musings on concept signal. Check out the link below! Link: arxiv.org/abs/2406.19370
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 10/11/2024
To further corroborate this claim, we (i) train models on CelebA & (ii) experiment with pretrained StableDiffusion models. We again find latent interventions can elicit arbitrary compositions that are not present in the training data (or are unlikely to be present for SD models)!
100
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 10/11/2024
Despite being a toy setting, we hypothesize this result will likely hold at scale with modern generative models: there exist latent capabilities that one, via mere input prompting, is unable to elicit! This can be concerning, for such capabilities may slip through during evals!
110
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 10/11/2024
We hypothesize the sudden turns mark *disentanglement*: the model can arbitrarily compose all concepts after this turn. But learning dynamics shows otherwise–what’s going on?! Turns out capabilities are *latent* at this point, but can be elicited via mere linear interventions!
100