Sign in

Eric Todd

@ericwtodd.bsky.social
507 followers 184 following 16 posts

CS PhD Student, Northeastern University - Machine Learning, Interpretability ericwtodd.github.io

PostsRepliesMedia
Eric Todd @ericwtodd.bsky.social · 22/01/2026
Another strategy infers meaning using sets. We have seen models keep track of "positive" and "negative" sets that let it narrow its understanding of a symbol using Sudoku-style cancellation. Red bars (a) show the positive set and blue boxes (b) show the negative.
160
Eric Todd @ericwtodd.bsky.social · 22/01/2026
What in-context mechanisms do we find, other than copying? The first one is the "identity rule". Here, the answer is the same as the question after eliminating a recognized "identity" from the question, like "ab=a". @taylorwwebb.bsky.social has seen this in LLMs too! bsky.app/profile/tay...
170
Eric Todd @ericwtodd.bsky.social · 22/01/2026
Our work maps out several context-based algorithms (copy, identity, commutativity, cancellation, & associativity). We use targeted data distributions to measure and dissect each strategy. These five strategies explain almost all of our model's in-context performance!
180
Eric Todd @ericwtodd.bsky.social · 22/01/2026
If you pick a random puzzle (try one here: algebra.baulab.info), you'll see there's often more than one way to understand context. @nelhage.bsky.social & @neelnanda.bsky.social found LLMs infer meaning by induction-style copying, and that happens here too. But there are many other strategies.
160
Eric Todd @ericwtodd.bsky.social · 22/01/2026
Can you solve this algebra puzzle? 🧩 cb=c, ac=b, ab=? A small transformer can learn to solve problems like this! And since the letters don't have inherent meaning, this lets us study how context alone imparts meaning. Here's what we found:🧵⬇️
24811