Sign in

Isabelle Lee

@wordscompute.bsky.social
2.4K followers 534 following 56 posts

ml/nlp phding @ usc, currently visiting harvard; training & interpretability & reasoning iglee.me

PostsRepliesMedia
Reposted by Isabelle Lee
Naomi Saphra @nsaphra.bsky.social · 15/06/2026
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social
313734
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
really excited to head home for icml:) and attending the co-located FAR.ai alignment workshop (for the first time)! would love to meet others interested in training & interpretability
far.ai
FAR.AI: Frontier Alignment Research
FAR.AI is an AI safety research non-profit facilitating technical breakthroughs and fostering global collaboration.
030
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
also, blog: iglee.me/papers/inte... 7/6
iglee.me
Rigorous Interpretation Is a Form of Evaluation
If held to scientific standards of falsifiability, reproducibility, and predictability, interpretation can become evaluation.
010
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
work w/ Emmy Liu, Cathy Jiao @brihi.bsky.social, Dani Yogatama, Fazl Barez, @saxon.me since i'm headed home for icml, presented by amazing @brihi.bsky.social! this was my first time writing a position paper, which turned into a grant, which i'm turning into multiple projects 🙂 stay tuned 6/6
110
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
3. Predicting failures. A distinction: scientific prediction (not the ML kind) is how scientists validate our understanding. A hypothesis proves its strength w/ predictive power. Used as eval, interp can predict failures from internals. Meaning, we generate eval from interp. 5/6
100
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
2. Reproducibility detects faulty mechanisms. If we were to actually act on it, we want our claim to identify mechanisms robustly against variations wrt input, method, etc. Our claim needs to be reproducible under specified conditions. 4/6
100
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
1. Falsifiability enables debugging. First argued by Leavitt & Morcos, it has to produce hypotheses that can be proven wrong. And if it can, we can then act on it, to trace back to the source of error and attempt a fix. 3/6
100
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
What does it mean for interp to meet scientific standards? We argue that it has to meet 3 criteria: falsifiability, reproducibility and predictability. 2/6
100
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
Benchmarks can be superficial, but model explanations and evaluations are fundamentally intertwined. What if we used interpretability as principled, scientific evaluation? If it met scientific standards? arxiv.org/abs/2605.05508 coming to EvalEval at ACL as oral 🧵 1/6
The text lists authors associated with the topic of rigorous interpretation as evaluation, affiliated with various universities.
1141
Reposted by Isabelle Lee
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Our ICML 2025 workshop on Actionable Interpretability drew massive interest. But the same questions kept coming up: What does "actionable" mean? Is it achievable? How? We're ready to answer. 🧵
12410
Reposted by Isabelle Lee
Sarah Liaw @sarahliaw.bsky.social · 12/02/2026
Excited to share our new dataset, FOL-Traces! We introduce a large-scale dataset of programmatically verified FOL reasoning traces for studying structured logical inference + process fidelity. Happy to hear thoughts from others working on reasoning in LLMs! Check it out here 👇
041
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
paper: arxiv.org/abs/2505.14932 dataset: huggingface.co/datasets/fo... work w/ @sarahliaw.bsky.social and Dani Yogatama If you want to chat about interpretability & training dynamics & reasoning and munch on mezzes, come hang out with me in Rabat 🇲🇦🙃 9/9
huggingface.co
fol-traces/fol-traces · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
000
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
I wanted to study reasoning acquisition in training by complexity + process fidelity but wasn't able to find a dataset. So we built one that's rigorously annotated and large enough to train a small LM. Now I’m excited about what we can do with it 8/9
100
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
a harder task- last step prediction: ¬(¬Sunny(x) ∧ Breezy(x)) ↔ [MASK] or last two step prediction. Most LLMs only achieve <50% accuracy on both tasks. (n.b. since FOL is verifiable, we define correct as any generation that's equivalent to expression.) 7/9
Bar graph displaying the accuracy percentages of various models on two-step prediction tasks, with distinct colors for each step.Bar chart displaying accuracy percentages for various models across three complexity thresholds: 10-19, 20-29, and 30+.
110
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
e.g. masked prediction. we mask an operator randomly and have LLMs guess: ¬(¬Sunny(x) ∧ Breezy(x)) ↔ (Sunny(x) [MASK] Breezy(x)). LLMs are correct ~45.7% on average: 6/9
Table displaying model accuracies for components, operators, and predicates prediction tasks, with various metrics for performance evaluation.
100
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
...resulting in a bunch of reasoning traces that are verifiably correct with measurable programmatic complexity. And we find that they're very hard for LLMs! Let's consider an example w/ de Morgan's law: ¬(¬Sunny(x) ∧ Breezy(x)) ↔ (Sunny(x) ∨ Breezy(x)) 5/9
100
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
So how do we strike a balance? We propose using First-Order Logic (FOL) as a middle ground. We 1. programmatically, randomly generate a bunch of FOL expressions 2. progressively simplify them, verifying their equivalence 3. chain them together 4. NL instantiate them w/ LLMs 4/9
Flowchart illustrating the process of using First-Order Logic, integrating human rules, symbolic generation, and LLM instantiation for reasoning examples.
100
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
We mostly interface with LLMs with words but evaluating NL reasoning is messy. On the other hand, something like math reasoning gives us concrete, objectively correct answers. But it’s narrow/doesn’t look like NL. 3/9
100
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
There are many evals and benchmarks in this field, but natural language (NL) reasoning is tricky--meaning depends on context (commonsense), shared assumptions (pragmatics), and what’s unsaid (abduction). Pattern shortcuts/heuristics ≠ logical inference. 2/9
110
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
New dataset 🗂️ coming to #eacl What is (correct) reasoning in LLMs? How do you rigorously define/measure process fidelity? How might we study its acquisition in large scale training? We made a gigantic, verifiably correct reasoning traces of first order logic expressions! 1/9
Title highlights "FOL-Traces," a dataset for evaluating logical reasoning in language models, emphasizing rigorous testing and performance metrics.
140
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
one of my new years "considerations" is to be less silent #onhere. so i guess i'll be #here and maybe also #there til february 15th
gemini summarized my google search when i was tryna look for an anti-new years resolution blog post. it says in highlight: "Approximately 80% to 88% of New Year's resolutions fail by mid-February, ..."
000
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
If you're interested in interpretability driven evaluations, I'd love to hear from you! And stay tuned for more work from us :)
030
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
Really excited to receive Coefficient Giving's Technical AI Safety Research Grant via Berkeley Existential Risk Initiative w/ @nsaphra.bsky.social! We aim to predict potential AI model failures before impact--before deployment, using interpretability.
161
Reposted by Isabelle Lee
Fazl Barez @fbarez.bsky.social · 01/07/2025
Excited to share our paper: "Chain-of-Thought Is Not Explainability"! We unpack a critical misconception in AI: models explaining their steps (CoT) aren't necessarily revealing their true reasoning. Spoiler: the transparency can be an illusion. (1/9) 🧵
28531
Isabelle Lee @wordscompute.bsky.social · 10/05/2025
weve reached that point in this submission cycle, no amount of coffee will do 😞🙂‍↔️😞
010
Isabelle Lee @wordscompute.bsky.social · 29/03/2025
INCOMING
020
Isabelle Lee @wordscompute.bsky.social · 29/03/2025
titled: peer review
a leaf falls on moo deng the pygmy hippo , blocking her visionmoo deng is upset presumably because she can’t see!
071
Reposted by Isabelle Lee
Naomi Saphra @nsaphra.bsky.social · 27/03/2025
Life update: I'm starting as faculty at Boston University @bucds.bsky.social in 2026! BU has SCHEMES for LM interpretability & analysis, I couldn't be more pumped to join a burgeoning supergroup w/ @najoung.bsky.social @amuuueller.bsky.social. Looking for my first students, so apply and reach out!
CDS building which looks like a jenga tower
3524213
Isabelle Lee @wordscompute.bsky.social · 15/03/2025
or if you're awesome and happen to be in sf, also message me
010
Isabelle Lee @wordscompute.bsky.social · 15/03/2025
pls message me if you wanna meet up for coffee and chat about ai/physics/llms/interpretability
100
Isabelle Lee @wordscompute.bsky.social · 15/03/2025
really excited to be headed to OFC in SF! so excited to revisit optical physics 😀
210
Reposted by Isabelle Lee
aaditya6284.bsky.social @aaditya6284.bsky.social · 11/03/2025
Transformers employ different strategies through training to minimize loss, but how do these tradeoff and why? Excited to share our newest work, where we show remarkably rich competitive and cooperative interactions (termed "coopetition") as a transformer learns. Read on 🔎⏬
184
Isabelle Lee @wordscompute.bsky.social · 05/03/2025
i use the same template and need help getting a butterfly button help
000
Reposted by Isabelle Lee
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
New paper–accepted as *spotlight* at #ICLR2025! 🧵👇 We show a competition dynamic between several algorithms splits a toy model’s ICL abilities into four broad phases of train/test settings! This means ICL is akin to a mixture of different algorithms, not a monolithic ability.
2325
Reposted by Isabelle Lee
Quanta Magazine @quantamagazine.org · 15/02/2025
Starlings move in undulating curtains across the sky. Forests of bamboo blossom at once. But some individuals don’t participate in these mystifying synchronized behaviors — and scientists are learning that they may be as important as those that do.
buff.ly
Out-of-Sync ‘Loners’ May Secretly Protect Orderly Swarms
Studies of collective behavior usually focus on how crowds of organisms coordinate their actions. But what if the individuals that don’t participate have just as much to tell us?
23310
Reposted by Isabelle Lee
Margaret Mitchell @mmitchell.bsky.social · 06/02/2025
New piece out! We explain why Fully Autonomous Agents Should Not be Developed, breaking “AI Agent” down into its components & examining through ethical values. With @evijit.io, @giadapistilli.com and @sashamtl.bsky.social huggingface.co/papers/2502....
huggingface.co
Paper page - Fully Autonomous AI Agents Should Not be Developed
Join the discussion on this paper page
414148
Reposted by Isabelle Lee
Quanta Magazine @quantamagazine.org · 05/02/2025
Brian Hie harnessed the powerful parallels between DNA and human language to create an AI tool that interprets genomes. Read his conversation with Ingrid Wickelgren: www.quantamagazine.org/the-poetry-f...
quantamagazine.org
The Poetry Fan Who Taught an LLM to Read and Write DNA | Quanta Magazine
By treating DNA as a language, Brian Hie’s “ChatGPT for genomes” could pick up patterns that humans can’t see, accelerating biological design.
14014
Reposted by Isabelle Lee
Valérie Castin @vcastin.bsky.social · 31/01/2025
How do tokens evolve as they are processed by a deep Transformer? With José A. Carrillo, @gabrielpeyre.bsky.social and @pierreablin.bsky.social, we tackle this in our new preprint: A Unified Perspective on the Dynamics of Deep Transformers arxiv.org/abs/2501.18322 ML and PDE lovers, check it out!
29616
Isabelle Lee @wordscompute.bsky.social · 26/01/2025
it’s finally raining in la:)
050
Isabelle Lee @wordscompute.bsky.social · 09/01/2025
001
Isabelle Lee @wordscompute.bsky.social · 09/01/2025
part of me wants to quip, is this why i quit smoking, but i think im actually getting a lil scared. hope we get thru the next few days okay cause feels like theres very little we can do here rn
110
Isabelle Lee @wordscompute.bsky.social · 09/01/2025
when i lived in seattle, fires were a summer expectation at a distance. here, it feels very different, to see it actually closing in on us
100
Isabelle Lee @wordscompute.bsky.social · 09/01/2025
i go on a really long walk almost every day, and at a high point in silverlake, i saw fire from all sides. and it's harder to breathe. and everything is orange.
160
Reposted by Isabelle Lee
Andrew Lee @ajyl.bsky.social · 05/01/2025
New paper <3 Interested in inference-time scaling? In-context Learning? Mech Interp? LMs can solve novel in-context tasks, with sufficient examples (longer contexts). Why? Bc they dynamically form *in-context representations*! 1/N
25316
Reposted by Isabelle Lee
Los Angeles Times @latimes.com · 09/01/2025
Hollywood High School will serve as an evacuation site for the Sunset fire in Hollywood, KTLA reported. The school is at 1521 Highland Ave. www.latimes.com/california/s...
latimes.com
Sunset fire in Hollywood Hills: Evacuations, shelter
An evacuation zone was established between the 101 Freeway and Laurel Canyon and between Mulholland Drive and Hollywood Boulevard.
231131455
Reposted by Isabelle Lee
Bálint Máté @balintmate.bsky.social · 17/12/2024
hello bluesky! we have a new preprint on solvation free energies: tl;dr: We define an interpolating density by its sampling process, and learn the corresponding equilibrium potential with score matching. arxiv.org/abs/2410.15815 with @francois.fleuret.org and @tbereau.bsky.social (1/n)
13310
Reposted by Isabelle Lee
Kate @katef.bsky.social · 15/12/2024
look at our sheep
3336
Reposted by Isabelle Lee
Arnaud Doucet @arnauddoucet.bsky.social · 15/12/2024
The slides of my NeurIPS lecture "From Diffusion Models to Schrödinger Bridges - Generative Modeling meets Optimal Transport" can be found here drive.google.com/file/d/1eLa3...
drive.google.com
BreimanLectureNeurIPS2024_Doucet.pdf
932868
Reposted by Isabelle Lee
Jennifer Hu @jennhu.bsky.social · 11/12/2024
Slides from the tutorial are now posted here! neurips.cc/media/neurip...
neurips.cc
0177
Reposted by Isabelle Lee
Sakana AI @sakanaai.bsky.social · 10/12/2024
An Evolved Universal Transformer Memory sakana.ai/namm/ Introducing Neural Attention Memory Models (NAMM), a new kind of neural memory system for Transformers that not only boost their performance and efficiency but are also transferable to other foundation models without any additional training!
14115