Sign in

Isabelle Lee

@wordscompute.bsky.social
2.4K followers 534 following 56 posts

ml/nlp phding @ usc, currently visiting harvard; training & interpretability & reasoning iglee.me

PostsRepliesMedia
Reposted by Isabelle Lee
Naomi Saphra @nsaphra.bsky.social · 15/06/2026
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social
313734
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
really excited to head home for icml:) and attending the co-located FAR.ai alignment workshop (for the first time)! would love to meet others interested in training & interpretability
far.ai
FAR.AI: Frontier Alignment Research
FAR.AI is an AI safety research non-profit facilitating technical breakthroughs and fostering global collaboration.
030
Isabelle Lee @wordscompute.bsky.social · 14/06/2026
Benchmarks can be superficial, but model explanations and evaluations are fundamentally intertwined. What if we used interpretability as principled, scientific evaluation? If it met scientific standards? arxiv.org/abs/2605.05508 coming to EvalEval at ACL as oral 🧵 1/6
The text lists authors associated with the topic of rigorous interpretation as evaluation, affiliated with various universities.
1141
Reposted by Isabelle Lee
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Our ICML 2025 workshop on Actionable Interpretability drew massive interest. But the same questions kept coming up: What does "actionable" mean? Is it achievable? How? We're ready to answer. 🧵
12410
Reposted by Isabelle Lee
Sarah Liaw @sarahliaw.bsky.social · 12/02/2026
Excited to share our new dataset, FOL-Traces! We introduce a large-scale dataset of programmatically verified FOL reasoning traces for studying structured logical inference + process fidelity. Happy to hear thoughts from others working on reasoning in LLMs! Check it out here 👇
041
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
New dataset 🗂️ coming to #eacl What is (correct) reasoning in LLMs? How do you rigorously define/measure process fidelity? How might we study its acquisition in large scale training? We made a gigantic, verifiably correct reasoning traces of first order logic expressions! 1/9
Title highlights "FOL-Traces," a dataset for evaluating logical reasoning in language models, emphasizing rigorous testing and performance metrics.
140
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
one of my new years "considerations" is to be less silent #onhere. so i guess i'll be #here and maybe also #there til february 15th
gemini summarized my google search when i was tryna look for an anti-new years resolution blog post. it says in highlight: "Approximately 80% to 88% of New Year's resolutions fail by mid-February, ..."
000
Isabelle Lee @wordscompute.bsky.social · 11/02/2026
Really excited to receive Coefficient Giving's Technical AI Safety Research Grant via Berkeley Existential Risk Initiative w/ @nsaphra.bsky.social! We aim to predict potential AI model failures before impact--before deployment, using interpretability.
161
Reposted by Isabelle Lee
Fazl Barez @fbarez.bsky.social · 01/07/2025
Excited to share our paper: "Chain-of-Thought Is Not Explainability"! We unpack a critical misconception in AI: models explaining their steps (CoT) aren't necessarily revealing their true reasoning. Spoiler: the transparency can be an illusion. (1/9) 🧵
28531
Isabelle Lee @wordscompute.bsky.social · 10/05/2025
weve reached that point in this submission cycle, no amount of coffee will do 😞🙂‍↔️😞
010
Isabelle Lee @wordscompute.bsky.social · 29/03/2025
INCOMING
020
Isabelle Lee @wordscompute.bsky.social · 29/03/2025
titled: peer review
a leaf falls on moo deng the pygmy hippo , blocking her visionmoo deng is upset presumably because she can’t see!
071
Reposted by Isabelle Lee
Naomi Saphra @nsaphra.bsky.social · 27/03/2025
Life update: I'm starting as faculty at Boston University @bucds.bsky.social in 2026! BU has SCHEMES for LM interpretability & analysis, I couldn't be more pumped to join a burgeoning supergroup w/ @najoung.bsky.social @amuuueller.bsky.social. Looking for my first students, so apply and reach out!
CDS building which looks like a jenga tower
3524213
Isabelle Lee @wordscompute.bsky.social · 15/03/2025
really excited to be headed to OFC in SF! so excited to revisit optical physics 😀
210
Reposted by Isabelle Lee
aaditya6284.bsky.social @aaditya6284.bsky.social · 11/03/2025
Transformers employ different strategies through training to minimize loss, but how do these tradeoff and why? Excited to share our newest work, where we show remarkably rich competitive and cooperative interactions (termed "coopetition") as a transformer learns. Read on 🔎⏬
184
Reposted by Isabelle Lee
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
New paper–accepted as *spotlight* at #ICLR2025! 🧵👇 We show a competition dynamic between several algorithms splits a toy model’s ICL abilities into four broad phases of train/test settings! This means ICL is akin to a mixture of different algorithms, not a monolithic ability.
2325
Reposted by Isabelle Lee
Quanta Magazine @quantamagazine.org · 15/02/2025
Starlings move in undulating curtains across the sky. Forests of bamboo blossom at once. But some individuals don’t participate in these mystifying synchronized behaviors — and scientists are learning that they may be as important as those that do.
buff.ly
Out-of-Sync ‘Loners’ May Secretly Protect Orderly Swarms
Studies of collective behavior usually focus on how crowds of organisms coordinate their actions. But what if the individuals that don’t participate have just as much to tell us?
23310
Reposted by Isabelle Lee
Margaret Mitchell @mmitchell.bsky.social · 06/02/2025
New piece out! We explain why Fully Autonomous Agents Should Not be Developed, breaking “AI Agent” down into its components & examining through ethical values. With @evijit.io, @giadapistilli.com and @sashamtl.bsky.social huggingface.co/papers/2502....
huggingface.co
Paper page - Fully Autonomous AI Agents Should Not be Developed
Join the discussion on this paper page
414148
Reposted by Isabelle Lee
Quanta Magazine @quantamagazine.org · 05/02/2025
Brian Hie harnessed the powerful parallels between DNA and human language to create an AI tool that interprets genomes. Read his conversation with Ingrid Wickelgren: www.quantamagazine.org/the-poetry-f...
quantamagazine.org
The Poetry Fan Who Taught an LLM to Read and Write DNA | Quanta Magazine
By treating DNA as a language, Brian Hie’s “ChatGPT for genomes” could pick up patterns that humans can’t see, accelerating biological design.
14014
Reposted by Isabelle Lee
Valérie Castin @vcastin.bsky.social · 31/01/2025
How do tokens evolve as they are processed by a deep Transformer? With José A. Carrillo, @gabrielpeyre.bsky.social and @pierreablin.bsky.social, we tackle this in our new preprint: A Unified Perspective on the Dynamics of Deep Transformers arxiv.org/abs/2501.18322 ML and PDE lovers, check it out!
29616
Isabelle Lee @wordscompute.bsky.social · 26/01/2025
it’s finally raining in la:)
050
Isabelle Lee @wordscompute.bsky.social · 09/01/2025
i go on a really long walk almost every day, and at a high point in silverlake, i saw fire from all sides. and it's harder to breathe. and everything is orange.
160
Reposted by Isabelle Lee
Andrew Lee @ajyl.bsky.social · 05/01/2025
New paper <3 Interested in inference-time scaling? In-context Learning? Mech Interp? LMs can solve novel in-context tasks, with sufficient examples (longer contexts). Why? Bc they dynamically form *in-context representations*! 1/N
25316
Reposted by Isabelle Lee
Los Angeles Times @latimes.com · 09/01/2025
Hollywood High School will serve as an evacuation site for the Sunset fire in Hollywood, KTLA reported. The school is at 1521 Highland Ave. www.latimes.com/california/s...
latimes.com
Sunset fire in Hollywood Hills: Evacuations, shelter
An evacuation zone was established between the 101 Freeway and Laurel Canyon and between Mulholland Drive and Hollywood Boulevard.
231131455
Reposted by Isabelle Lee
Bálint Máté @balintmate.bsky.social · 17/12/2024
hello bluesky! we have a new preprint on solvation free energies: tl;dr: We define an interpolating density by its sampling process, and learn the corresponding equilibrium potential with score matching. arxiv.org/abs/2410.15815 with @francois.fleuret.org and @tbereau.bsky.social (1/n)
13310
Reposted by Isabelle Lee
Kate @katef.bsky.social · 15/12/2024
look at our sheep
3336
Reposted by Isabelle Lee
Arnaud Doucet @arnauddoucet.bsky.social · 15/12/2024
The slides of my NeurIPS lecture "From Diffusion Models to Schrödinger Bridges - Generative Modeling meets Optimal Transport" can be found here drive.google.com/file/d/1eLa3...
drive.google.com
BreimanLectureNeurIPS2024_Doucet.pdf
932868
Reposted by Isabelle Lee
Jennifer Hu @jennhu.bsky.social · 11/12/2024
Slides from the tutorial are now posted here! neurips.cc/media/neurip...
neurips.cc
0177
Reposted by Isabelle Lee
Sakana AI @sakanaai.bsky.social · 10/12/2024
An Evolved Universal Transformer Memory sakana.ai/namm/ Introducing Neural Attention Memory Models (NAMM), a new kind of neural memory system for Transformers that not only boost their performance and efficiency but are also transferable to other foundation models without any additional training!
14115
Reposted by Isabelle Lee
Naomi Saphra @nsaphra.bsky.social · 11/12/2024
Tomorrow (Dec 12) poster #2311! Go talk to @emalach.bsky.social and the other authors at #NeurIPS, say hi from me!
0161
Reposted by Isabelle Lee
Ethan Mollick @emollick.bsky.social · 10/12/2024
Sometimes our anthropocentric assumptions about how intelligence "should" work (like using language for reasoning) may be holding AI back. Letting AI reason in its own native "language" in latent space could unlock new capabilities, improving reasoning over Chain of Thought. arxiv.org/pdf/2412.06769
59415
Reposted by Isabelle Lee
Andrew Lampinen @lampinen.bsky.social · 10/12/2024
What counts as in-context learning (ICL)? Typically, you might think of it as learning a task from a few examples. However, we’ve just written a perspective (arxiv.org/abs/2412.03782) suggesting interpreting a much broader spectrum of behaviors as ICL! Quick summary thread: 1/7
arxiv.org
The broader spectrum of in-context learning
The ability of language models to learn a task from a few examples in context has generated substantial interest. Here, we provide a perspective that situates this type of supervised few-shot learning...
212232
Reposted by Isabelle Lee
Ferenc Huszár @inference.vc · 06/12/2024
Can language models transcend the limitations of training data? We train LMs on a formal grammar, then prompt them OUTSIDE of this grammar. We find that LMs often extrapolate logical rules and apply them OOD, too. Proof of a useful inductive bias. Check it out at NeurIPS: nips.cc/virtual/2024...
nips.cc
NeurIPS Poster Rule Extrapolation in Language Modeling: A Study of Compositional Generalization on OOD PromptsNeurIPS 2024
71138
Reposted by Isabelle Lee
Manlio De Domenico @manlius.bsky.social · 12/11/2024
The BlueSky team created a great tool to help newcomers: starter packs. Here a quick starter pack for #complexity and network scientists + feeds to quickly join the community! #NetSky Please, help to share (and if you are not in the list, get in touch and I will add you) go.bsky.app/KMfiTU2
7416884
Reposted by Isabelle Lee
Kevin Slote @kevinslote.bsky.social · 04/12/2024
go.bsky.app/Gf4uKHG Let me know if you want to be added.
244
Isabelle Lee @wordscompute.bsky.social · 03/12/2024
neural population models are so cool/wish i knew more about them ✨😲
140
Isabelle Lee @wordscompute.bsky.social · 03/12/2024
debating which is scarier, overleaf being down since 4am vs emergency martial law in korea
070
Reposted by Isabelle Lee
Martin Mundt @martinmundt.bsky.social · 27/11/2024
Is generalisation a process, an operation, or a product? 🤨 Read about the different ways generalisation is defined, parallels between humans & machines, methods & evaluation in our new paper: arxiv.org/abs/2411.15626 co-authored with many smart minds as a product of Dagstuhl 🙏🎉
092
Reposted by Isabelle Lee
vaguely reassuring state machines @happyautomata.com · 25/11/2024
😮
A randomly generated finite state machine, made of the single emoji 😮
710926
Reposted by Isabelle Lee
Kate @katef.bsky.social · 24/11/2024
friends i made a Moomins starter pack. that is, for those of us who are moomins. i don't know if this is good
8668
Reposted by Isabelle Lee
Laura @lauraruis.bsky.social · 20/11/2024
How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this: Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢 🧵⬇️
36851139
Reposted by Isabelle Lee
Maria Antoniak @mariaa.bsky.social · 22/11/2024
I made a ranked version of your feed, just for fun! Posts and replies from the last 12 hours using the HN algorithm.
2144
Reposted by Isabelle Lee
eleutherai.bsky.social @eleutherai.bsky.social · 22/11/2024
The latest from our interpretability team: there is an ambiguity in prior work on the linear representation hypothesis: Is a linear representation a linear function (that preserves the origin) or an affine function (that does not)? This distinction matters in practice. arxiv.org/abs/2411.09003
arxiv.org
Refusal in LLMs is an Affine Function
We propose affine concept editing (ACE) as an approach for steering language models' behavior by intervening directly in activations. We begin with an affine decomposition of model activation vectors ...
4837
Isabelle Lee @wordscompute.bsky.social · 18/11/2024
made something with friends on friendsgiving
A bunch of strings!! The fabric kind, not the code kind. Many different shades of teal from light to dark.Great wave off kanagawa inspired little craft project, wave embroidered on a canvas
070
Reposted by Isabelle Lee
Tim G. J. Rudner @timrudner.bsky.social · 17/11/2024
I made a starter pack for researchers in probabilistic machine learning. DM/reply if you want to be added! go.bsky.app/DuCtJqC
6610536
Isabelle Lee @wordscompute.bsky.social · 18/11/2024
i guess i should add my profile pic soon? Face tbd
000
Reposted by Isabelle Lee
Alicia Curth @aliciacurth.bsky.social · 18/11/2024
From double descent to grokking, deep learning sometimes works in unpredictable ways.. or does it? For NeurIPS(my final PhD paper!), @alanjeffares.bsky.social & I explored if&how smart linearisation can help us better understand&predict numerous odd deep learning phenomena — and learned a lot..🧵1/n
717434
Isabelle Lee @wordscompute.bsky.social · 18/11/2024
Hello world (begrudgingly)
230