Sign in

Gabriele Sarti

@gsarti.com
1.8K followers 1K following 223 posts

Open-source interpretability to seize the means of prediction. Postdoc @ Northeastern, @ndif-team.bsky.social w/ @davidbau.bsky.social. gsarti.com

PostsRepliesMedia
Reposted by Gabriele Sarti
Martin Tutek @mtutek.bsky.social · 09/09/2026
Take a break from reading up on Navier-Stokes drama and add October 29th to your calendar, as we have a stellar lineup of speakers at this years' BlackboxNLP!
051
Reposted by Gabriele Sarti
BlackboxNLP @blackboxnlp.bsky.social · 09/09/2026
🎉 Thrilled to announce the keynote speakers for BlackboxNLP 2026 @emnlpmeeting, with three incredible perspectives on interpretability: 🔍 Ivan Titov 🔍 Sheridan Feucht @sfeucht.bsky.social 🔍 Michael Hahn @m-hahn.bsky.social Join us on October 29th! 🇭🇺
0104
Gabriele Sarti @gsarti.com · 04/09/2026
The report is finally out, titanic effort by @zouhar.bsky.social and the team, check it out!
160
Reposted by Gabriele Sarti
BlackboxNLP @blackboxnlp.bsky.social · 31/08/2026
Due to the unprecedented amount of submissions to BlackboxNLP 2026, we will announce the decisions for Main, Special and ARR commitments on September 2 AoE. Thanks for your patience!
063
Reposted by Gabriele Sarti
Martin Tutek @mtutek.bsky.social · 22/08/2026
The @blackboxnlp.bsky.social reproducibility track is in an urgent need for reviewers! If you can review a paper (or two), these bite-sized papers deserve some attention, and reviewing them will leave you fulfilled! 🥰 Please reach out if you can help out! RTs appreciated 👐
035
Gabriele Sarti @gsarti.com · 17/08/2026
We are still looking for emergency reviewers for BlackboxNLP 2026. Please sign up at the link below if you have time to review 1-2 papers before Aug 20 AoE! 🙏
033
Gabriele Sarti @gsarti.com · 23/07/2026
An unprecedented number of submissions means an unprecedented need for reviewers! If you have published work in interpretability, please consider signing up to review ⬇️
062
Gabriele Sarti @gsarti.com · 17/07/2026
Happy to be a small part of this colossal effort led by @zouharvi.bsky.social to reinvent the future of machine translation evaluation! 🔡 Join us in building the LTB and become a coauthor, help wanted!
073
Reposted by Gabriele Sarti
BlackboxNLP @blackboxnlp.bsky.social · 16/07/2026
⏳ The BlackboxNLP 2026 Reproducibility Challenge deadline has been extended to July 24 (AoE) ⏳ If you've been working on a robustness check, ablation, or replication of recent NLP interpretability work, you have a bit more time to get your submission in.
186
Reposted by Gabriele Sarti
Daniel Scalena @danielsc4.it · 01/07/2026
I'd never have guessed models commit to their final answer this early, often within the first 20% of reasoning, across math/logic tasks and model families. The rest is mostly hedging that doesn't change their mind. And turns out they encode this internally, we can decode it! 🧵👇
093
Gabriele Sarti @gsarti.com · 01/07/2026
🚨New paper led by @danielsc4.it & @saracandussio.bsky.social studying answer commitment across CoT for math & logic reasoning! Turns out, models often converge to final answers early in the CoT, then tend to "fake" hedging & re-checks - and we can tell apart mid/final guesses quite robustly! See 👇
1183
Gabriele Sarti @gsarti.com · 01/07/2026
Our @ndif-team.bsky.social was thrilled to sponsor the BlackboxNLP repro challenge! Interpretability needs more robust & generalizable findings - if you're working in this space, this should be on your radar!
030
Gabriele Sarti @gsarti.com · 18/06/2026
It was very nice to be back in the Netherlands and EAMT to receive this prize! Thanks to everyone involved 🤗
0160
Reposted by Gabriele Sarti
Yonatan Belinkov @boknilev.bsky.social · 17/06/2026
Are you wondering if LLM interpretability results generalize, reproduce, etc.? Check out the reproducibility challenge and submit your work reproducing papers in this area: bsky.app/profile/blac...
1156
Reposted by Gabriele Sarti
Tomer Ullman @tomerullman.bsky.social · 15/06/2026
"the new toaster says 'I Love You' when you put the toast in, and 'I'm Sorry' when it burns it. This causes some people to get angry b/c they don't think the toaster means it, and others to develop unhealthy attachments b/c they think it does. the solution is to make the toaster TRULY sorry"
2349
Reposted by Gabriele Sarti
Naomi Saphra @nsaphra.bsky.social · 15/06/2026
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social
313734
Reposted by Gabriele Sarti
David Bau @davidbau.bsky.social · 14/06/2026
I recently spoke with Yascha Mounk about how researchers look inside AI to understand how it is thinking. Here is the podcast: writing.yaschamounk.com/p/david-bau-2
writing.yaschamounk.com
David Bau on How—and Whether—Artificial Intelligence Thinks
Yascha Mounk and David Bau examine the mysterious internal processes that drive AI behavior—and why they may be fundamentally alien.
162
Reposted by Gabriele Sarti
peterdoohan.bsky.social @peterdoohan.bsky.social · 11/06/2026
How do brains plan actions towards goals? To get at this question we studied mice navigating complex mazes as goals changed on every trial 🧵 Work with @thomasakam.bsky.social @behrenstimb.bsky.social @kristorpjensen.bsky.social now on BioRxiv: www.biorxiv.org/content/10.6...
411846
Reposted by Gabriele Sarti
Geoffrey Irving @girving.bsky.social · 10/06/2026
We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵 sequent.org/launch
4172
Reposted by Gabriele Sarti
Aaron Mueller @amuuueller.bsky.social · 10/06/2026
The New England Mechanistic Interpretability (NEMI) workshop is coming to BU on Aug. 14! Join us for talks, a panel, food, and plenty of opportunities to connect with the many great researchers in the area. Register and help spread the word!
0176
Reposted by Gabriele Sarti
Ehud Reiter @ehudreiter.bsky.social · 08/06/2026
New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...
ehudreiter.com
I am worried by NLP research culture
In most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have…
1144
Reposted by Gabriele Sarti
David Bau @davidbau.bsky.social · 05/06/2026
"You're right to call me on that!" Can you catch an AI in the act of lying? Register below to enter our AI lie-detection contest. AI lies are a big problem. The frontier labs have all worked hard to fight AI deception. They all try to monitor their AIs for it.
173
Gabriele Sarti @gsarti.com · 05/06/2026
Check out our latest work led by @veraneplenbroek.bsky.social! Conversation topics seem much better explanations for LLMs behavior over latent sociodemographic info about the user - but what does the choice of topic tell us about the speaker in the first place?
051
Gabriele Sarti @gsarti.com · 04/06/2026
Check out our new hackathon for building the best AI lie detector on the market! A lot of interesting questions and great prizes for participants, apply early if you want in! :)
041
Reposted by Gabriele Sarti
Martin Tutek @mtutek.bsky.social · 02/06/2026
With the large influx of submissions and a faster pace of research, reproducibility is more important than ever. With this reproducibility challenge, we want to put the focus on best practices wrt. baselines🧱, ablations🌈, eval🔎 and generalizability🗺️ of interpretability!
073
Gabriele Sarti @gsarti.com · 28/05/2026
Despite the huge inflow of researchers, much of the work in interpretability remains anecdotal. Our new repro challenge at BlackboxNLP (co-located with EMNLP 2026) aims to attract work challenging common assumptions and showing failure/success cases of popular methods. Negative results welcome!
181
Reposted by Gabriele Sarti
Daniel Lowd @dlowd.com · 22/05/2026
As a computer scientist, we INVENTED the phrase "artificial intelligence." It was never exclusively yours. In both fact and fiction, AI has always been awesome, and beautiful, and, yes, problematic — and it still is.
2653
Gabriele Sarti @gsarti.com · 13/05/2026
Excellent survey on causal interpretability by @amuuueller.bsky.social and many BauLab members, don't miss it!
1133
Reposted by Gabriele Sarti
Prof Dynarski @dynarski.bsky.social · 09/05/2026
IMO a key skill of a good scientist is moving comfortably between the specific & general e.g., relentless in understand the nerdy details of the data (including its coding) WHILE holding onto the big picture of the hypotheses being tested with the data
17610
Reposted by Gabriele Sarti
Ai2 @ai2.bsky.social · 08/05/2026
Today we’re releasing EMO, a new mixture-of-experts (MoE) model trained so modular structure emerges directly from data without human-defined priors. EMO can use a small subset of its experts for a given task while keeping near full-model performance. 🧵
217223
Reposted by Gabriele Sarti
Paul Röttger @paul-rottger.bsky.social · 27/04/2026
New paper w/ UK AISI: Millions of people now use AI to help them write and communicate. In three experiments (14k participants, 3m+ human ratings) we show that AI writing assistance systematically distorts writer personas – their perceived beliefs, personality, and identity. 🧵
24413
Reposted by Gabriele Sarti
Harrison Ritz @hritz.bsky.social · 24/04/2026
🚨 Tom Griffiths has a podcast where he interviews cognitive scientists podcasts.apple.com/ca/podcast/t... This just went to the top of my list.
podcasts.apple.com
The Cognition Project
Science Podcast · How can we study the mind, something we can never see or touch? This podcast tells the story of how psychologists, neuroscientists, computer scientists, linguists, and philosophers w...
06416
Reposted by Gabriele Sarti
David Bau @davidbau.bsky.social · 20/04/2026
2026 is a whirlwind year for AI. Underlying it all is the greatest scientific mystery of our age. How does a neural network think? I talked w Oliver Whang in NYTimes Magazine, on how AI interpretability is a tangle of structure waiting to be unraveled: www.nytimes.com/2026/04/15/...
2132
Reposted by Gabriele Sarti
David Bau @davidbau.bsky.social · 08/04/2026
Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top".
1133
Gabriele Sarti @gsarti.com · 02/04/2026
Mfw fiddling with probes all day but patching experiments don't pan out
051
Gabriele Sarti @gsarti.com · 26/03/2026
Thank you for having me! Next time in person! 🤗
050
Reposted by Gabriele Sarti
David Bau @davidbau.bsky.social · 25/03/2026
Calling attention to an exciting "deception detection" hackathon we're planning this summer! w @NDIF and @CadenzaLabs. Recruiting red teams now, blue teams later. Red teams, time is short: proposals due Mar 31. $10K stipend + compute, $15K finals prize. nnsight.net/blog/2026/0...
252
Reposted by Gabriele Sarti
Micah Benson @micahben.bsky.social · 25/03/2026
I truly believe the rapid advances in the mech interp subfield have something real to offer AI ethics researchers: A chance to look beyond the HOW of evals to the WHY, a first pass at a technical solution when we see the opportunity, a new avenue for showing failures that prove models are not gods
173
Reposted by Gabriele Sarti
Nathan Godey @nthngdy.bsky.social · 12/03/2026
🧵New paper: "Lost in Backpropagation: The LM Head is a Gradient Bottleneck" The output layer of LLMs destroys 95-99% of your training signal during backpropagation, and this significantly slows down pretraining 👇
610815
Gabriele Sarti @gsarti.com · 23/03/2026
Check out David's NetHack port! "Complexity does not yield to speed. Judgment remains essential. The work of deciding what matters, of seeing what is hidden, of knowing when your own metrics are lying to you: this is the work that remains, and it is the work worth learning."
030
Gabriele Sarti @gsarti.com · 19/03/2026
This is an important project! If you believe alignment faking is true, you should at least entertain the possibility of misalignment faking before drawing your conclusions. Especially true if researchers fishing for misaligned behaviors are the ones running the evals!
1211
Gabriele Sarti @gsarti.com · 18/03/2026
BlackboxNLP is back once again at EMNLP'26! Very happy to be part of the team again, and excited for our new reproducibility track! Check it out ⬇️
0143
Gabriele Sarti @gsarti.com · 18/03/2026
tired: meta omni-translation to 1600 low-resource languages wired: kagi translate english to mechinterp
1181
Reposted by Gabriele Sarti
Avery Yen @averyyen.bsky.social · 15/03/2026
I'm calling it DeepSeek's new 1T parameter model (V4)? The style, content, and length of the reasoning are extremely similar.
153
Gabriele Sarti @gsarti.com · 14/03/2026
My contribution to model welfare efforts for today
2250
Gabriele Sarti @gsarti.com · 10/03/2026
This morning I happened to hang out around the Harvard med school café and all conversations I overheard were about LLMs med assistants and XAI 🫡
170
Reposted by Gabriele Sarti
Martin Wattenberg @wattenberg.bsky.social · 07/03/2026
I want to talk about why AI-based mass surveillance is so dangerous, and why I would oppose it no matter which party or president is in office.
34910
Reposted by Gabriele Sarti
Antonin Poché @antoninpoche.bsky.social · 04/03/2026
🔥Super excited to share our new demo website for 🪄Interpreto! 🖼️It is basically an explanation gallery showcasing attribution and concept-based explanations for classification and generation. 🎮Play with it: for-sight-ai.github.io/interpreto-d... We will keep improving it, so stay tuned!
193
Reposted by Gabriele Sarti
Alessio Miaschi @alessiomiaschi.bsky.social · 02/03/2026
Great wrap-up for #EVALITA2026! 🔥 Glad to have helped organize this edition and to see many interesting discussions! Great response to our task Cruciverb-IT (with Ciaccio C., @gsarti.com, Dell’Orletta F., @malvinanissim.bsky.social)! Thanks to all co-organizers and @ailc-nlp.bsky.social! #NLProc
072
Gabriele Sarti @gsarti.com · 27/02/2026
Great release from our engineering team! A lot of the major pain points have been addressed, and this is our first step towards supporting interpretability workflows on more realistic scenarios! Check it out!
051