Reposted by Gabriele SartiMartin Tutek @mtutek.bsky.social · 09/09/2026Take a break from reading up on Navier-Stokes drama and add October 29th to your calendar, as we have a stellar lineup of speakers at this years' BlackboxNLP! 051
Reposted by Gabriele SartiBlackboxNLP @blackboxnlp.bsky.social · 09/09/2026🎉 Thrilled to announce the keynote speakers for BlackboxNLP 2026 @emnlpmeeting, with three incredible perspectives on interpretability: 🔍 Ivan Titov 🔍 Sheridan Feucht @sfeucht.bsky.social 🔍 Michael Hahn @m-hahn.bsky.social Join us on October 29th! 🇭🇺 0104
Gabriele Sarti @gsarti.com · 04/09/2026The report is finally out, titanic effort by @zouhar.bsky.social and the team, check it out! 160
Reposted by Gabriele SartiBlackboxNLP @blackboxnlp.bsky.social · 31/08/2026Due to the unprecedented amount of submissions to BlackboxNLP 2026, we will announce the decisions for Main, Special and ARR commitments on September 2 AoE. Thanks for your patience! 063
Reposted by Gabriele SartiMartin Tutek @mtutek.bsky.social · 22/08/2026The @blackboxnlp.bsky.social reproducibility track is in an urgent need for reviewers! If you can review a paper (or two), these bite-sized papers deserve some attention, and reviewing them will leave you fulfilled! 🥰 Please reach out if you can help out! RTs appreciated 👐 035
Gabriele Sarti @gsarti.com · 17/08/2026We are still looking for emergency reviewers for BlackboxNLP 2026. Please sign up at the link below if you have time to review 1-2 papers before Aug 20 AoE! 🙏 033
Gabriele Sarti @gsarti.com · 23/07/2026An unprecedented number of submissions means an unprecedented need for reviewers! If you have published work in interpretability, please consider signing up to review ⬇️ 062
Gabriele Sarti @gsarti.com · 17/07/2026Happy to be a small part of this colossal effort led by @zouharvi.bsky.social to reinvent the future of machine translation evaluation! 🔡 Join us in building the LTB and become a coauthor, help wanted! 073
Reposted by Gabriele SartiBlackboxNLP @blackboxnlp.bsky.social · 16/07/2026⏳ The BlackboxNLP 2026 Reproducibility Challenge deadline has been extended to July 24 (AoE) ⏳ If you've been working on a robustness check, ablation, or replication of recent NLP interpretability work, you have a bit more time to get your submission in. 186
Reposted by Gabriele SartiDaniel Scalena @danielsc4.it · 01/07/2026I'd never have guessed models commit to their final answer this early, often within the first 20% of reasoning, across math/logic tasks and model families. The rest is mostly hedging that doesn't change their mind. And turns out they encode this internally, we can decode it! 🧵👇 093
Gabriele Sarti @gsarti.com · 01/07/2026🚨New paper led by @danielsc4.it & @saracandussio.bsky.social studying answer commitment across CoT for math & logic reasoning! Turns out, models often converge to final answers early in the CoT, then tend to "fake" hedging & re-checks - and we can tell apart mid/final guesses quite robustly! See 👇 1183
Gabriele Sarti @gsarti.com · 01/07/2026Our @ndif-team.bsky.social was thrilled to sponsor the BlackboxNLP repro challenge! Interpretability needs more robust & generalizable findings - if you're working in this space, this should be on your radar! 030
Gabriele Sarti @gsarti.com · 18/06/2026It was very nice to be back in the Netherlands and EAMT to receive this prize! Thanks to everyone involved 🤗 0160
Reposted by Gabriele SartiYonatan Belinkov @boknilev.bsky.social · 17/06/2026Are you wondering if LLM interpretability results generalize, reproduce, etc.? Check out the reproducibility challenge and submit your work reproducing papers in this area: bsky.app/profile/blac... 1156
Reposted by Gabriele SartiTomer Ullman @tomerullman.bsky.social · 15/06/2026"the new toaster says 'I Love You' when you put the toast in, and 'I'm Sorry' when it burns it. This causes some people to get angry b/c they don't think the toaster means it, and others to develop unhealthy attachments b/c they think it does. the solution is to make the toaster TRULY sorry" 2349
Reposted by Gabriele SartiNaomi Saphra @nsaphra.bsky.social · 15/06/2026We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social 313734
Reposted by Gabriele SartiDavid Bau @davidbau.bsky.social · 14/06/2026I recently spoke with Yascha Mounk about how researchers look inside AI to understand how it is thinking. Here is the podcast: writing.yaschamounk.com/p/david-bau-2writing.yaschamounk.comDavid Bau on How—and Whether—Artificial Intelligence ThinksYascha Mounk and David Bau examine the mysterious internal processes that drive AI behavior—and why they may be fundamentally alien. 162
Reposted by Gabriele Sartipeterdoohan.bsky.social @peterdoohan.bsky.social · 11/06/2026How do brains plan actions towards goals? To get at this question we studied mice navigating complex mazes as goals changed on every trial 🧵 Work with @thomasakam.bsky.social @behrenstimb.bsky.social @kristorpjensen.bsky.social now on BioRxiv: www.biorxiv.org/content/10.6... 411846
Reposted by Gabriele SartiGeoffrey Irving @girving.bsky.social · 10/06/2026We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵 sequent.org/launch 4172
Reposted by Gabriele SartiAaron Mueller @amuuueller.bsky.social · 10/06/2026The New England Mechanistic Interpretability (NEMI) workshop is coming to BU on Aug. 14! Join us for talks, a panel, food, and plenty of opportunities to connect with the many great researchers in the area. Register and help spread the word! 0176
Reposted by Gabriele SartiEhud Reiter @ehudreiter.bsky.social · 08/06/2026New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...ehudreiter.comI am worried by NLP research cultureIn most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have… 1144
Reposted by Gabriele SartiDavid Bau @davidbau.bsky.social · 05/06/2026"You're right to call me on that!" Can you catch an AI in the act of lying? Register below to enter our AI lie-detection contest. AI lies are a big problem. The frontier labs have all worked hard to fight AI deception. They all try to monitor their AIs for it. 173
Gabriele Sarti @gsarti.com · 05/06/2026Check out our latest work led by @veraneplenbroek.bsky.social! Conversation topics seem much better explanations for LLMs behavior over latent sociodemographic info about the user - but what does the choice of topic tell us about the speaker in the first place? 051
Gabriele Sarti @gsarti.com · 04/06/2026Check out our new hackathon for building the best AI lie detector on the market! A lot of interesting questions and great prizes for participants, apply early if you want in! :) 041
Reposted by Gabriele SartiMartin Tutek @mtutek.bsky.social · 02/06/2026With the large influx of submissions and a faster pace of research, reproducibility is more important than ever. With this reproducibility challenge, we want to put the focus on best practices wrt. baselines🧱, ablations🌈, eval🔎 and generalizability🗺️ of interpretability! 073
Gabriele Sarti @gsarti.com · 28/05/2026Despite the huge inflow of researchers, much of the work in interpretability remains anecdotal. Our new repro challenge at BlackboxNLP (co-located with EMNLP 2026) aims to attract work challenging common assumptions and showing failure/success cases of popular methods. Negative results welcome! 181
Reposted by Gabriele SartiDaniel Lowd @dlowd.com · 22/05/2026As a computer scientist, we INVENTED the phrase "artificial intelligence." It was never exclusively yours. In both fact and fiction, AI has always been awesome, and beautiful, and, yes, problematic — and it still is. 2653
Gabriele Sarti @gsarti.com · 13/05/2026Excellent survey on causal interpretability by @amuuueller.bsky.social and many BauLab members, don't miss it! 1133
Reposted by Gabriele SartiProf Dynarski @dynarski.bsky.social · 09/05/2026IMO a key skill of a good scientist is moving comfortably between the specific & general e.g., relentless in understand the nerdy details of the data (including its coding) WHILE holding onto the big picture of the hypotheses being tested with the data 17610
Reposted by Gabriele SartiAi2 @ai2.bsky.social · 08/05/2026Today we’re releasing EMO, a new mixture-of-experts (MoE) model trained so modular structure emerges directly from data without human-defined priors. EMO can use a small subset of its experts for a given task while keeping near full-model performance. 🧵 217223
Reposted by Gabriele SartiPaul Röttger @paul-rottger.bsky.social · 27/04/2026New paper w/ UK AISI: Millions of people now use AI to help them write and communicate. In three experiments (14k participants, 3m+ human ratings) we show that AI writing assistance systematically distorts writer personas – their perceived beliefs, personality, and identity. 🧵 24413
Reposted by Gabriele SartiHarrison Ritz @hritz.bsky.social · 24/04/2026🚨 Tom Griffiths has a podcast where he interviews cognitive scientists podcasts.apple.com/ca/podcast/t... This just went to the top of my list.podcasts.apple.comThe Cognition ProjectScience Podcast · How can we study the mind, something we can never see or touch? This podcast tells the story of how psychologists, neuroscientists, computer scientists, linguists, and philosophers w... 06416
Reposted by Gabriele SartiDavid Bau @davidbau.bsky.social · 20/04/20262026 is a whirlwind year for AI. Underlying it all is the greatest scientific mystery of our age. How does a neural network think? I talked w Oliver Whang in NYTimes Magazine, on how AI interpretability is a tangle of structure waiting to be unraveled: www.nytimes.com/2026/04/15/... 2132
Reposted by Gabriele SartiDavid Bau @davidbau.bsky.social · 08/04/2026Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top". 1133
Gabriele Sarti @gsarti.com · 02/04/2026Mfw fiddling with probes all day but patching experiments don't pan out 051
Reposted by Gabriele SartiDavid Bau @davidbau.bsky.social · 25/03/2026Calling attention to an exciting "deception detection" hackathon we're planning this summer! w @NDIF and @CadenzaLabs. Recruiting red teams now, blue teams later. Red teams, time is short: proposals due Mar 31. $10K stipend + compute, $15K finals prize. nnsight.net/blog/2026/0... 252
Reposted by Gabriele SartiMicah Benson @micahben.bsky.social · 25/03/2026I truly believe the rapid advances in the mech interp subfield have something real to offer AI ethics researchers: A chance to look beyond the HOW of evals to the WHY, a first pass at a technical solution when we see the opportunity, a new avenue for showing failures that prove models are not gods 173
Reposted by Gabriele SartiNathan Godey @nthngdy.bsky.social · 12/03/2026🧵New paper: "Lost in Backpropagation: The LM Head is a Gradient Bottleneck" The output layer of LLMs destroys 95-99% of your training signal during backpropagation, and this significantly slows down pretraining 👇 610815
Gabriele Sarti @gsarti.com · 23/03/2026Check out David's NetHack port! "Complexity does not yield to speed. Judgment remains essential. The work of deciding what matters, of seeing what is hidden, of knowing when your own metrics are lying to you: this is the work that remains, and it is the work worth learning." 030
Gabriele Sarti @gsarti.com · 19/03/2026This is an important project! If you believe alignment faking is true, you should at least entertain the possibility of misalignment faking before drawing your conclusions. Especially true if researchers fishing for misaligned behaviors are the ones running the evals! 1211
Gabriele Sarti @gsarti.com · 18/03/2026BlackboxNLP is back once again at EMNLP'26! Very happy to be part of the team again, and excited for our new reproducibility track! Check it out ⬇️ 0143
Gabriele Sarti @gsarti.com · 18/03/2026tired: meta omni-translation to 1600 low-resource languages wired: kagi translate english to mechinterp 1181
Reposted by Gabriele SartiAvery Yen @averyyen.bsky.social · 15/03/2026I'm calling it DeepSeek's new 1T parameter model (V4)? The style, content, and length of the reasoning are extremely similar. 153
Gabriele Sarti @gsarti.com · 10/03/2026This morning I happened to hang out around the Harvard med school café and all conversations I overheard were about LLMs med assistants and XAI 🫡 170
Reposted by Gabriele SartiMartin Wattenberg @wattenberg.bsky.social · 07/03/2026I want to talk about why AI-based mass surveillance is so dangerous, and why I would oppose it no matter which party or president is in office. 34910
Reposted by Gabriele SartiAntonin Poché @antoninpoche.bsky.social · 04/03/2026🔥Super excited to share our new demo website for 🪄Interpreto! 🖼️It is basically an explanation gallery showcasing attribution and concept-based explanations for classification and generation. 🎮Play with it: for-sight-ai.github.io/interpreto-d... We will keep improving it, so stay tuned! 193
Reposted by Gabriele SartiAlessio Miaschi @alessiomiaschi.bsky.social · 02/03/2026Great wrap-up for #EVALITA2026! 🔥 Glad to have helped organize this edition and to see many interesting discussions! Great response to our task Cruciverb-IT (with Ciaccio C., @gsarti.com, Dell’Orletta F., @malvinanissim.bsky.social)! Thanks to all co-organizers and @ailc-nlp.bsky.social! #NLProc 072
Gabriele Sarti @gsarti.com · 27/02/2026Great release from our engineering team! A lot of the major pain points have been addressed, and this is our first step towards supporting interpretability workflows on more realistic scenarios! Check it out! 051