Sign in

Gabriele Sarti

@gsarti.com
1.8K followers 1K following 223 posts

Open-source interpretability to seize the means of prediction. Postdoc @ Northeastern, @ndif-team.bsky.social w/ @davidbau.bsky.social. gsarti.com

PostsRepliesMedia
Gabriele Sarti @gsarti.com · 18/03/2026
tired: meta omni-translation to 1600 low-resource languages wired: kagi translate english to mechinterp
1181
Gabriele Sarti @gsarti.com · 14/03/2026
My contribution to model welfare efforts for today
2250
Gabriele Sarti @gsarti.com · 23/02/2026
Our research report on red-teaming stateful OpenClaw agents in the BauLab is finally out! 🥳 This awesome effort was led by @natalieshapira.bsky.social and involved 6 ClawBots and 20 researchers from various institutions. Check it out ➡️ agentsofchaos.baulab.info
0144
Gabriele Sarti @gsarti.com · 22/02/2026
A totally normal morning in Cambridge, MA
0120
Gabriele Sarti @gsarti.com · 21/02/2026
To accompany our paper, we also released a @hf.co space to explore agent trajectories and decoded cognitive maps from our eval outputs! Find it here: huggingface.co/spaces/proje...
1101
Gabriele Sarti @gsarti.com · 19/02/2026
Finally, we probe for multi-step plans, with a novel probe. - Pre-reasoning activations better predict longer horizons - Post-reasoning activations better predict the immediate next move. Additional evidence that reasoning seem to sharpen representations toward short-term decision-making
100
Gabriele Sarti @gsarti.com · 19/02/2026
Interestingly, we also look at how reasoning reorganises information in activations. Post reasoning cognitive map quality drops (≈75% to ≈60%), especially for agent and goal tiles, suggesting an information shift from spatial structure to task-directed action selection.
100
Gabriele Sarti @gsarti.com · 19/02/2026
But superficial failures in pursuing the task might be the product of deeper, faulty beliefs about the state of the environment! We probe model activations to decode the agent's "cognitive map" of the grid, showing that agent and goal positions are reliably but coarsely encoded by the model.
100
Gabriele Sarti @gsarti.com · 19/02/2026
We then test for instrumental and implicit goals. The agent succeeds nearly perfectly with instrumental goals (get key > unlock door > go to goal). However, task-irrelevant elements (e.g. a key without a door) can act as a distractors, likely due to training-induced semantic priors.
100
Gabriele Sarti @gsarti.com · 19/02/2026
We also test the agent robustness by applying iso-difficulty transformations (rotations, reflections, start–goal swaps), finding that performance remains statistically unchanged. This suggests the agent responds to task-relevant structure rather than superficial grid features.
110
Gabriele Sarti @gsarti.com · 19/02/2026
From a purely behavioral perspective, we see that performance scales cleanly with task difficulty. As grid size, obstacle density, and distance to the goal increase, action accuracy decreases and divergence from the optimal policy increases. This is consistent with capability-bounded goal pursuit.
121
Gabriele Sarti @gsarti.com · 19/02/2026
When we say an AI agent is “goal-directed”, what do we actually mean? In our new work, we study this question by combining behavioural and interpretability analysis in a language model agent navigating 2D grid worlds. Blog: projecttelos.substack.com/p/a-behaviou... Paper: arxiv.org/abs/2602.08964
1124
Gabriele Sarti @gsarti.com · 04/02/2026
Surely a good omen, thanks Gemini
0493
Gabriele Sarti @gsarti.com · 01/02/2026
Had a bit of fun with @anthropic.com Claude artifacts tonight and ended up with two demos that feel pretty useful for language learners: an assistive reader that lets you export/practice new words, and a translation helper that explains your mistakes. Find them here: gsarti.com/langlearn
030
Gabriele Sarti @gsarti.com · 30/01/2026
No notes, matches expectations!
1151
Gabriele Sarti @gsarti.com · 09/01/2026
Happy to announce I will be mentoring a SPAR project this Spring! ✨Check out the programme and apply by Jan 14th to work with me on understanding and mitigating implicit personalization in LLMs, i.e. how models form hidden beliefs about users that shape their responses.
170
Gabriele Sarti @gsarti.com · 04/01/2026
📣 I'm starting a postdoc at Northeastern University, where I will work on open-source NN interpretability with @davidbau.bsky.social and the @ndif-team.bsky.social. In 2026, we'll grow the NDIF ecosystem and democratize access to interpretability methods for academics and domain experts! 🚀
1270
Gabriele Sarti @gsarti.com · 16/12/2025
Big news! 🗞️ I defended my PhD thesis "From Insights to Impact: Actionable Interpretability for Neural Machine Translation" @rug.nl @gronlp.bsky.social I'm grateful to my advisors @arianna-bis.bsky.social @malvinanissim.bsky.social and to everyone who played a role in this journey! 🎉 #PhDone
2502
Gabriele Sarti @gsarti.com · 21/11/2025
Kinda crazy the improvement from Nano banana (left) to NB Pro (right): "Create an infographic explaining how model components contribute to the prediction process of a decoder-only Transformer LLM. Use the residual stream view of the Transformer by Elhage et al. (2021) in your presentation."
070
Gabriele Sarti @gsarti.com · 07/11/2025
Wrapping up my oral presentations today with our TACL paper "QE4PE: Quality Estimation for Human Post-editing" at the Interpretability morning session #EMNLP2025 (Room A104, 11:45 China time)! Paper: arxiv.org/abs/2503.03044 Slides/video/poster: underline.io/lecture/1315...
2101
Gabriele Sarti @gsarti.com · 02/10/2025
The session ended with Claude committing harakiri by deleting all DOM elements (including the chatbox for interacting with it) except the two beautiful sticky notes I asked it to make. I consider this first playing session a success!
020
Gabriele Sarti @gsarti.com · 02/10/2025
Unforeseen development
120
Gabriele Sarti @gsarti.com · 02/10/2025
What could go wrong when asking Claude to make an Imagine demo within Claude Imagine and using it to play Tic Tac Toe? When notified about the error, the model promptly adds "Sorry about that. Continue playing..." to the interface 😂
140
Gabriele Sarti @gsarti.com · 23/09/2025
I picked this expecting something close to the familiar sci-fi shorts style of Ted Chiang, but I ended up enjoying Ken Liu even more! His combination of fantastic elements with Chinese and East Asian culture and history is quite unique. Top picks: State Change, The Literomancer, The Paper Menagerie.
050
Gabriele Sarti @gsarti.com · 23/09/2025
Now with sleek flyers to test your skills in Italian crossword solving! 🤗 Join our #EVALITA2026 task!
011
Gabriele Sarti @gsarti.com · 15/09/2025
It is again the time of year when I beg @aclmeeting.bsky.social execs to rethink the current streaming platform system. For my #EMNLP2025 submissions, I am *required* to upload 2 video recordings + 2 posters + 2 slide decks. Why force both posters and talks for all? Nonsense.
0162
Gabriele Sarti @gsarti.com · 26/08/2025
TFW milk producers use semantic versioning better than LLM providers
020
Gabriele Sarti @gsarti.com · 20/08/2025
@zouharvi.bsky.social recommended this and I finally gave it a shot. Excellent read for all academics, and esp. early career people, tracing back many issues in the research landscape to a misplaced system of incentives. Will be my go-to textbook if I ever teach a research practices 101 class!
130
Gabriele Sarti @gsarti.com · 02/08/2025
After a long streak on nonfiction, I landed on "Harry Potter meets the Roman Empire". Loved the flavorful worldbuilding, the mechanics of will usage (the fantasy component, very coherent with the overall plot) and the charismatic characters. Looking forward to Part 2!
120
Gabriele Sarti @gsarti.com · 28/07/2025
Interesting stats for ACL first authors' country of affiliation! (2024 vs 2025)
250
Gabriele Sarti @gsarti.com · 21/07/2025
As a European I was very curious to read this book, which I saw heralded as the de-facto manifesto of US pro-growth libertarians. A lot of no-nonsense points, esp. on risk-taking in science, but I'm left a bit uneasy by the utopistic techno-solutionism. The elephant in the room: abundance for whom?
120
Gabriele Sarti @gsarti.com · 01/07/2025
Empire of AI was a great overview of the modern history of AI and the challenges brought by the cuttroath competition of industrial superpowers. Unlike other works toeing the "AI is all bullshit" line, Hao takes a critical but constructive stance that makes for a refreshing read. Highly recommended!
250
Gabriele Sarti @gsarti.com · 01/07/2025
Keeping up with the animal intelligence theme, Other Minds was an interesting read, although nothing special from a narrative standpoint
100
Gabriele Sarti @gsarti.com · 30/05/2025
Finally, we correlate metrics with the num. of annotators marking each token as error as more robust gold labels. We find that >3 annotations are enough for robust metric rankings, with still some margin to human-level performance → use multiple annotation sets for WQE eval! 7/
120
Gabriele Sarti @gsarti.com · 30/05/2025
XCOMETs underperform because they do not match translators' subjective error annotation propensity. Using the granular p(error) value from XCOMET significantly boost their performance when calibration is possible → desirable for a fair evaluation 6/
121
Gabriele Sarti @gsarti.com · 30/05/2025
We test predictive uncertainty (entropy, logprobs MCD avg/var), vocab projections (logitlens variants) and context mixing (attention entropy), plus XCOMET (@nunonmg.bsky.social) as supervised baselines → while costly, MCD is competitive with the 11B trained model! 5/
110
Gabriele Sarti @gsarti.com · 30/05/2025
While most evals employ a single annotation set as reference, we use our recent QE4PE dataset (arxiv.org/abs/2503.03044) to obtain up to 6 word-level error annotations per segment from different post-edits. This allows us to set a "human-level agreement baseline" for the task. 4/
110
Gabriele Sarti @gsarti.com · 30/05/2025
📢 New paper: Can unsupervised metrics extracted from MT models detect their translation errors reliably? Do annotators even *agree* on what constitutes an error? 🧐 We compare uncertainty- and interp-based WQE metrics across 12 directions, with some surprising findings! 🧵 1/
1163
Gabriele Sarti @gsarti.com · 28/05/2025
The concept-driven painting demo by Goodfire is very cool! paint.goodfire.ai Here's a manually drawn dragon with a lion head sitting atop of a pyramid in the middle of the sea with a planet in the top-right corner :)
030
Gabriele Sarti @gsarti.com · 27/05/2025
[🇮🇹 posting] Playing with Claude 4 on verbalized rebuses (aclanthology.org/2024.clicit-...), it's quite fascinating how in this case it thought that the definition "Ora è compact" (Now it's compact) could be a wordplay, and it considered compact versions of "Ora" (Now) as potential solutions 😄
020
Gabriele Sarti @gsarti.com · 13/05/2025
Excited to finally start using this gorgeous LLaMA-powered notebook I got in an indie shop in Singapore during EMNLP'23!
060
Gabriele Sarti @gsarti.com · 27/03/2025
Hats off to the GDM team for reporting negative results on SAEs while being very invested on those lines! But I can't help feeling like... www.alignmentforum.org/posts/4uXCAJ...
0120
Gabriele Sarti @gsarti.com · 05/03/2025
Fictions was my first Borges—I know, better late than never—and it was definitely a trip. Some stories, like Pierre Menard, were quite forgettable. But most of them, like the Library of Babel, Funes and Death and The Compass, were works of art. Planning to read the Aleph in the near future.
230
Gabriele Sarti @gsarti.com · 25/02/2025
Finally in Toulouse 🇫🇷 where I'll collaborate with @fannyjrd.bsky.social @antoninpoche.bsky.social and the DEEL/FOR teams at IRT & ANITI on an exciting interpretability project. Stay tuned! 🔍
0153
Gabriele Sarti @gsarti.com · 17/02/2025
"Thoughts of the Crooked-Headed Fly" was fascinating dive into the cognitive mechanisms of small animals and what perception/sensing can tell us about consciousness. As a cogsci layman, I found the concepts of efferent copy and corollary discharge super interesting! Will def read more on that.
170
Gabriele Sarti @gsarti.com · 31/01/2025
Cool! If QE4PE taught me something, it's how to select challenging data. First shot 😄
110
Gabriele Sarti @gsarti.com · 30/01/2025
Reminds me of:
110
Gabriele Sarti @gsarti.com · 15/01/2025
"Exhalations" by Ted Chiang was a great collection of sci-fi short stories with excellent plots and world-building. I loved "The Truth of Fact, the Truth of Feeling" (recommended to all AI researchers!), and the concepts behind "Oomphalos" and "The Merchant and the Alchemist's Gate". Recommended!
4151
Gabriele Sarti @gsarti.com · 10/12/2024
Despite that, most translators found word-level QE "too confusing" and "not quite accurate enough", regardless of the editing modality. Paradoxically, unsupervised highlights were sometimes rated better than those based on previous human edits! 🤔 6/
110
Gabriele Sarti @gsarti.com · 10/12/2024
While the positive effect of word-level QE on post-edit quality might not be evident at a macroscopic level (figure), we find that translators having access to highlights were 16-25% more likely to correct critical errors than those in the No Highlight modality. 5/
110