Gabriele Sarti @gsarti.com · 18/03/2026tired: meta omni-translation to 1600 low-resource languages wired: kagi translate english to mechinterp 1181
Gabriele Sarti @gsarti.com · 23/02/2026Our research report on red-teaming stateful OpenClaw agents in the BauLab is finally out! 🥳 This awesome effort was led by @natalieshapira.bsky.social and involved 6 ClawBots and 20 researchers from various institutions. Check it out ➡️ agentsofchaos.baulab.info 0144
Gabriele Sarti @gsarti.com · 21/02/2026To accompany our paper, we also released a @hf.co space to explore agent trajectories and decoded cognitive maps from our eval outputs! Find it here: huggingface.co/spaces/proje... 1101
Gabriele Sarti @gsarti.com · 19/02/2026Finally, we probe for multi-step plans, with a novel probe. - Pre-reasoning activations better predict longer horizons - Post-reasoning activations better predict the immediate next move. Additional evidence that reasoning seem to sharpen representations toward short-term decision-making 100
Gabriele Sarti @gsarti.com · 19/02/2026Interestingly, we also look at how reasoning reorganises information in activations. Post reasoning cognitive map quality drops (≈75% to ≈60%), especially for agent and goal tiles, suggesting an information shift from spatial structure to task-directed action selection. 100
Gabriele Sarti @gsarti.com · 19/02/2026But superficial failures in pursuing the task might be the product of deeper, faulty beliefs about the state of the environment! We probe model activations to decode the agent's "cognitive map" of the grid, showing that agent and goal positions are reliably but coarsely encoded by the model. 100
Gabriele Sarti @gsarti.com · 19/02/2026We then test for instrumental and implicit goals. The agent succeeds nearly perfectly with instrumental goals (get key > unlock door > go to goal). However, task-irrelevant elements (e.g. a key without a door) can act as a distractors, likely due to training-induced semantic priors. 100
Gabriele Sarti @gsarti.com · 19/02/2026We also test the agent robustness by applying iso-difficulty transformations (rotations, reflections, start–goal swaps), finding that performance remains statistically unchanged. This suggests the agent responds to task-relevant structure rather than superficial grid features. 110
Gabriele Sarti @gsarti.com · 19/02/2026From a purely behavioral perspective, we see that performance scales cleanly with task difficulty. As grid size, obstacle density, and distance to the goal increase, action accuracy decreases and divergence from the optimal policy increases. This is consistent with capability-bounded goal pursuit. 121
Gabriele Sarti @gsarti.com · 19/02/2026When we say an AI agent is “goal-directed”, what do we actually mean? In our new work, we study this question by combining behavioural and interpretability analysis in a language model agent navigating 2D grid worlds. Blog: projecttelos.substack.com/p/a-behaviou... Paper: arxiv.org/abs/2602.08964 1124
Gabriele Sarti @gsarti.com · 01/02/2026Had a bit of fun with @anthropic.com Claude artifacts tonight and ended up with two demos that feel pretty useful for language learners: an assistive reader that lets you export/practice new words, and a translation helper that explains your mistakes. Find them here: gsarti.com/langlearn 030
Gabriele Sarti @gsarti.com · 09/01/2026Happy to announce I will be mentoring a SPAR project this Spring! ✨Check out the programme and apply by Jan 14th to work with me on understanding and mitigating implicit personalization in LLMs, i.e. how models form hidden beliefs about users that shape their responses. 170
Gabriele Sarti @gsarti.com · 04/01/2026📣 I'm starting a postdoc at Northeastern University, where I will work on open-source NN interpretability with @davidbau.bsky.social and the @ndif-team.bsky.social. In 2026, we'll grow the NDIF ecosystem and democratize access to interpretability methods for academics and domain experts! 🚀 1270
Gabriele Sarti @gsarti.com · 16/12/2025Big news! 🗞️ I defended my PhD thesis "From Insights to Impact: Actionable Interpretability for Neural Machine Translation" @rug.nl @gronlp.bsky.social I'm grateful to my advisors @arianna-bis.bsky.social @malvinanissim.bsky.social and to everyone who played a role in this journey! 🎉 #PhDone 2502
Gabriele Sarti @gsarti.com · 21/11/2025Kinda crazy the improvement from Nano banana (left) to NB Pro (right): "Create an infographic explaining how model components contribute to the prediction process of a decoder-only Transformer LLM. Use the residual stream view of the Transformer by Elhage et al. (2021) in your presentation." 070
Gabriele Sarti @gsarti.com · 07/11/2025Wrapping up my oral presentations today with our TACL paper "QE4PE: Quality Estimation for Human Post-editing" at the Interpretability morning session #EMNLP2025 (Room A104, 11:45 China time)! Paper: arxiv.org/abs/2503.03044 Slides/video/poster: underline.io/lecture/1315... 2101
Gabriele Sarti @gsarti.com · 02/10/2025The session ended with Claude committing harakiri by deleting all DOM elements (including the chatbox for interacting with it) except the two beautiful sticky notes I asked it to make. I consider this first playing session a success! 020
Gabriele Sarti @gsarti.com · 02/10/2025What could go wrong when asking Claude to make an Imagine demo within Claude Imagine and using it to play Tic Tac Toe? When notified about the error, the model promptly adds "Sorry about that. Continue playing..." to the interface 😂 140
Gabriele Sarti @gsarti.com · 23/09/2025I picked this expecting something close to the familiar sci-fi shorts style of Ted Chiang, but I ended up enjoying Ken Liu even more! His combination of fantastic elements with Chinese and East Asian culture and history is quite unique. Top picks: State Change, The Literomancer, The Paper Menagerie. 050
Gabriele Sarti @gsarti.com · 23/09/2025Now with sleek flyers to test your skills in Italian crossword solving! 🤗 Join our #EVALITA2026 task! 011
Gabriele Sarti @gsarti.com · 15/09/2025It is again the time of year when I beg @aclmeeting.bsky.social execs to rethink the current streaming platform system. For my #EMNLP2025 submissions, I am *required* to upload 2 video recordings + 2 posters + 2 slide decks. Why force both posters and talks for all? Nonsense. 0162
Gabriele Sarti @gsarti.com · 26/08/2025TFW milk producers use semantic versioning better than LLM providers 020
Gabriele Sarti @gsarti.com · 20/08/2025@zouharvi.bsky.social recommended this and I finally gave it a shot. Excellent read for all academics, and esp. early career people, tracing back many issues in the research landscape to a misplaced system of incentives. Will be my go-to textbook if I ever teach a research practices 101 class! 130
Gabriele Sarti @gsarti.com · 02/08/2025After a long streak on nonfiction, I landed on "Harry Potter meets the Roman Empire". Loved the flavorful worldbuilding, the mechanics of will usage (the fantasy component, very coherent with the overall plot) and the charismatic characters. Looking forward to Part 2! 120
Gabriele Sarti @gsarti.com · 28/07/2025Interesting stats for ACL first authors' country of affiliation! (2024 vs 2025) 250
Gabriele Sarti @gsarti.com · 21/07/2025As a European I was very curious to read this book, which I saw heralded as the de-facto manifesto of US pro-growth libertarians. A lot of no-nonsense points, esp. on risk-taking in science, but I'm left a bit uneasy by the utopistic techno-solutionism. The elephant in the room: abundance for whom? 120
Gabriele Sarti @gsarti.com · 01/07/2025Empire of AI was a great overview of the modern history of AI and the challenges brought by the cuttroath competition of industrial superpowers. Unlike other works toeing the "AI is all bullshit" line, Hao takes a critical but constructive stance that makes for a refreshing read. Highly recommended! 250
Gabriele Sarti @gsarti.com · 01/07/2025Keeping up with the animal intelligence theme, Other Minds was an interesting read, although nothing special from a narrative standpoint 100
Gabriele Sarti @gsarti.com · 30/05/2025Finally, we correlate metrics with the num. of annotators marking each token as error as more robust gold labels. We find that >3 annotations are enough for robust metric rankings, with still some margin to human-level performance → use multiple annotation sets for WQE eval! 7/ 120
Gabriele Sarti @gsarti.com · 30/05/2025XCOMETs underperform because they do not match translators' subjective error annotation propensity. Using the granular p(error) value from XCOMET significantly boost their performance when calibration is possible → desirable for a fair evaluation 6/ 121
Gabriele Sarti @gsarti.com · 30/05/2025We test predictive uncertainty (entropy, logprobs MCD avg/var), vocab projections (logitlens variants) and context mixing (attention entropy), plus XCOMET (@nunonmg.bsky.social) as supervised baselines → while costly, MCD is competitive with the 11B trained model! 5/ 110
Gabriele Sarti @gsarti.com · 30/05/2025While most evals employ a single annotation set as reference, we use our recent QE4PE dataset (arxiv.org/abs/2503.03044) to obtain up to 6 word-level error annotations per segment from different post-edits. This allows us to set a "human-level agreement baseline" for the task. 4/ 110
Gabriele Sarti @gsarti.com · 30/05/2025📢 New paper: Can unsupervised metrics extracted from MT models detect their translation errors reliably? Do annotators even *agree* on what constitutes an error? 🧐 We compare uncertainty- and interp-based WQE metrics across 12 directions, with some surprising findings! 🧵 1/ 1163
Gabriele Sarti @gsarti.com · 28/05/2025The concept-driven painting demo by Goodfire is very cool! paint.goodfire.ai Here's a manually drawn dragon with a lion head sitting atop of a pyramid in the middle of the sea with a planet in the top-right corner :) 030
Gabriele Sarti @gsarti.com · 27/05/2025[🇮🇹 posting] Playing with Claude 4 on verbalized rebuses (aclanthology.org/2024.clicit-...), it's quite fascinating how in this case it thought that the definition "Ora è compact" (Now it's compact) could be a wordplay, and it considered compact versions of "Ora" (Now) as potential solutions 😄 020
Gabriele Sarti @gsarti.com · 13/05/2025Excited to finally start using this gorgeous LLaMA-powered notebook I got in an indie shop in Singapore during EMNLP'23! 060
Gabriele Sarti @gsarti.com · 27/03/2025Hats off to the GDM team for reporting negative results on SAEs while being very invested on those lines! But I can't help feeling like... www.alignmentforum.org/posts/4uXCAJ... 0120
Gabriele Sarti @gsarti.com · 05/03/2025Fictions was my first Borges—I know, better late than never—and it was definitely a trip. Some stories, like Pierre Menard, were quite forgettable. But most of them, like the Library of Babel, Funes and Death and The Compass, were works of art. Planning to read the Aleph in the near future. 230
Gabriele Sarti @gsarti.com · 25/02/2025Finally in Toulouse 🇫🇷 where I'll collaborate with @fannyjrd.bsky.social @antoninpoche.bsky.social and the DEEL/FOR teams at IRT & ANITI on an exciting interpretability project. Stay tuned! 🔍 0153
Gabriele Sarti @gsarti.com · 17/02/2025"Thoughts of the Crooked-Headed Fly" was fascinating dive into the cognitive mechanisms of small animals and what perception/sensing can tell us about consciousness. As a cogsci layman, I found the concepts of efferent copy and corollary discharge super interesting! Will def read more on that. 170
Gabriele Sarti @gsarti.com · 31/01/2025Cool! If QE4PE taught me something, it's how to select challenging data. First shot 😄 110
Gabriele Sarti @gsarti.com · 15/01/2025"Exhalations" by Ted Chiang was a great collection of sci-fi short stories with excellent plots and world-building. I loved "The Truth of Fact, the Truth of Feeling" (recommended to all AI researchers!), and the concepts behind "Oomphalos" and "The Merchant and the Alchemist's Gate". Recommended! 4151
Gabriele Sarti @gsarti.com · 10/12/2024Despite that, most translators found word-level QE "too confusing" and "not quite accurate enough", regardless of the editing modality. Paradoxically, unsupervised highlights were sometimes rated better than those based on previous human edits! 🤔 6/ 110
Gabriele Sarti @gsarti.com · 10/12/2024While the positive effect of word-level QE on post-edit quality might not be evident at a macroscopic level (figure), we find that translators having access to highlights were 16-25% more likely to correct critical errors than those in the No Highlight modality. 5/ 110