Sign in

Antonin Poché

@antoninpoche.bsky.social
101 followers 115 following 53 posts

PhD Student doing XAI for NLP at @ANITI_Toulouse, IRIT, and IRT Saint Exupery. 🛠️ Interpreto & Xplique library development team member. antoninpoche.github.io

PostsRepliesMedia
Reposted by Antonin Poché
Nils Feldhus @nfel.bsky.social · 22/09/2026
Chain-of-Thought reasoning can sound plausible while being unfaithful. Is there a way to test that directly by looking inside the model? Our new work led by @qiaw99.bsky.social w/ @apepa.bsky.social et al. 📜 Pre-print: arxiv.org/abs/2609.23065
151
Antonin Poché @antoninpoche.bsky.social · 09/09/2026
I am both excited🔥and worried❄️. 🔥We got a paper accepted to @blackboxnlp.bsky.social reproducibility track. ❄️It reproduces and destroys my own paper. So I basically have 1 PhD year, my scientific integrity, and interpreto left 😅 By the way, I describe the paper in this thread: 🧵1/9
2141
Reposted by Antonin Poché
Vilém Zouhar @zouhar.bsky.social · 04/09/2026
Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
arxiv.org
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for ...
34612
Antonin Poché @antoninpoche.bsky.social · 19/08/2026
🥱 Am I the only one being bored of MI paper studing probing/steering/circuits for a single concept? I feel like you will always be able to find a probe or a circuit of some kind which works better than random.
010
Reposted by Antonin Poché
Nils Feldhus @nfel.bsky.social · 17/08/2026
Looking for emergency reviewers for BlackboxNLP 2026. Please sign up through the form linked in the post below if you have time to review 1-2 papers before Aug 20 AoE. bsky.app/profile/blac...
025
Antonin Poché @antoninpoche.bsky.social · 11/08/2026
Better late than never, I want to share that 1 month ago we released Interpreto 0.5.0 🪄 🔥It basically improves the whole library and adds new features. This version was used in our FAccT tutorial 🗣️ and ACL Demo 💻 Let's detail things a bit 🧵
100
Reposted by Antonin Poché
BlackboxNLP @blackboxnlp.bsky.social · 21/07/2026
With a high volume of submissions this year, we're recruiting additional reviewers for BlackboxNLP 2026! 🗓 Review deadline: August 17 (AoE) 🗓 Review load: 2-4 papers 📝 Sign up: forms.gle/tha166UQZYqX...
065
Reposted by Antonin Poché
Vilém Zouhar @zouhar.bsky.social · 17/07/2026
There are many things machine translation still can't do. Help us steer the next direction by contributing hard-to-translate inputs (and be on a cool paper).
2148
Antonin Poché @antoninpoche.bsky.social · 06/07/2026
Tomorrow (Monday, 6th), I am presenting Interpreto at ACL San Diego! 🛬 Interpreto is an open-source library for interpreting language models (attributions, probes, SAEs...). I will be in Demo Session E: Grand Hall from 11 am to 1 pm. 📺Look for the TVs, and you will find me!
120
Reposted by Antonin Poché
Cas (Stephen Casper) @scasper.bsky.social · 02/07/2026
docs.google.com/forms/d/e/1...
docs.google.com
Guess which AI companies tend to make models that exhibit corporate loyalty
We (Lennart and Cas) started wondering a few months ago whether popular frontier AI models tend to exhibit corporate loyalties. In a pre-registered experiment, we elicited open-ended discussions from 21 models from 7 companies on 206 negative news stories using 25 prompt templates to assess how favorably each model discusses controversies from each company. The results: ---> Strong evidence (p<10^-5) that models from 4 companies tend to differentially downplay company controversies. ---> No evidence for models from 3 companies. Your task: Guess which
142
Antonin Poché @antoninpoche.bsky.social · 27/06/2026
Yesterday, with @fannyjrd.bsky.social we were delighted to share how explainability can be used for fairness in a @facct.bsky.social tutorial!! We loved the questions and interactions! Check Fanny's thread for the link. The conference is really fun thus far! Thanks to the organizers!
050
Antonin Poché @antoninpoche.bsky.social · 19/06/2026
I am currently in Quebec city, I will be in Montreal next week for FAccT and the week after in San Diego for ACL. Don't hesitate to reach out to meet and talk about interpretability & co 😀
130
Reposted by Antonin Poché
Zeerak Talat زیرک طلعت (they/them) @zeerak.bsky.social · 12/06/2026
NLP reviews suck! But why is not clear– @aclrollingreview.bsky.social provides *a lot* of guidance for how to review, but in that, first principles get lost. So @adamlopez.bsky.social and I have written down some of our thoughts on first principles.
medium.com
The Missing First Principles of Reviewing for ACL
Zeerak Talat & Adam Lopez, University of Edinburgh
1236
Antonin Poché @antoninpoche.bsky.social · 23/03/2026
I forgot to say that I am currently in Berlin till the end of the month. 🇩🇪 If you want to have a beer/tea, please send a message!🍻 I am visiting the XplaiNLP group (@nfel.bsky.social...) at @tuberlin.bsky.social 🤗 Next Thursday, I will talk about Concept-based explanations at @dfki.bsky.social
020
Reposted by Antonin Poché
Nils Feldhus @nfel.bsky.social · 19/03/2026
Can Persona Prompting function as a lens on social reasoning? In our #EACL2026 work (led by @jingyng.bsky.social), we investigate how it impacts the quality of model outputs and rationales. 🗞️ arXiv: arxiv.org/abs/2601.20757 Come and find us (Jing, Moritz, Elisabeth, myself) in 🇲🇦 Rabat next week!
1102
Reposted by Antonin Poché
BlackboxNLP @blackboxnlp.bsky.social · 18/03/2026
BlackboxNLP will be co-located with EMNLP 2026 in 🇭🇺 Budapest 🇭🇺 this October! This edition will feature a special reproducibility track, investigating generalization and robustness of established results from interpretability research 👷‍♂️ Stay tuned for more details!
1167
Antonin Poché @antoninpoche.bsky.social · 04/03/2026
🔥Super excited to share our new demo website for 🪄Interpreto! 🖼️It is basically an explanation gallery showcasing attribution and concept-based explanations for classification and generation. 🎮Play with it: for-sight-ai.github.io/interpreto-d... We will keep improving it, so stay tuned!
193
Antonin Poché @antoninpoche.bsky.social · 23/01/2026
Pleasently surprised to see our blog post trending on HuggingFace 🤗 Well, @fannyjrd.bsky.social did a great job! 🚀 If you missed it, check it out: huggingface.co/blog/Fannyjr... It's a didactic presentation of our new library: 🪄 Interpreto: github.com/FOR-sight-ai...
131
Reposted by Antonin Poché
Gabriele Sarti @gsarti.com · 21/01/2026
It was an honor to be part of this awesome project! Interpreto is a great up-and-coming tool for concept-based interpretability analyses of NLP models, check it out!
081
Reposted by Antonin Poché
Fanny Jourdan @fannyjrd.bsky.social · 20/01/2026
🎉 I’m thrilled to announce the release of Interpreto: a user-friendly, open-source toolbox to make NLP model interpretability accessible, practical, and rigorous. github.com/FOR-sight-ai... 🧵1/5
github.com
GitHub - FOR-sight-ai/interpreto: 🪄 Interpreto is an interpretability toolbox for LLMs
🪄 Interpreto is an interpretability toolbox for LLMs - FOR-sight-ai/interpreto
171
Antonin Poché @antoninpoche.bsky.social · 20/01/2026
🔥I am super excited for the official release of an open-source library we've been working on for about a year! 🪄interpreto is an interpretability toolbox for HF language models🤗. In both generation and classification! Why do you need it, and for what? 1/8 (links at the end)
1229
Reposted by Antonin Poché
🎃 🕯️ Monica Valentinelli is Sipping Tea. 🐈‍⬛ ☕ @booksofm.com · 20/11/2025
If you use GMail, AI (Gemini) was turned on yesterday by default and now scans all of your content for machine learning. To turn off, go to Settings>General and scroll down. Uncheck the box for "Smart features." There's other "Smart" add-ons as well, but that's the one that reads your content.
319106777932
Reposted by Antonin Poché
Thomas Fel @thomasfel.bsky.social · 14/10/2025
🕳️🐇 𝙄𝙣𝙩𝙤 𝙩𝙝𝙚 𝙍𝙖𝙗𝙗𝙞𝙩 𝙃𝙪𝙡𝙡 – 𝙋𝙖𝙧𝙩 𝙄 (𝑃𝑎𝑟𝑡 𝐼𝐼 𝑡𝑜𝑚𝑜𝑟𝑟𝑜𝑤) 𝗔𝗻 𝗶𝗻𝘁𝗲𝗿𝗽𝗿𝗲𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗱𝗲𝗲𝗽 𝗱𝗶𝘃𝗲 𝗶𝗻𝘁𝗼 𝗗𝗜𝗡𝗢𝘃𝟮, one of vision’s most important foundation models. And today is Part I, buckle up, we're exploring some of its most charming features. :)
23612
Reposted by Antonin Poché
Mel Andrews @bayesianboy.bsky.social · 05/10/2025
expressing appreciation for this scientific diagram
3507
Antonin Poché @antoninpoche.bsky.social · 25/07/2025
🔥 I am super excited to be presenting a poster at #ACL2025 in Vienna next week! 🌏 This is my first big conference! 📅 Tuesday morning, 10:30–12:00, during Poster Session 2. 💬 If you're around, feel free to message me. I would be happy to connect, chat, or have a drink!
151
Reposted by Antonin Poché
Naomi Saphra @nsaphra.bsky.social · 10/07/2025
🚨 New preprint! 🚨 Everyone loves causal interp. It’s coherently defined! It makes testable predictions about mechanistic interventions! But what if we had a different objective: predicting model behavior not under mechanistic interventions, but on unseen input data?
36312
Antonin Poché @antoninpoche.bsky.social · 16/05/2025
🔥ConSim has been accepted to the #ACL2025 main conference! 🙏 Thanks again to my amazing co-authors: @alon_jacovi, Agustin Picard, @VictorBoutin, and @Fannyjrd_. Work done in DEEL and FOR from IRT St Exupéry and @ANITI_Toulouse. See you in Vienna 📅 For more information, check out my last post:
141
Reposted by Antonin Poché
Gabriele Sarti @gsarti.com · 15/05/2025
BlackboxNLP is back! 💥 Happy to be part of the organizing team for this year, and super excited for our new shared task using the excellent MIB Benchmark, check it out! blackboxnlp.github.io/2025/task/
062
Reposted by Antonin Poché
Mor Geva @megamor2.bsky.social · 31/03/2025
🎉 Our Actionable Interpretability workshop has been accepted to #ICML2025! 🎉 > Follow @actinterp.bsky.social > Website actionable-interpretability.github.io @talhaklay.bsky.social @anja.re @mariusmosbach.bsky.social @sarah-nlp.bsky.social @iftenney.bsky.social Paper submission deadline: May 9th!
34116
Reposted by Antonin Poché
Nathan Kalman-Lamb @nkalamb.bsky.social · 29/03/2025
Hundreds of international students have just received an email telling them their visas have been revoked. The ‘justification’ is campus activism or social media posts. timesofindia.indiatimes.com/world/us/hun...
The biggest reason government officials aren't giving any specifics about the criteria by which these arrests and deportations are selected, is that the criteria is "pro-Israel think-tanks and advocacy organizations created lists of troublesome individuals and gave them to us."
20756233060
Reposted by Antonin Poché
Chris Olah @colah.bsky.social · 27/03/2025
Can we understand the mechanisms of a frontier AI model? 📝 Blog post: www.anthropic.com/research/tra... 🧪 "Biology" paper: transformer-circuits.pub/2025/attribu... ⚙️ Methods paper: transformer-circuits.pub/2025/attribu... Featuring basic multi-step reasoning, planning, introspection and more!
transformer-circuits.pub
On the Biology of a Large Language Model
412528
Reposted by Antonin Poché
Seth Abramson @sethabramson.bsky.social · 20/03/2025
Jawdropping. You would expect this in a dictatorship, not the United States. This country is unrecognizable.
1388186157653
Reposted by Antonin Poché
David Bau @davidbau.bsky.social · 16/03/2025
What will be the linchpin for AI dominance? Read our NSF/OSTP recommendations written with Goodfire's Tom McGrath tommcgrath.github.io, Transluce's Sarah Schwettmann cogconfluence.com, MIT's Dylan Hadfield-Menell @dhadfieldmenell.bsky.social TLDR; Dominance comes from **interpretability** 🧵 ↘️
1218
Reposted by Antonin Poché
Tom Aarsen @tomaarsen.com · 10/03/2025
An assembly of 18 European companies, labs, and universities have banded together to launch 🇪🇺 EuroBERT! It's a state-of-the-art multilingual encoder for 15 European languages, designed to be finetuned for retrieval, classification, etc. Details in 🧵
57820
Antonin Poché @antoninpoche.bsky.social · 25/02/2025
Super excited to welcome @gsarti.com in Toulouse with @fannyjrd.bsky.social and Thomas Mullor ! We will be working on a new library for interpretability 😀
030
Antonin Poché @antoninpoche.bsky.social · 31/01/2025
🚀 Thrilled to share our new paper (the first of my PhD)! How can we compare concept-based #XAI methods in #NLProc? ConSim (arxiv.org/abs/2501.05855) provides the answer. Read the thread to find out which method is the most interpretable! 🧵1/7
662
Antonin Poché @antoninpoche.bsky.social · 24/01/2025
🤩 Thrilled to announce that I've started my PhD in #XAI for #NLProc under the supervision of Pr. Nicholas Asher, @philmuller.bsky.social, and @fannyjrd.bsky.social! My project? Improve the transparency of LLMs through interactive explanations and user-tailored explanations. 🚀
052