Sign in

Anja Reusch

@anja.re
66 followers 117 following 0 posts

👩‍💻 Postdoc @ Technion, interested in Interpretability in IR 🔎 and NLP 💬

PostsRepliesMedia
Reposted by Anja Reusch
LION Lab @lionlab.bsky.social · 17/09/2026
🦁LION Lab is hiring! 🧑‍🔬One fully funded PhD student (TVL-13 100%) for 3 years 💉Topic: Interpretability for Protein Language Models 👥Advised by @weissweiler.bsky.social together with Clara Schoeder 🌍Leipzig, Germany 🔗Apply by Oct 15: lionlabnlp.github.io/jobs/ai4pf/ Please share! #NLProc #NLP
0107
Reposted by Anja Reusch
LION Lab @lionlab.bsky.social · 26/08/2026
Yesterday we had @anja.re from Technion visiting us to give a talk on Interpretability for Generative Information Retrieval. The talk was a cool example to show that interpretability can also be actionable. Thank you very much for joining us, Anja!
052
Reposted by Anja Reusch
Itay Itzhak @ COLM 🍁 @itay-itzhak.bsky.social · 09/07/2026
We vibe-tested our own paper, but the reviewers gave us the actual metrics: "From Feelings to Metrics" has been accepted to #COLM2026! 🎉 We turn "this model just feels better" into structured, user-aware evaluation! #COLM people - let’s grab a coffee and vibe-test some research in person! ☕️🌴
031
Reposted by Anja Reusch
Hadas Orgad @hadasorgad.bsky.social · 10/06/2026
📢 We’re looking for reviewers for the Actionable Interpretability workshop @actinterp.bsky.social! If you’re interested in helping review submitted papers, please sign up here: forms.gle/VpLJpkM6zw3V... Your expertise would be greatly appreciated!
043
Reposted by Anja Reusch
Aaron Mueller @amuuueller.bsky.social · 10/06/2026
The New England Mechanistic Interpretability (NEMI) workshop is coming to BU on Aug. 14! Join us for talks, a panel, food, and plenty of opportunities to connect with the many great researchers in the area. Register and help spread the word!
0176
Reposted by Anja Reusch
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Our ICML 2025 workshop on Actionable Interpretability drew massive interest. But the same questions kept coming up: What does "actionable" mean? Is it achievable? How? We're ready to answer. 🧵
12410
Reposted by Anja Reusch
Martin Tutek @mtutek.bsky.social · 08/10/2025
🤔What happens when LLM agents choose between achieving their goals and avoiding harm to humans in realistic management scenarios? Are LLMs pragmatic or prefer to avoid human harm? 🚀 New paper out: ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs🚀🧵
182
Reposted by Anja Reusch
Aaron Mueller @amuuueller.bsky.social · 17/07/2025
If you're at #ICML2025, chat with me, @sarah-nlp.bsky.social, Atticus, and others at our poster 11am - 1:30pm at East #1205! We're establishing a 𝗠echanistic 𝗜nterpretability 𝗕enchmark. We're planning to keep this a living benchmark; come by and share your ideas/hot takes!
0133
Reposted by Anja Reusch
BlackboxNLP @blackboxnlp.bsky.social · 24/06/2025
Working on circuit discovery in LMs? Consider submitting your work to the MIB Shared Task, part of #BlackboxNLP at @emnlpmeeting.bsky.social 2025! The goal: benchmark existing MI methods and identify promising directions to precisely and concisely recover causal pathways in LMs >>
154
Reposted by Anja Reusch
Actionable Interpretability Workshop ICML2025 @actinterp.bsky.social · 20/05/2025
🚨 We're looking for more reviewers for the workshop! 📆 Review period: May 24-June 7 If you're passionate about making interpretability useful and want to help shape the conversation, we'd love your input. 💡🔍 Self-nominate here: docs.google.com/forms/d/e/1F...
An image with the Vancouver skyline and the words "sign up to review". At the top are the logos of both the Actionable Interpretability workshop (a magnifying glass) and the ICML conference (a brain).
055
Reposted by Anja Reusch
Tal Haklay @talhaklay.bsky.social · 14/05/2025
We knew many of you wanted to submit to our Actionable Interpretability workshop, but we didn’t expect to crash Overleaf! 😏🍃 Only 5 days left ⏰! Got a paper accepted to ICML that fits our theme? Submit it to our conference track! 👉 @actinterp.bsky.social
142
Reposted by Anja Reusch
Hadas Orgad @hadasorgad.bsky.social · 03/05/2025
Deadline extended! ⏳ The Actionable Interpretability Workshop at #ICML2025 has moved its submission deadline to May 19th. More time to submit your work 🔍🧠✨ Don’t miss out!
043
Reposted by Anja Reusch
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!
Logo for MIB: A Mechanistic Interpretability Benchmark
15115
Reposted by Anja Reusch
Tal Haklay @talhaklay.bsky.social · 07/04/2025
🚨 Call for Papers is Out! The First Workshop on 𝐀𝐜𝐭𝐢𝐨𝐧𝐚𝐛𝐥𝐞 𝐈𝐧𝐭𝐞𝐫𝐩𝐫𝐞𝐭𝐚𝐛𝐢𝐥𝐢𝐭𝐲 will be held at ICML 2025 in Vancouver! 📅 Submission Deadline: May 9 Follow us >> @ActInterp 🧠Topics of interest include: 👇
153
Reposted by Anja Reusch
Sarah Wiegreffe @sarah-nlp.bsky.social · 03/04/2025
Have work on the actionable impact of interpretability findings? Consider submitting to our Actionable Interpretability workshop at ICML! See below for more info. Website: actionable-interpretability.github.io Deadline: May 9
02010
Reposted by Anja Reusch
Mor Geva @megamor2.bsky.social · 31/03/2025
🎉 Our Actionable Interpretability workshop has been accepted to #ICML2025! 🎉 > Follow @actinterp.bsky.social > Website actionable-interpretability.github.io @talhaklay.bsky.social @anja.re @mariusmosbach.bsky.social @sarah-nlp.bsky.social @iftenney.bsky.social Paper submission deadline: May 9th!
34116
Reposted by Anja Reusch
Aaron Mueller @amuuueller.bsky.social · 11/03/2025
(ICLR) How do LLMs perform arithmetic operations? Do they implement robust algorithms, or rely on heuristics? We find that they rely on a "bag of heuristics" that work well—but on a limited range of inputs. Led by Yaniv Nikankin: arxiv.org/abs/2410.21272
arxiv.org
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
Do large language models (LLMs) solve reasoning tasks by learning robust generalizable algorithms, or do they memorize training data? To investigate this question, we use arithmetic reasoning as a rep...
141
Reposted by Anja Reusch
Tal Haklay @talhaklay.bsky.social · 06/03/2025
1/13 LLM circuits tell us where the computation happens inside the model—but the computation varies by token position, a key detail often ignored! We propose a method to automatically find position-aware circuits, improving faithfulness while keeping circuits compact. 🧵👇
1268
Reposted by Anja Reusch
Adi Simhi @adisimhi.bsky.social · 19/02/2025
🚨New arXiv preprint!🚨 LLMs can hallucinate - but did you know they can do so with high certainty even when they know the correct answer? 🤯 We find those hallucinations in our latest work with @itay-itzhak.bsky.social, @fbarez.bsky.social, @gabistanovsky.bsky.social and Yonatan Belinkov
32110
Reposted by Anja Reusch
Martin Tutek @mtutek.bsky.social · 21/02/2025
🚨🚨 New preprint 🚨🚨 Ever wonder whether verbalized CoTs correspond to the internal reasoning process of the model? We propose a novel parametric faithfulness approach, which erases information contained in CoT steps from the model parameters to assess CoT faithfulness. arxiv.org/abs/2502.14829
arxiv.org
Measuring Faithfulness of Chains of Thought by Unlearning Reasoning Steps
When prompted to think step-by-step, language models (LMs) produce a chain of thought (CoT), a sequence of reasoning steps that the model supposedly used to produce its prediction. However, despite mu...
24813
Reposted by Anja Reusch
Jeremy Howard @howard.fm · 19/12/2024
I'll get straight to the point. We trained 2 new models. Like BERT, but modern. ModernBERT. Not some hypey GenAI thing, but a proper workhorse model, for retrieval, classification, etc. Real practical stuff. It's much faster, more accurate, longer context, and more useful. 🧵
19620147