Sign in

Yonatan Belinkov

@boknilev.bsky.social
140 followers 392 following 21 posts

Associate professor of computer science at Technion belinkov.com

PostsRepliesMedia
Reposted by Yonatan Belinkov
BlackboxNLP @blackboxnlp.bsky.social · 01/07/2026
🏆 Announcing the NDIF Best Paper Award for the BlackboxNLP 2026 Reproducibility Challenge! @ndif-team.bsky.social Reproduce an interp. finding with nnsight + NDIF, open-source it, and push it further. Winners will receive a $500 prize, and will be invited to present at the workshop!
183
Yonatan Belinkov @boknilev.bsky.social · 17/06/2026
Are you wondering if LLM interpretability results generalize, reproduce, etc.? Check out the reproducibility challenge and submit your work reproducing papers in this area: bsky.app/profile/blac...
1156
Reposted by Yonatan Belinkov
BlackboxNLP @blackboxnlp.bsky.social · 28/05/2026
📣 Announcing the BlackboxNLP 2026 Reproducibility Challenge! A new track dedicated to rigorous robustness checks of NLP interpretability work - stress-testing baselines, ablations, generalizability, and evaluation.
1165
Reposted by Yonatan Belinkov
Martin Tutek @mtutek.bsky.social · 17/03/2026
How does an LLM’s past influence its future?🤔 In new work, led by @adisimhi.bsky.social, together with @fbarez.bsky.social @boknilev.bsky.social and Shay Cohen, we find conversational history creates a latent "geometric trap" which makes old habits e.g. hallucinations hard to break!
2245
Reposted by Yonatan Belinkov
BlackboxNLP @blackboxnlp.bsky.social · 18/03/2026
BlackboxNLP will be co-located with EMNLP 2026 in 🇭🇺 Budapest 🇭🇺 this October! This edition will feature a special reproducibility track, investigating generalization and robustness of established results from interpretability research 👷‍♂️ Stay tuned for more details!
1167
Reposted by Yonatan Belinkov
Martin Tutek @mtutek.bsky.social · 08/10/2025
🤔What happens when LLM agents choose between achieving their goals and avoiding harm to humans in realistic management scenarios? Are LLMs pragmatic or prefer to avoid human harm? 🚀 New paper out: ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs🚀🧵
182
Yonatan Belinkov @boknilev.bsky.social · 07/10/2025
Traveling to #COLM2025 this week, and here's some work from our group and collaborators: Cognitive biases, hidden knowledge, CoT faithfulness, model editing, and LM4Science See the thread for details and reach out if you'd like to discuss more!
161
Reposted by Yonatan Belinkov
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
What's the right unit of analysis for understanding LLM internals? We explore in our mech interp survey (a major update from our 2024 ms). We’ve added more recent work and more immediately actionable directions for future work. Now published in Computational Linguistics!
24115
Yonatan Belinkov @boknilev.bsky.social · 01/10/2025
Opportunities to join my group in fall 2026: * PhD applications direct or via ELLIS @ellis.eu (ellis.eu/news/ellis-p...) * Post-doc applications direct or via Azrieli (azrielifoundation.org/fellows/inte...) or Zuckerman (zuckermanstem.org/ourprograms/...)
031
Yonatan Belinkov @boknilev.bsky.social · 22/09/2025
Excited to join @KempnerInst this year! Get in touch if you're in the Boston area and want to chat about anything related to AI interpretability, robustness, interventions, safety, multi-modality, protein/DNA LMs, new architectures, multi-agent communication, or anything else you're excited about!
020
Yonatan Belinkov @boknilev.bsky.social · 18/09/2025
@robinjia.bsky.social speaking at @kempnerinstitute.bsky.social on Auditing, Dissecting, and Evaluating LLMs
011
Reposted by Yonatan Belinkov
Martin Tutek @mtutek.bsky.social · 21/08/2025
Thrilled that FUR was accepted to @emnlpmeeting.bsky.social Main🎉 In case you can’t wait so long to hear about it in person, it will also be presented as an oral at @interplay-workshop.bsky.social @colmweb.org 🥳 FUR is a parametric test assessing whether CoTs faithfully verbalize latent reasoning.
1133
Yonatan Belinkov @boknilev.bsky.social · 12/08/2025
BlackboxNLP is the workshop on interpreting and analyzing NLP models (including LLMs, VLMs, etc). We accept full (archival) papers and extended abstracts. The workshop is highly attended and is a great exposure for your finished work or feedback on work in progress. #emnlp2025 at Sujhou, China!
051
Yonatan Belinkov @boknilev.bsky.social · 13/07/2025
Join our Discord for discussions and a bunch of simple submission ideas you can try! discord.gg/n5uwjQcxPR Participants will have the option to write a system description paper that gets published.
discord.gg
Join the BlackboxNLP Shared Task Discord Server!
Check out the BlackboxNLP Shared Task community on Discord - hang out with 83 other members and enjoy free voice and text chat.
020
Reposted by Yonatan Belinkov
BlackboxNLP @blackboxnlp.bsky.social · 09/07/2025
Have you started working on your submission for the MIB shared task yet? Tell us what you’re exploring! New featurization methods? Circuit pruning? Better feature attribution? We'd love to hear about it 👇
021
Reposted by Yonatan Belinkov
BlackboxNLP @blackboxnlp.bsky.social · 08/07/2025
Working on feature attribution, circuit discovery, feature alignment, or sparse coding? Consider submitting your work to the MIB Shared Task, part of this year’s #BlackboxNLP We welcome submissions of both existing methods and new or experimental POCs!
153
Reposted by Yonatan Belinkov
Dana Arad @danaarad.bsky.social · 26/06/2025
VLMs perform better on questions about text than when answering the same questions about images - but why? and how can we fix it? In a new project led by Yaniv (@YNikankin on the other app), we investigate this gap from an mechanistic perspective, and use our findings to close a third of it! 🧵
164
Reposted by Yonatan Belinkov
Dana Arad @danaarad.bsky.social · 27/05/2025
Tried steering with SAEs and found that not all features behave as expected? Check out our new preprint - "SAEs Are Good for Steering - If You Select the Right Features" 🧵
2186
Reposted by Yonatan Belinkov
tomerashuach.bsky.social @tomerashuach.bsky.social · 27/05/2025
🚨New paper at #ACL2025 Findings! REVS: Unlearning Sensitive Information in LMs via Rank Editing in the Vocabulary Space. LMs memorize and leak sensitive data—emails, SSNs, URLs from their training. We propose a surgical method to unlearn it. 🧵👇w/ @boknilev.bsky.social @mtutek.bsky.social 1/8
162
Yonatan Belinkov @boknilev.bsky.social · 15/05/2025
Interested in mechanistic interpretability and care about evaluation? Please consider submitting to our shared task at #blackboxNLP this year!
041
Reposted by Yonatan Belinkov
Ana Marasović @anamarasovic.bsky.social · 04/05/2025
Slides available here: docs.google.com/presentation...
docs.google.com
Repl4NLP Keynote – Ana Marasovic
If You Want Reasoning, Look Inside Ana Marasović Repl4NLP 2025
1255
Yonatan Belinkov @boknilev.bsky.social · 24/04/2025
Excited about the release of MIB, a mechanistic Interpretability benchmark! Come talk to us at #iclr2025 and consider submitting to the leaderboard. We’re also planning a shared task around it at #blackboxNLP this year, located with #emnlp2025
140
Reposted by Yonatan Belinkov
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
We release many public resources, including: 🌐 Website: mib-bench.github.io 📄 Data: huggingface.co/collections/... 💻 Code: github.com/aaronmueller... 📊 Leaderboard: Coming very soon!
mib-bench.github.io
MIB – Project Page
131
Reposted by Yonatan Belinkov
Tal Haklay @talhaklay.bsky.social · 06/03/2025
1/13 LLM circuits tell us where the computation happens inside the model—but the computation varies by token position, a key detail often ignored! We propose a method to automatically find position-aware circuits, improving faithfulness while keeping circuits compact. 🧵👇
1268
Reposted by Yonatan Belinkov
Martin Tutek @mtutek.bsky.social · 21/02/2025
If erasing information from CoT steps adversely affects the prediction of the model, indicating that such explanations are parametrically faithful.
141
Reposted by Yonatan Belinkov
Martin Tutek @mtutek.bsky.social · 21/02/2025
It has been amazing to work with @fatemehc.bsky.social, @anamarasovic.bsky.social and Yonatan Belinkov on this incredibly important topic. I look forward to further works on the parametric faithfulness route! Codebase (& data): github.com/technion-cs-...
github.com
GitHub - technion-cs-nlp/parametric-faithfulness
Contribute to technion-cs-nlp/parametric-faithfulness development by creating an account on GitHub.
062
Reposted by Yonatan Belinkov
Adi Simhi @adisimhi.bsky.social · 19/02/2025
🚨New arXiv preprint!🚨 LLMs can hallucinate - but did you know they can do so with high certainty even when they know the correct answer? 🤯 We find those hallucinations in our latest work with @itay-itzhak.bsky.social, @fbarez.bsky.social, @gabistanovsky.bsky.social and Yonatan Belinkov
32110
Reposted by Yonatan Belinkov
Michael Hanna @michaelwhanna.bsky.social · 19/12/2024
Sentences are partially understood before they're fully read. How do LMs incrementally interpret their inputs? In a new paper, @amuuueller.bsky.social and I use mech interp tools to study how LMs process structurally ambiguous sentences. We show LMs rely on both syntactic & spurious features! 1/10
17115