Sign in

Dana Arad

@danaarad.bsky.social
67 followers 231 following 23 posts

NLP Researcher | CS PhD Candidate @ Technion

PostsRepliesMedia
Reposted by Dana Arad
BlackboxNLP @blackboxnlp.bsky.social · 21/07/2026
With a high volume of submissions this year, we're recruiting additional reviewers for BlackboxNLP 2026! 🗓 Review deadline: August 17 (AoE) 🗓 Review load: 2-4 papers 📝 Sign up: forms.gle/tha166UQZYqX...
065
Reposted by Dana Arad
Nils Feldhus @nfel.bsky.social · 17/08/2026
Looking for emergency reviewers for BlackboxNLP 2026. Please sign up through the form linked in the post below if you have time to review 1-2 papers before Aug 20 AoE. bsky.app/profile/blac...
025
Reposted by Dana Arad
BlackboxNLP @blackboxnlp.bsky.social · 18/03/2026
BlackboxNLP will be co-located with EMNLP 2026 in 🇭🇺 Budapest 🇭🇺 this October! This edition will feature a special reproducibility track, investigating generalization and robustness of established results from interpretability research 👷‍♂️ Stay tuned for more details!
1167
Dana Arad @danaarad.bsky.social · 20/08/2025
Now accepted to EMNLP Main Conference!
040
Dana Arad @danaarad.bsky.social · 12/08/2025
Submit your work to #BlackboxNLP 2025!
030
Dana Arad @danaarad.bsky.social · 10/08/2025
Excited to spend the rest of the summer visiting @davidbau.bsky.social's lab at Northeastern! If you’re in the area and want to chat about interpretability, let me know ☕️
091
Reposted by Dana Arad
Itay Itzhak @ COLM 🍁 @itay-itzhak.bsky.social · 27/07/2025
In Vienna for #ACL2025, and already had my first (vegan) Austrian sausage! Now hungry for discussing: – LLMs behavior – Interpretability – Biases & Hallucinations – Why eval is so hard (but so fun) Come say hi if that’s your vibe too!
031
Dana Arad @danaarad.bsky.social · 23/07/2025
10 days to go! Still time to run your method and submit!
011
Dana Arad @danaarad.bsky.social · 13/07/2025
Three weeks is plenty of time to submit your method!
000
Dana Arad @danaarad.bsky.social · 09/07/2025
What are you working on for the MIB shared task? Check out the full task description here: blackboxnlp.github.io/2025/task/
blackboxnlp.github.io
BlackboxNLP 2025
The Eight Workshop on Analyzing and Interpreting Neural Networks for NLP
000
Reposted by Dana Arad
BlackboxNLP @blackboxnlp.bsky.social · 07/07/2025
New to mechanistic interpretability? The MIB shared task is a great opportunity to experiment: ✅ Clean setup ✅ Open baseline code ✅ Standard evaluation Join the discord server for ideas and discussions: discord.gg/n5uwjQcxPR
093
Dana Arad @danaarad.bsky.social · 26/06/2025
VLMs perform better on questions about text than when answering the same questions about images - but why? and how can we fix it? In a new project led by Yaniv (@YNikankin on the other app), we investigate this gap from an mechanistic perspective, and use our findings to close a third of it! 🧵
164
Reposted by Dana Arad
BlackboxNLP @blackboxnlp.bsky.social · 24/06/2025
Working on circuit discovery in LMs? Consider submitting your work to the MIB Shared Task, part of #BlackboxNLP at @emnlpmeeting.bsky.social 2025! The goal: benchmark existing MI methods and identify promising directions to precisely and concisely recover causal pathways in LMs >>
154
Reposted by Dana Arad
BlackboxNLP @blackboxnlp.bsky.social · 23/06/2025
Have you heard about this year's shared task? 📢 Mechanistic Interpretability (MI) is quickly advancing, but comparing methods remains a challenge. This year at #BlackboxNLP, we're introducing a shared task to rigorously evaluate MI methods in language models 🧵
1164
Reposted by Dana Arad
Aaron Mueller @amuuueller.bsky.social · 27/05/2025
SAEs have been found to massively underperform supervised methods for steering neural networks. In new work led by @danaarad.bsky.social, we find that this problem largely disappears if you select the right features!
1161
Dana Arad @danaarad.bsky.social · 27/05/2025
Tried steering with SAEs and found that not all features behave as expected? Check out our new preprint - "SAEs Are Good for Steering - If You Select the Right Features" 🧵
2186
Reposted by Dana Arad
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!
Logo for MIB: A Mechanistic Interpretability Benchmark
15115
Reposted by Dana Arad
Martin Tutek @mtutek.bsky.social · 21/02/2025
🚨🚨 New preprint 🚨🚨 Ever wonder whether verbalized CoTs correspond to the internal reasoning process of the model? We propose a novel parametric faithfulness approach, which erases information contained in CoT steps from the model parameters to assess CoT faithfulness. arxiv.org/abs/2502.14829
arxiv.org
Measuring Faithfulness of Chains of Thought by Unlearning Reasoning Steps
When prompted to think step-by-step, language models (LMs) produce a chain of thought (CoT), a sequence of reasoning steps that the model supposedly used to produce its prediction. However, despite mu...
24813
Reposted by Dana Arad
Adi Simhi @adisimhi.bsky.social · 19/02/2025
🚨New arXiv preprint!🚨 LLMs can hallucinate - but did you know they can do so with high certainty even when they know the correct answer? 🤯 We find those hallucinations in our latest work with @itay-itzhak.bsky.social, @fbarez.bsky.social, @gabistanovsky.bsky.social and Yonatan Belinkov
32110
Reposted by Dana Arad
Sweta Karlekar @swetakar.bsky.social · 19/11/2024
If you’re interested in mechanistic interpretability, I just found this starter pack and wanted to boost it (thanks for creating it @butanium.bsky.social !). Excited to have a mech interp community on bluesky 🎉 go.bsky.app/LisK3CP
3368
Reposted by Dana Arad
Maria Antoniak @mariaa.bsky.social · 04/11/2024
A starter pack for #NLP #NLProc researchers! 🎉 go.bsky.app/SngwGeS
4525199