Hadas Orgad @hadasorgad.bsky.social · 15/06/2026We’re extending the Actionable Interpretability workshop @actinterp.bsky.social submission deadline by 3 days! New deadline: June 24th. Looking forward to your submissions ;) Link in thread 101
Hadas Orgad @hadasorgad.bsky.social · 10/06/2026📢 We’re looking for reviewers for the Actionable Interpretability workshop @ActInterp ! If you’re interested in helping review submitted papers, please sign up here: forms.gle/7pihaQuSQ2Wq... Your expertise would be greatly appreciated!forms.gleReviewer Form - Actionable Interpretability WorkshopThis form collects information on reviewers for the workshop Actionable Interpretability @ COLM 2026. Please take note of the details for the review process: Important Dates: Review Start: June 25... 000
Hadas Orgad @hadasorgad.bsky.social · 10/06/2026📢 We’re looking for reviewers for the Actionable Interpretability workshop @actinterp.bsky.social! If you’re interested in helping review submitted papers, please sign up here: forms.gle/VpLJpkM6zw3V... Your expertise would be greatly appreciated! 043
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026New paper: LLMs encode harmful content generation in a distinct, unified mechanism Using weight pruning, we find that harmful generation depends on a tiny subset of the weights that are shared across harm types and separate from benign capabilities. 🧵 172
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026Our ICML 2025 workshop on Actionable Interpretability drew massive interest. But the same questions kept coming up: What does "actionable" mean? Is it achievable? How? We're ready to answer. 🧵 12410
Hadas Orgad @hadasorgad.bsky.social · 03/05/2025Deadline extended! ⏳ The Actionable Interpretability Workshop at #ICML2025 has moved its submission deadline to May 19th. More time to submit your work 🔍🧠✨ Don’t miss out! 043
Reposted by Hadas OrgadAaron Mueller @amuuueller.bsky.social · 23/04/2025Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark! 15115