Sign in

Hadas Orgad

@hadasorgad.bsky.social
61 followers 2 following 26 posts
PostsRepliesMedia
Hadas Orgad @hadasorgad.bsky.social · 15/06/2026
You can submit any relevant work that wasn't already accepted to another venue, and in any case won't be presented before the workshop day (October 9th). COLM accepted papers will go through a fast track (details in the link) actionable-interpretability.github.io/cfp/
actionable-interpretability.github.io
Call for Papers
Deadline: June 24 2026 AOE
000
Hadas Orgad @hadasorgad.bsky.social · 15/06/2026
We’re extending the Actionable Interpretability workshop @actinterp.bsky.social submission deadline by 3 days! New deadline: June 24th. Looking forward to your submissions ;) Link in thread
101
Hadas Orgad @hadasorgad.bsky.social · 10/06/2026
📢 We’re looking for reviewers for the Actionable Interpretability workshop @ActInterp ! If you’re interested in helping review submitted papers, please sign up here: forms.gle/7pihaQuSQ2Wq... Your expertise would be greatly appreciated!
forms.gle
Reviewer Form - Actionable Interpretability Workshop
This form collects information on reviewers for the workshop Actionable Interpretability @ COLM 2026. Please take note of the details for the review process: Important Dates: Review Start: June 25...
000
Hadas Orgad @hadasorgad.bsky.social · 10/06/2026
📢 We’re looking for reviewers for the Actionable Interpretability workshop @actinterp.bsky.social! If you’re interested in helping review submitted papers, please sign up here: forms.gle/VpLJpkM6zw3V... Your expertise would be greatly appreciated!
043
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
Wonderful collaboration with @boyiwei.bsky.social Kaden Zheng @wattenberg.bsky.social @peterhenderson.bsky.social Seraphina Goldfarb-Tarrant and @boknilev.bsky.social
010
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
Curious? Read more in our new preprint: arxiv.org/abs/2604.09544
arxiv.org
Large Language Models Generate Harmful Content Using a Distinct,...
Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains...
110
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
We use weight pruning here as a causal probe of model internals, not a deployment-ready defense. But — this opens a path toward *mechanistic alignment*: considering the mechanisms behind harmful behavior, rather than training behavioral guardrails on top of them.
110
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
And it's specifically the generative mechanism we remove: fine-tuning on harmful examples can mostly restore it—confirming our point that the underlying knowledge is largely intact.
100
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
Generating harmful content is dissociated from "understanding" it. Pruned models retain nearly full ability to detect, explain, and refuse harmful requests. Refusal and generation are *double dissociated*: pruning one leaves the other intact.
110
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
And here as well, it generalizes between domains (plus, we observe significant intersection of weight sets).
110
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
This also explains emergent misalignment ( @BetleyJan et al.) If harmful behaviors share weights, then fine-tuning one narrow domain can unintentionally affect others. → Pruning a narrow misaligned domain substantially reduces emergent misalignment.
130
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
This compression seems to be a product of alignment training. Aligned models show far greater separation between harmful and benign weights than their pretrained counterparts. Alignment reshapes the internals, even when behavioral guardrails remain brittle. [more in the paper]
130
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
The mechanism is unified across harm types. E.g., prune weights identified from malware generation → hate speech drops too. Different harms rely on a shared underlying mechanism. These weights also heavily overlap.
120
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
Paper >> arxiv.org/abs/2604.09544 We use weight pruning to probe model internals. Result: pruning ~0.0005% of model parameters, harmful generation drops dramatically, while general capabilities remain largely intact.
110
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
New paper: LLMs encode harmful content generation in a distinct, unified mechanism Using weight pruning, we find that harmful generation depends on a tiny subset of the weights that are shared across harm types and separate from benign capabilities. 🧵
172
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Full paper >> actionable-interpretability-guide.github.io/paper.pdf Blog >> actionable-interpretability-guide.github.io
actionable-interpretability-guide.github.io
000
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Joint work w/ amazing collaborators @fbarez.bsky.social @talhaklay.bsky.social @wordscompute.bsky.social @mariusmosbach.bsky.social @anja.re @nsaphra.bsky.social @byron.bsky.social @sarah-nlp.bsky.social @profericwong.bsky.social @iftenney.bsky.social @megamor2.bsky.social
actionable-interpretability-guide.github.io
100
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
We’re not saying all interpretability work must be immediately actionable— curiosity-driven research still matters. But actionability is a high bar: understanding that works outside the lab. To make your next project more actionable, use our checklist >>
120
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Actionable interpretability is worth aiming for. We identified five domains where answering *why* unlocks a fundamental advantage.
100
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Interpretability isn't actionable (yet) for three reasons: → Papers aren't expected to demonstrate applications → Insights are shown in oversimplified settings without real baselines → Methods require domain expertise
100
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Why haven't insights from interpretability transformed AI yet? Because we're not prioritizing actionable insights. Full paper >> actionable-interpretability-guide.github.io/paper.pdf Blog - The Hitchhiker's Guide 🧭 to Actionable Interpretability >> actionable-interpretability-guide.github.io/
130
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Our ICML 2025 workshop on Actionable Interpretability drew massive interest. But the same questions kept coming up: What does "actionable" mean? Is it achievable? How? We're ready to answer. 🧵
12410
Hadas Orgad @hadasorgad.bsky.social · 03/05/2025
Deadline extended! ⏳ The Actionable Interpretability Workshop at #ICML2025 has moved its submission deadline to May 19th. More time to submit your work 🔍🧠✨ Don’t miss out!
043
Reposted by Hadas Orgad
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!
Logo for MIB: A Mechanistic Interpretability Benchmark
15115
Hadas Orgad @hadasorgad.bsky.social · 31/03/2025
• Model Innovation – Designs and training inspired by interpretability. • Impact Measurement – Benchmarks for real-world effectiveness. • Critical Perspectives – Feasibility, limits, and future directions. Website >>> actionable-interpretability.github.io
actionable-interpretability.github.io
General Information
ICML 2025 - Vancouver
030
Hadas Orgad @hadasorgad.bsky.social · 31/03/2025
• Real-world Applications – Tackling bias, hallucinations, adversarial threats, and use in critical domains like healthcare, finance and cybersecurity. • Method Comparison – Interpretability vs. alternative methods such as fine-tuning, prompting, etc.
120
Hadas Orgad @hadasorgad.bsky.social · 31/03/2025
We aim to foster discussions on how interpretability research can inform concrete improvements in model design, safety, and robustness. Topics of interest: ⬇️
130