Sign in

Tal Haklay

@talhaklay.bsky.social
62 followers 327 following 28 posts

NLP | Interpretability | PhD student at the Technion

PostsRepliesMedia
Tal Haklay @talhaklay.bsky.social · 22/05/2025
Project page >> peap-circuits.github.io Arxiv >> arxiv.org/abs/2502.04577
peap-circuits.github.io
Position-aware Automatic Circuit Discovery – Project Page
010
Tal Haklay @talhaklay.bsky.social · 22/05/2025
Our paper "Position-Aware Automatic Circuit Discovery" got accepted to ACL! 🎉 Huge thanks to my collaborators🙏 @hadasorgad.bsky.social @davidbau.bsky.social @amuuueller.bsky.social @boknilev.bsky.social See you in Vienna! 🇦🇹 #ACL2025 @aclmeeting.bsky.social
1132
Reposted by Tal Haklay
Actionable Interpretability Workshop ICML2025 @actinterp.bsky.social · 20/05/2025
🚨 We're looking for more reviewers for the workshop! 📆 Review period: May 24-June 7 If you're passionate about making interpretability useful and want to help shape the conversation, we'd love your input. 💡🔍 Self-nominate here: docs.google.com/forms/d/e/1F...
An image with the Vancouver skyline and the words "sign up to review". At the top are the logos of both the Actionable Interpretability workshop (a magnifying glass) and the ICML conference (a brain).
055
Tal Haklay @talhaklay.bsky.social · 14/05/2025
Website & CFP >> actionable-interpretability.github.io
actionable-interpretability.github.io
General Information
July 19 - ICML 2025 - Vancouver
010
Tal Haklay @talhaklay.bsky.social · 14/05/2025
We knew many of you wanted to submit to our Actionable Interpretability workshop, but we didn’t expect to crash Overleaf! 😏🍃 Only 5 days left ⏰! Got a paper accepted to ICML that fits our theme? Submit it to our conference track! 👉 @actinterp.bsky.social
142
Reposted by Tal Haklay
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
This was a huge collaboration with many great folks! If you get a chance, be sure to talk to Atticus Geiger, @sarah-nlp.bsky.social, @danaarad.bsky.social, Iván Arcuschin, @adambelfki.bsky.social, @yiksiu.bsky.social, Jaden Fiotto-Kaufmann, @talhaklay.bsky.social, @michaelwhanna.bsky.social, ...
181
Tal Haklay @talhaklay.bsky.social · 07/04/2025
000
Tal Haklay @talhaklay.bsky.social · 07/04/2025
100
Tal Haklay @talhaklay.bsky.social · 07/04/2025
Website >> actionable-interpretability.github.io
actionable-interpretability.github.io
General Information
ICML 2025 - Vancouver
000
Tal Haklay @talhaklay.bsky.social · 07/04/2025
6. Position papers: Critical discussions on the feasibility, limitations, and future directions of actionable interpretability research. We also invite perspectives that question whether actionability should be a goal of interpretability research.
100
Tal Haklay @talhaklay.bsky.social · 07/04/2025
5. Developing realistic benchmarking and assessment methods to measure the real-world impact of interpretability insights, particularly in production environments and large-scale models.
100
Tal Haklay @talhaklay.bsky.social · 07/04/2025
4. Incorporating interpretability–often focusing on micro-level decision analysis–into more complex scenarios, like reasoning processes or multi-turn interactions.
100
Tal Haklay @talhaklay.bsky.social · 07/04/2025
3. New model architectures, training paradigms or design choices informed by interpretability findings.
100
Tal Haklay @talhaklay.bsky.social · 07/04/2025
2. Comparative analyses of interpretability-based approaches versus alternative techniques like fine-tuning, prompting, and more.
100
Tal Haklay @talhaklay.bsky.social · 07/04/2025
1.Practical applications of interpretability insights to address key challenges in AI such as hallucinations, biases, and adversarial robustness, as well as applications in high-stakes, less-explored domains like healthcare, finance, and cybersecurity.
110
Tal Haklay @talhaklay.bsky.social · 07/04/2025
🚨 Call for Papers is Out! The First Workshop on 𝐀𝐜𝐭𝐢𝐨𝐧𝐚𝐛𝐥𝐞 𝐈𝐧𝐭𝐞𝐫𝐩𝐫𝐞𝐭𝐚𝐛𝐢𝐥𝐢𝐭𝐲 will be held at ICML 2025 in Vancouver! 📅 Submission Deadline: May 9 Follow us >> @ActInterp 🧠Topics of interest include: 👇
153
Tal Haklay @talhaklay.bsky.social · 31/03/2025
Amazing news: our workshop was accepted to ICML 2025! Interpretability research sheds light on how models work—but too often, those insights don’t translate into actions that improve them. Our workshop aims to challenge the interpretability community to go further.
020
Tal Haklay @talhaklay.bsky.social · 06/03/2025
13/13 This work was done in collaboration with @hadasorgad.bsky.social , @davidbau.bsky.social , @amuuueller.bsky.social and @boknilev.bsky.social. 💡 Thoughts? Questions? Let’s discuss! Website >> peap-circuits.github.io Arxiv >> arxiv.org/abs/2502.04577
arxiv.org
Position-aware Automatic Circuit Discovery
A widely used strategy to discover and understand language model mechanisms is circuit analysis. A circuit is a minimal subgraph of a model's computation graph that executes a specific task. We identi...
010
Tal Haklay @talhaklay.bsky.social · 06/03/2025
12/13 We evaluate our automatic pipeline across three datasets and two models, demonstrating that: 1️⃣ Our pipeline discovers circuits with a better tradeoff between size and faithfulness compared to EAP. 2️⃣ Our pipeline produces results comparable to those obtained when human experts define a schema.
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
11/13 But where does this schema come from? And how do we determine the boundaries of each span within each example? Sounds like we just added more work for researchers! 😅 Actually, we show that an LLM (Claude) can do a pretty decent job at defining a schema and tagging all examples accordingly.
120
Tal Haklay @talhaklay.bsky.social · 06/03/2025
10/13 After defining a schema, we construct an abstract computation graph where each span type corresponds to a single token position. We then map attribution scores from example-specific computation graphs to the abstract graph and identify circuits within it.
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
9/13 To address this problem, we introduce the concept of a 𝙙𝙖𝙩𝙖𝙨𝙚𝙩 𝙨𝙘𝙝𝙚𝙢𝙖, which defines token spans with similar semantics across examples in the dataset.
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
8/13 But you may notice an issue... What if the examples in a dataset vary in length and structure? Discovering a circuit in such cases is not straightforward, leading many researchers to focus only on datasets with uniform length and structure.
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
7/13 First improvement : We introduce 𝗣𝗼𝘀𝗶𝘁𝗶𝗼𝗻𝗮𝗹 𝗘𝗱𝗴𝗲 𝗔𝘁𝘁𝗿𝗶𝗯𝘂𝘁𝗶𝗼𝗻 𝗣𝗮𝘁𝗰𝗵𝗶𝗻𝗴 (𝗣𝗘𝗔𝗣) —an extension of EAP that allows us to discover circuits that differentiate between token positions. The key advancement? Our approach uncovers "attention edges", revealing dependencies missed by previous methods.
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
6/13 The Problem: Automatic circuit discovery methods like Edge Attribution Patching (EAP) and EAP-IP implicitly assume that circuits are position-invariant—they do not differentiate between components at different token positions. As a result, the circuit may include irrelevant components.
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
5/13 Since the IOI circuit was first discovered, many new techniques for discovering circuits have emerged, with a clear trend of being automated and efficient. Automated methods offer the advantage of scaling more easily and being less susceptible to human biases.
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
4/13 Early circuit discovery techniques relied on manual causal analysis to identify circuits. Here’s an example of a well-studied circuit in the IOI task by Wang et al. Notice how different components play crucial roles at different token positions—this is expected!
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
3/13 What is a circuit? A circuit is a minimal subgraph of a model’s computation graph that executes a specific task. Circuit analysis helps us understand how the model operates and which components (e.g., MLPs, attention heads) are involved.
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
2/13 Check out the full paper >> arxiv.org/abs/2502.04577 Website >> peap-circuits.github.io Or continue in this thread for paper highlights! 🧵👇
arxiv.org
Position-aware Automatic Circuit Discovery
A widely used strategy to discover and understand language model mechanisms is circuit analysis. A circuit is a minimal subgraph of a model's computation graph that executes a specific task. We identi...
110
Tal Haklay @talhaklay.bsky.social · 06/03/2025
1/13 LLM circuits tell us where the computation happens inside the model—but the computation varies by token position, a key detail often ignored! We propose a method to automatically find position-aware circuits, improving faithfulness while keeping circuits compact. 🧵👇
1268
Reposted by Tal Haklay
Martin Tutek @mtutek.bsky.social · 21/02/2025
🚨🚨 New preprint 🚨🚨 Ever wonder whether verbalized CoTs correspond to the internal reasoning process of the model? We propose a novel parametric faithfulness approach, which erases information contained in CoT steps from the model parameters to assess CoT faithfulness. arxiv.org/abs/2502.14829
arxiv.org
Measuring Faithfulness of Chains of Thought by Unlearning Reasoning Steps
When prompted to think step-by-step, language models (LMs) produce a chain of thought (CoT), a sequence of reasoning steps that the model supposedly used to produce its prediction. However, despite mu...
24813