Sign in

Alessandro Stolfo

@alestolfo.bsky.social
371 followers 65 following 2 posts

PhD @ ETHZ - LLM Interpretability alestolfo.github.io

PostsRepliesMedia
Reposted by Alessandro Stolfo
Yucheng Sun @yuchengsun.bsky.social · 18/07/2025
1/6: Can we use an LLM’s hidden activations to predict and prevent wrong predictions? When it comes to arithmetic, yes! I’m presenting new work w/ @alestolfo.bsky.social “Probing for Arithmetic Errors in LMs” @ #ICML2025 Act Interp WS 🧵 below
511
Reposted by Alessandro Stolfo
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!
Logo for MIB: A Mechanistic Interpretability Benchmark
15115
Alessandro Stolfo @alestolfo.bsky.social · 15/04/2025
Our paper "Improving Instruction-Following in Language Models through Activation Steering” has been accepted to #ICLR2025! We're also excited to share that our public GitHub repo is now live. Code: github.com/microsoft/ll... Camera-ready: arxiv.org/abs/2410.12877
182