Sign in

Explainable AI

@xai-research.bsky.social
183 followers 118 following 3 posts

Explainable/Interpretable AI researchers and enthusiasts - DM to join the XAI Slack! Blue Sky and Slack maintained by Nick Kroeger

PostsRepliesMedia
Reposted by Explainable AI
Marianne de Heer Kloots @mdhk.net · 19/08/2025
Had such a great time presenting our tutorial on Interpretability Techniques for Speech Models at #Interspeech2025! 🔍 For anyone looking for an introduction to the topic, we've now uploaded all materials to the website: interpretingdl.github.io/speech-inter...
interpretingdl.github.io
Interpretability Techniques for Speech Models — Tutorial @ Interspeech 2025
24015
Reposted by Explainable AI
Antonin Poché @antoninpoche.bsky.social · 25/07/2025
🔥 I am super excited to be presenting a poster at #ACL2025 in Vienna next week! 🌏 This is my first big conference! 📅 Tuesday morning, 10:30–12:00, during Poster Session 2. 💬 If you're around, feel free to message me. I would be happy to connect, chat, or have a drink!
151
Reposted by Explainable AI
Naomi Saphra @nsaphra.bsky.social · 12/06/2025
ACL paper alert! What structure is lost when using linearizing interp methods like Shapley? We show the nonlinear interactions between features reflect structures described by the sciences of syntax, semantics, and phonology.
35412
Reposted by Explainable AI
Sweta Mahajan @swetamahajan.bsky.social · 18/07/2025
🚨Deadline Extension Alert! Our Non-proceedings track is open till August 15th for the eXCV workshop at ICCV. Our nectar track accepts published papers, as is. More info at: excv-workshop.github.io @iccv.bsky.social #ICCV2025
155
Reposted by Explainable AI
Dana Arad @danaarad.bsky.social · 23/07/2025
10 days to go! Still time to run your method and submit!
011
Reposted by Explainable AI
Jennifer Hu @jennhu.bsky.social · 16/07/2025
Excited to announce the first workshop on CogInterp: Interpreting Cognition in Deep Learning Models @ NeurIPS 2025! 📣 How can we interpret the algorithms and representations underlying complex behavior in deep learning models? 🌐 coginterp.github.io/neurips2025/ 1/4
coginterp.github.io
Home
First Workshop on Interpreting Cognition in Deep Learning Models (NeurIPS 2025)
15819
Reposted by Explainable AI
Aaron Mueller @amuuueller.bsky.social · 17/07/2025
If you're at #ICML2025, chat with me, @sarah-nlp.bsky.social, Atticus, and others at our poster 11am - 1:30pm at East #1205! We're establishing a 𝗠echanistic 𝗜nterpretability 𝗕enchmark. We're planning to keep this a living benchmark; come by and share your ideas/hot takes!
0133
Reposted by Explainable AI
Katharina Prasse @katharinaprasse.bsky.social · 17/07/2025
Poster is up and we are looking forward to the #ICML2025 poster session. Come join @patrickknab.bsky.social and me at Poster #W-214 presenting our work with @smarton.bsky.social, Christian Bartelt, and @margretkeuper.bsky.social @margretkeuper.bsky.social #UniMa
012
Reposted by Explainable AI
Sarah Wiegreffe @sarah-nlp.bsky.social · 17/07/2025
I am at #ICML2025! 🇨🇦🏞️ Catch me: 1️⃣ Presenting this paper👇 tomorrow 11am-1:30pm at East #1205 2️⃣ At the Actionable Interpretability @actinterp.bsky.social workshop on Saturday in East Ballroom A (I’m an organizer!)
131
Reposted by Explainable AI
Julian Minder @jkminder.bsky.social · 17/07/2025
Causal Abstraction, the theory behind DAS, tests if a network realizes a given algorithm. We show (w/ @denissutter.bsky.social, T. Hofmann, @tpimentel.bsky.social ) that the theory collapses without the linear representation hypothesis—a problem we call the non-linear representation dilemma.
152
Reposted by Explainable AI
Katharina Prasse @katharinaprasse.bsky.social · 15/07/2025
Join us on Thursday 11-13 in poster hall West #214 to discuss image segments as concepts. #ICML2025 @patrickknab.bsky.social @smarton.bsky.social Christian bartelt @margretkeuper.bsky.social @keuper-labs.bsky.social
022
Reposted by Explainable AI
Harry Thasarathan @hthasarathan.bsky.social · 15/07/2025
🌌🛰️🔭Want to explore universal visual features? Check out our interactive demo of concepts learned from our #ICML2025 paper "Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment". Come see our poster at 4pm on Tuesday in East Exhibition hall A-B, E-1208!
1126
Reposted by Explainable AI
Computer Vision and Machine Learning at MPI Informatics @cvml.mpi-inf.mpg.de · 13/07/2025
Papers being presented from our group at #ICML2025! Congratulations to all the authors! To know more, visit us in the poster sessions! A 🧵with more details: @icmlconf.bsky.social @mpi-inf.mpg.de
Papers accepted at ICML 2025 from the Computer Vision and Machine Learning Department at the Max Planck Institute for Informatics.
1215
Reposted by Explainable AI
Ulrike Luxburg @ulrikeluxburg.bsky.social · 11/07/2025
Our #ICML position paper: #XAI is similar to applied statistics: it uses summary statistics in an attempt to answer real world questions. But authors need to state how concretely (!) their XAI statistics contributes to answer which concrete (!) question! arxiv.org/abs/2402.02870
062
Reposted by Explainable AI
Naomi Saphra @nsaphra.bsky.social · 10/07/2025
🚨 New preprint! 🚨 Everyone loves causal interp. It’s coherently defined! It makes testable predictions about mechanistic interventions! But what if we had a different objective: predicting model behavior not under mechanistic interventions, but on unseen input data?
36312
Reposted by Explainable AI
Sweta Mahajan @swetamahajan.bsky.social · 10/07/2025
Introducing the speakers for the eXCV workshop at ICCV, Hawaii. Get ready for many stimulating and insightful talks and discussions. Our Non-proceedings track is still open! Paper submission deadline: July 18, 2025 More info at: excv-workshop.github.io @iccv.bsky.social #ICCV2025
064
Reposted by Explainable AI
Gunnar König @gunnark.bsky.social · 07/07/2025
In many XAI applications, it is crucial to determine whether features contribute individually or only when combined. However, existing methods fail to reveal cooperations since they entangle individual contributions with those made via interactions and dependencies. We show how to disentangle them!
1173
Reposted by Explainable AI
Dana Arad @danaarad.bsky.social · 27/05/2025
Tried steering with SAEs and found that not all features behave as expected? Check out our new preprint - "SAEs Are Good for Steering - If You Select the Right Features" 🧵
2186
Reposted by Explainable AI
nikhil07prakash.bsky.social @nikhil07prakash.bsky.social · 24/06/2025
How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!
25918
Reposted by Explainable AI
Dana Arad @danaarad.bsky.social · 26/06/2025
VLMs perform better on questions about text than when answering the same questions about images - but why? and how can we fix it? In a new project led by Yaniv (@YNikankin on the other app), we investigate this gap from an mechanistic perspective, and use our findings to close a third of it! 🧵
164
Reposted by Explainable AI
BlackboxNLP @blackboxnlp.bsky.social · 23/06/2025
Have you heard about this year's shared task? 📢 Mechanistic Interpretability (MI) is quickly advancing, but comparing methods remains a challenge. This year at #BlackboxNLP, we're introducing a shared task to rigorously evaluate MI methods in language models 🧵
1164
Reposted by Explainable AI
Oliver Eberle @eberleoliver.bsky.social · 20/06/2025
Our position paper on algorithmic explanations is out—excited to share it! 🙌 Proud of this collaborative effort toward a scientifically grounded understanding of generative AI. @tuberlin.bsky.social @bifold.berlin @msftresearch.bsky.social @UCSD & @UCLA
1187
Reposted by Explainable AI
Laura Kopf @lkopf.bsky.social · 19/06/2025
🔍 When do neurons encode multiple concepts? We introduce PRISM, a framework for extracting multi-concept feature descriptions to better understand polysemanticity. 📄 Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework arxiv.org/abs/2506.15538 🧵 (1/7)
13611
Reposted by Explainable AI
Oliver Eberle @eberleoliver.bsky.social · 19/06/2025
🚨 New preprint! Excited to share our work on extracting and evaluating the potentially many feature descriptions of language models 👉 arxiv.org/abs/2506.15538
0184
Reposted by Explainable AI
Visual Inference Lab @visinf.bsky.social · 13/12/2024
Want to learn about how model design choices affect the attribution quality of vision models? Visit our #NeurIPS2024 poster on Friday afternoon (East Exhibition Hall A-C #2910)! Paper: arxiv.org/abs/2407.11910 Code: github.com/visinf/idsds
1217
Reposted by Explainable AI
Marianne de Heer Kloots @mdhk.net · 13/06/2025
The @interspeech.bsky.social early registration deadline is coming up in a few days! Want to learn how to analyze the inner workings of speech processing models? 🔍 Check out the programme for our tutorial: interpretingdl.github.io/speech-inter... & sign up through the conference registration form!
interpretingdl.github.io
Interpretability Techniques for Speech Models — Tutorial @ Interspeech 2025
12710
Reposted by Explainable AI
Sohee Yang @soheeyang.bsky.social · 13/06/2025
🚨 New Paper 🚨 How effectively do reasoning models reevaluate their thought? We find that: - Models excel at identifying unhelpful thoughts but struggle to recover from them - Smaller models can be more robust - Self-reevaluation ability is far from true meta-cognitive awareness 1/N 🧵
1123
Reposted by Explainable AI
Sweta Mahajan @swetamahajan.bsky.social · 14/06/2025
Join us in advancing explainability in the rapidly evolving landscape of AI! We welcome papers on novel explanation methods, scaling existing xAI approaches, innovations for foundational models, and beyond.
152
Reposted by Explainable AI
Sweta Mahajan @swetamahajan.bsky.social · 14/06/2025
Call for papers is out for the 2nd Explainable Computer Vision - Quo Vadis?
 (eXCV) workshop in #XAI at #ICCV2025! More details can be found here: excv-workshop.github.io Submission Site: Coming Soon! @iccv.bsky.social @xai-research.bsky.social
132
Reposted by Explainable AI
Hendrik Strobelt @henstr.bsky.social · 11/06/2025
It's on! In its 8th year, the original VISxAI workshop wants your contributions for explaining ML/AI/GenAI principles. Please consider to submit. Deadline: July 30 webpage: visxai.io #ML #AI #GenAI #Visualization
076
Reposted by Explainable AI
sqIRL Lab @sqirllab.bsky.social · 02/06/2025
Our lab got two papers accepted at #ECMLPKDD2025 on the topics of #Interpretability for Spiking NNs and self-supervised representation learning with embedded interpretability . Congrats to Jasper, Hamed, Fabian and our collaborators. #SNN #SIM #AI #ML #neuromorphic #xai #interpretableML
031
Reposted by Explainable AI
Computer Vision and Machine Learning at MPI Informatics @cvml.mpi-inf.mpg.de · 03/06/2025
Exciting work on making neural networks interpretable from our department by Jonas Fischer!
042
Reposted by Explainable AI
Nishant Subramani @ ACL @nsubramani23.bsky.social · 04/06/2025
🚨 Check out our new #interpretability paper: 🕵🏽 Model Internal Sleuthing led by the amazing @bearseascape.bsky.social who is an undergrad at @scsatcmu.bsky.social @ltiatcmu.bsky.social
041
Reposted by Explainable AI
Tim van Erven @timvanerven.nl · 30/05/2025
Deadlines for PhD and Postdoc vacancies coming up: applications open until Monday June 2!
055
Reposted by Explainable AI
Yoav Gur Arieh @yoav.ml · 29/05/2025
New Paper Alert! Can we precisely erase conceptual knowledge from LLM parameters? Most methods are shallow, coarse, or overreach, adversely affecting related or general knowledge. We introduce🪝𝐏𝐈𝐒𝐂𝐄𝐒 — a general framework for Precise In-parameter Concept EraSure. 🧵 1/
152
Reposted by Explainable AI
Chhavi Yadav @chhaviyadav.bsky.social · 28/05/2025
Bringing to you my latest paper, to be presented at #ICML2025 this July -- ‘ExpProof : Operationalizing Explanations for Confidential Models with ZKPs’ Paper Link : arxiv.org/abs/2502.03773 Code : github.com/emlaufer/Exp...
072
Reposted by Explainable AI
Eliana Pastor @elianapastor.bsky.social · 23/05/2025
Today, we held the first XAI seminar at the Explainable and Trustworthy AI course at PoliTO 🚀 Alan Perotti (CENTAI) presented his work on Operationalising XAI, describing practical challenges in XAI! Thank you again, Alan! Great start to our series!
051
Reposted by Explainable AI
Gabriele Sarti @gsarti.com · 23/05/2025
Prev. work showed that SAEs can be helpful beyond interpretability, to surgically condition model behaviors. In this work, we take a small step towards making steering useful for MT personalization on literary works! Check out our thread! ⬇️ Paper: arxiv.org/abs/2505.16612
arxiv.org
Steering Large Language Models for Machine Translation Personalization
High-quality machine translation systems based on large language models (LLMs) have simplified the production of personalized translations reflecting specific stylistic constraints. However, these sys...
1202
Reposted by Explainable AI
kyunghyuncho.bsky.social @kyunghyuncho.bsky.social · 20/05/2025
this is probably not the complete picture of KD, but i can definitely sleep better after writing down and confirming this minimal working explanation. arXiv: arxiv.org/abs/2505.13111 (3/4)
arxiv.org
Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
Knowledge distillation (KD) is a core component in the training and deployment of modern generative models, particularly large language models (LLMs). While its empirical benefits are well documented-...
272
Reposted by Explainable AI
BlackboxNLP @blackboxnlp.bsky.social · 15/05/2025
BlackboxNLP, the leading workshop on interpretability and analysis of language models, will be co-located with EMNLP 2025 in Suzhou this November! 📆 This edition will feature a new shared task on circuits/causal variable localization in LMs, details here: blackboxnlp.github.io/2025/task
3218
Reposted by Explainable AI
Antonin Poché @antoninpoche.bsky.social · 31/01/2025
🚀 Thrilled to share our new paper (the first of my PhD)! How can we compare concept-based #XAI methods in #NLProc? ConSim (arxiv.org/abs/2501.05855) provides the answer. Read the thread to find out which method is the most interpretable! 🧵1/7
662
Reposted by Explainable AI
Zara Siddique @zarasiddique.bsky.social · 14/05/2025
🚨 NEW PAPER ALERT 🚨 Dialz: A Python Toolkit for Steering Vectors ArXiv: arxiv.org/abs/2505.06262 Docs: cardiffnlp.github.io/dialz/ Repo: github.com/cardiffnlp/d... A Python package to help you create, apply and visualise steering vectors for anything you want - from sycophancy to bias.
2268
Reposted by Explainable AI
Katharina Prasse @katharinaprasse.bsky.social · 08/05/2025
Happy to announce that we have an #ICML paper! Our DCBMs use cropped image segmentats for model interpretability. It was great collaborating with @patrickknab.bsky.social, @smarton.bsky.social , Christian Bartelt, and @margretkeuper.bsky.social @keuper-labs.bsky.social
074
Reposted by Explainable AI
José Oramas @jaom7.bsky.social · 08/05/2025
It is confirmed, the #AIMLAI workshop will be held jointly with @ecmlpkdd.org. We invite the submissions of long and short papers covering work around #interpretability and #explainability of #AI/#ML. Deadline: 14/06/25 CfP: shorturl.at/yYQ9G Website: shorturl.at/W9r1A #XAI #mechinterp #ECMLPKDD
012
Reposted by Explainable AI
Sébastien Destercke @sebastiendestercke.bsky.social · 07/05/2025
Getting coverage guarantees over functional surrogate models? This is what A. Gray and V. Gopakumar (they did the heavy lifting) have done using conformal predictions and zonotopes over reduced dimensions. It will be present at UAI 2025 (@auai.org), but a preview is here: arxiv.org/abs/2501.18426.
arxiv.org
Guaranteed confidence-band enclosures for PDE surrogates
We propose a method for obtaining statistically guaranteed confidence bands for functional machine learning techniques: surrogate models which map between function spaces, motivated by the need build ...
032
Reposted by Explainable AI
Sunnie S. Y. Kim ☀️ @sunniesuhyoung.bsky.social · 07/05/2025
📢 I successfully defended my PhD dissertation! Huge thanks to my committee (Olga @andresmh.com @jennwv.bsky.social @qveraliao.bsky.social @parastooabtahi.bsky.social) & everyone who supported me ❤️ 📢 Next I'll join Apple as a research scientist in the Responsible AI team led by @jeffreybigham.com!
Sunnie standing in front of her presentation celebrating the successful defense 🎉 Vera, Andrés, Sunnie, Olga, and Jenn (on Sunnie’s laptop screen) celebratingGroup photo of everyone who joined Sunnie’s dissertation defenseLauren, Sunnie, and Jeff (photo taken at CHI 2025)
6595
Reposted by Explainable AI
Tim van Erven @timvanerven.nl · 06/05/2025
And the video of Gunnar's talk is up on YouTube in case you missed it: youtu.be/7MrMjabTbuM @gunnark.bsky.social
youtu.be
Theory of Interpretable AI Seminar: Gunnar König
YouTube video by Theory of Interpretable AI Seminar
0133
Reposted by Explainable AI
Hadas Orgad @hadasorgad.bsky.social · 03/05/2025
Deadline extended! ⏳ The Actionable Interpretability Workshop at #ICML2025 has moved its submission deadline to May 19th. More time to submit your work 🔍🧠✨ Don’t miss out!
043
Reposted by Explainable AI
Sanghamitra Dutta @sanghamd.bsky.social · 07/04/2025
🔈 Sharing our recent paper on "Counterfactual Explanations for Model Ensembles Using Entropic Risk Measures" 🎉 🎉 Accepted at #AAMAS2025 Joint work with: Erfaun Noorani Pasan Dissanayake Faisal Hamman #Explainability #XAI #AlgorithmicRecourse #EntropicRisk arxiv.org/abs/2503.07934
arxiv.org
Counterfactual Explanations for Model Ensembles Using Entropic Risk Measures
Counterfactual explanations indicate the smallest change in input that can translate to a different outcome for a machine learning model. Counterfactuals have generated immense interest in high-stakes...
011
Reposted by Explainable AI
Sanghamitra Dutta @sanghamd.bsky.social · 02/05/2025
📢 Knowledge distillation trains smaller student models from complex teacher models. But are all teachers equally helpful? Can we formally quantify useful distillable knowledge? Our paper at #AISTATS2025 explains distillation using Partial Information Decomposition. arxiv.org/abs/2411.07483
arxiv.org
Quantifying Knowledge Distillation Using Partial Information Decomposition
Knowledge distillation deploys complex machine learning models in resource-constrained environments by training a smaller student model to emulate internal representations of a complex teacher model. ...
101