Sign in

Andreas Waldis

@tresiwald.bsky.social
178 followers 650 following 30 posts

Behavioral and Internal Interpretability 🔎 PostDoc Tübingen University | previously PhD Student at @ukplab.bsky.social, TU Darmstadt/Hochschule Luzern

PostsRepliesMedia
Andreas Waldis @tresiwald.bsky.social · 02/10/2026
Is an LM that acts Bayesian also Bayesian inside? ☝️ Only as far as its beliefs allow! Fine-tuned on an optimal Bayesian model, LMs hold and use better beliefs than usual fine-tuning on golden answers. Details 👇 or bayeslm.github.io #interpretability #nlproc (1/🧵)
171
Andreas Waldis @tresiwald.bsky.social · 01/06/2026
It was a pleasure to 🍸
030
Andreas Waldis @tresiwald.bsky.social · 18/05/2026
Do instructions affect how LMs process and produce language? ☝️Not the way you think! 😲LMs barely change task information when processing a task sample. Instead, instructions shape how this information is accessed and expressed when producing output tokens. #interpretability #nlproc (1/🧵)
1141
Reposted by Andreas Waldis
Vagrant Gautam @dippedrusk.com · 27/03/2026
I already presented some work on reference (names, pronouns, coreference resolution, pronoun fidelity, etc.) as a rich site to evaluate biases and commonsense reasoning, and our work on disentangling model behaviour and internals through aligned probing (led by @tresiwald.bsky.social).
151
Andreas Waldis @tresiwald.bsky.social · 25/03/2026
Excited to present this work together with @dippedrusk.com at #EACL. Join us in the poster session 1 (11:30-13:00) 🔥
Poster of the Paper Aligned Probing
041
Andreas Waldis @tresiwald.bsky.social · 03/03/2026
Thanks a lot to everyone for the support, guidance, mentoring, collaboration, and great moments over the past years! 🙏 Without you, this journey wouldn't have been such a pleasure — and now excited to see what the future brings! 🚀
151
Andreas Waldis @tresiwald.bsky.social · 27/01/2026
LMs that "know more" about toxicity are less toxic! Our #TACL 📄 connects behavior and internals: 💠 LMs amplify toxicity beyond humans 💠 Information about toxicity peaks in lower layers 💠 Bypassing these layers increases toxicity More details👇 #NLProc #interpretability (1/🧵)
simplified overview of our aligned probing setup, where we join the behavioral and internal evaluation of LMs' toxicity
1167
Reposted by Andreas Waldis
INTERPLAY Workshop@COLM '25 @interplay-workshop.bsky.social · 09/10/2025
✨ The schedule for our INTERPLAY workshop at COLM is live! ✨ 🗓️ October 10th, Room 518C 🔹 Invited talks from @sarah-nlp.bsky.social John Hewitt @amuuueller.bsky.social @kmahowald.bsky.social 🔹 Paper presentations and posters 🔹 Closing roundtable discussion. Join us in Montréal! @colmweb.org
Schedule for the INTERPLAY workshop at COLM on October 10th, Room 518C.

09:00 am: Opening
09:10 am: Invited Talks by Sarah Wiegreffe and John Hewitt
10:20 am: Paper Presentations

Lunch Break

01:00 pm: Invited Talks by Aaron Mueller and Kyle Mowhald
02:10 pm: Poster Session
03:20 pm: Roundtable Discussion
04:50 pm: Closing
044
Reposted by Andreas Waldis
INTERPLAY Workshop@COLM '25 @interplay-workshop.bsky.social · 08/07/2025
Missed a spot? If you have a pre-reviewed paper from ARR or COLM that focuses on the INTERPLAY between LM internals and behavior, there is a shortcut to presenting at our @colmweb.org workshop! ✨ Join us in Montréal! 🇨🇦 CfP: shorturl.at/sBomu OpenReview: shorturl.at/WwWhg #nlproc #interpretability
Call for Pre-Reviewed Papers, Interplay Workshop at COLM: July 10th - submissions due. July 24th - acceptance notification. October 10th - workshop day.
152
Reposted by Andreas Waldis
INTERPLAY Workshop@COLM '25 @interplay-workshop.bsky.social · 24/06/2025
Delighted that ✨Mor Geva (@megamor2.bsky.social) and ✨Anna Ivanova (@neuranna.bsky.social) will complete our speaker line-up and talk about the INTERPLAY of model internals and behavior. Be there and submit by June 30th 📄 shorturl.at/sBomu See you in 🇨🇦 @colmweb.org #nlproc #interpretability
Mor Geva and Anna Ivanova will talk at the INTERPLAY workshop.
053