Sign in

Zhengyang Shan

@shanzzyy.bsky.social
5 followers 3 following 13 posts

PhD @ Boston University | Researching interpretability & evaluation in large language models

PostsRepliesMedia
Reposted by Zhengyang Shan
Micah Benson @micahben.bsky.social · 10/06/2026
🧠🤖 The 2026 New England Mechanistic Interpretability (NEMI) Workshop will be Aug. 14 at Boston University! Help spread the word and join the New England mech interp community! Registration and submission info in thread:👇
BU campus and Boston skyline
1102
Zhengyang Shan @shanzzyy.bsky.social · 20/01/2026
Can steering remove LLM shortcuts without breaking legitimate LLM capabilities? In our @eaclmeeting.bsky.social paper, we show that conceptual bias is separable from concept detection; this means inference-time debiasing is possible with minimal capability loss.
630