LLMs have been shown to provide different predictions in clinical tasks when patient race is altered. Can SAEs spot this undue reliance on race? 🧵
Work w/ @byron.bsky.social
Link: arxiv.org/abs/2511.00177
PhD student @ Northeastern University, Clinical NLP hibaahsan.github.io she/her