Huge thanks to my amazing co-authors @butanium.bsky.social, Stewart Slocum, Helena Casademunt, @cameronholmes.bsky.social, Robert West @neelnanda.bsky.social
Paper: www.arxiv.org/abs/2510.13900
(9/9)
arxiv.org
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
Finetuning on narrow domains has become an essential tool to adapt Large Language Models (LLMs) to specific tasks and to create models with known unusual properties that are useful for research. We sh...