Chain-of-Thought reasoning can sound plausible while being unfaithful. Is there a way to test that directly by looking inside the model?
Our new work led by @qiaw99.bsky.social w/ @apepa.bsky.social et al.
📜 Pre-print: arxiv.org/abs/2609.23065
Postdoctoral Researcher @ University of Groningen / GroNLP (@gronlp.bsky.social) interested in the interpretability and analysis of language models. Ex- BIFOLD, TU Berlin, DFKI. nfelnlp.github.io