Reposted by Tal Linzen
Can LLMs introspect?
Anthropic said yes.
However, CDS PhD student Shashwat Singh, CDS Associate Professor Tal Linzen (@tallinzen.bsky.social) & CDS Faculty Fellow Shauli Ravfogel (@shauli.bsky.social) found the evidence falls short.
nyudatascience.medium.com/cds-research...
nyudatascience.medium.com
CDS Researchers Challenge Anthropic’s Evidence That Language Models Can Introspect
In 2025, Anthropic reported that its Claude models could detect when researchers injected a concept directly into their neural activity…