Sign in

Tal Linzen

@tallinzen.bsky.social
3.1K followers 80 following 24 posts

NYU professor, Google research scientist. Good at LaTeX.

PostsRepliesMedia
Reposted by Tal Linzen
NYU Center for Data Science @nyudatascience.bsky.social · 02/10/2026
Can LLMs introspect? Anthropic said yes. However, CDS PhD student Shashwat Singh, CDS Associate Professor Tal Linzen (@tallinzen.bsky.social) & CDS Faculty Fellow Shauli Ravfogel (@shauli.bsky.social) found the evidence falls short. nyudatascience.medium.com/cds-research...
nyudatascience.medium.com
CDS Researchers Challenge Anthropic’s Evidence That Language Models Can Introspect
In 2025, Anthropic reported that its Claude models could detect when researchers injected a concept directly into their neural activity…
13416
Reposted by Tal Linzen
Andreas Waldis @tresiwald.bsky.social · 02/10/2026
Is an LM that acts Bayesian also Bayesian inside? ☝️ Only as far as its beliefs allow! Fine-tuned on an optimal Bayesian model, LMs hold and use better beliefs than usual fine-tuning on golden answers. Details 👇 or bayeslm.github.io #interpretability #nlproc (1/🧵)
171
Reposted by Tal Linzen
Ryan Moulton @moultano.bsky.social · 28/09/2026
A political social network has latched on to a handful of politically active academics whose views are the vast minority in their field, who keep making wrong predictions and denying verifiable facts, but retain a large following. Bluesky recognizes this pattern with climate denial, but not with AI
1032644
Reposted by Tal Linzen
Dario Paape @dariopaape.eurosky.social · 25/09/2026
Updated preprint with @tallinzen.bsky.social and @shravanvasishth.bsky.social : "LLM surprisal is necessary but not sufficient to capture English garden-path effects: Evidence from joint latent modeling of reading paradigms" arxiv.org/abs/2602.04489
032
Reposted by Tal Linzen
NYU Center for Data Science @nyudatascience.bsky.social · 16/09/2026
To make AI better at imitating humans, limit its memory. Ex-CDS Faculty Fellow Nicholas Tomlin, CDS PhD @michahu.bsky.social, & CDS Prof @tallinzen.bsky.social show that capping a model's memory at 4 slots makes it a far better stand-in for a real person. nyudatascience.medium.com/forcing-ai-t...
nyudatascience.medium.com
Forcing AI to Forget: Constraining memory improves human simulation
Have you ever talked to an AI model and found the response so long and complex that you were unable to remember any of it? The model has no…
032
Reposted by Tal Linzen
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n
Overview of the paper. 
Title: The Emergent Symbolic Structure of Artificial Neural Networks
Authors: Tom McCoy, Paul Soulos, Tal Linzen, Paul Smolensky
Left: Neural networks encode information in vectors (there is then an image of a vector), yet they excel at tasks long thought to require symbolic structure (there is then an image of a symbolic representation, specifically a syntax tree). How do LLMs do it?
Right: We find that LLM representations can be closely approximated with symbolic structures. This approximation lets us edit the structure of an LLM’s output by editing the structure of its internal representations, as shown. There is then an image of two edits to LLMs. In the first one, the original input is 3 + 6 * 8, with an answer of 51. But if we swap the positions of the 3 and the 6, the output becomes 30. In the second one, the original input is a Python command repeating the list [Z, U] three times, producing [Z, U, Z, U, Z, U]. But if we edit the input in a way that adds a Q at the end of the input, the output becomes [Z, U, Q, Z, U, Q, Z, U, Q].
431888
Tal Linzen @tallinzen.bsky.social · 26/08/2026
Thanks to the National Science Foundation for featuring our new @pnas.org paper, and for the continued support for basic science projects like this one!
1183
Reposted by Tal Linzen
William Timkey @wtimkey.bsky.social · 10/08/2026
Why do we breeze through some sentences, but others make us slow down and reread? A popular answer is predictability: unexpected words are harder to process. In our new @pnas.org article, we used LMs as models of human prediction to ask how far this explanation can actually go🧵
1276
Reposted by Tal Linzen
Dario Paape @dariopaape.eurosky.social · 15/07/2026
Who's attending #CogSci2026 in Rio next week? 🇧🇷 I will be presenting joint work with @shravanvasishth.bsky.social and @tallinzen.bsky.social on LLM surprisal in garden-path sentences on Friday! arxiv.org/abs/2602.04489
0174
Reposted by Tal Linzen
Nick Tomlin @nickatomlin.bsky.social · 27/05/2026
This work was done with Qihan Wang, Michael Hu, @linguistbrian.bsky.social, and @tallinzen.bsky.social ! We’re excited about the potential for leveraging ideas from cogsci/linguistics and using them to improve user sims, which can be used to train models that collaborate better with real humans
041
Reposted by Tal Linzen
Nick Tomlin @nickatomlin.bsky.social · 27/05/2026
New paper! LLM memory keeps improving, but this makes them *worse* as user sims. If we want to build models that can, e.g., simulate realistic students to train chatbots to be better teachers, then these models need to be able to forget like humans do 📄: arxiv.org/abs/2605.25680
arxiv.org
Simulating Human Memory with Language Models
Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic memory experiments fr...
1251
Tal Linzen @tallinzen.bsky.social · 09/12/2025
excited that the Society for Computation in Linguistics (SCiL) will be colocated with #acl2026nlp this year, and I'm grateful to the National Science Foundation for helping support SCiL presenters' registration costs! (keynotes: Jenn Hu and Noah Smith deadline: Jan 30 conference: July 3 & 4)
060
Reposted by Tal Linzen
Brian Dillon @linguistbrian.bsky.social · 14/11/2025
Really big announcement! See @wtimkey.bsky.social's thread for the details on an exciting new preprint from the NYU-UMass Syntactic Ambiguity Processing group. It is the culmination of the team's research efforts over these last couple of years, and we're really happy with it.
1133
Reposted by Tal Linzen
William Timkey @wtimkey.bsky.social · 14/11/2025
New Preprint: osf.io/eq2ra Reading feels effortless, but it's actually quite complex under the hood. Most words are easy to process, but some words make us reread or linger. It turns out that LLMs can tell us about why, but only in certain cases... (1/n)
2135
Reposted by Tal Linzen
Shauli Ravfogel @shauli.bsky.social · 24/10/2025
New NeurIPS paper! Why do LMs represent concepts linearly? We focus on LMs's tendency to linearly separate true and false assertions, and provide an analysis of the truth circuit in a toy model. A joint work with Gilad Yehudai, @tallinzen.bsky.social, Joan Bruna and @albertobietti.bsky.social.
1265
Reposted by Tal Linzen
Jaap Jumelet @jumelet.bsky.social · 15/10/2025
🌍Introducing BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data! LLMs learn from vastly more data than humans ever experience. BabyLM challenges this paradigm by focusing on developmentally plausible data We extend this effort to 45 new languages!
14416
Reposted by Tal Linzen
Cameron Buckner @cameronbuckner.bsky.social · 15/10/2025
Another banger from @tallinzen.bsky.social . Also fits with some of the criticisms of Centaur and my faculty-based approach generally; if you want LLMs to model human cognition, give them more architecture akin to human faculty psychology like long and short-term memory. arxiv.org/abs/2510.05141
arxiv.org
To model human linguistic prediction, make LLMs less superhuman
When people listen to or read a sentence, they actively make predictions about upcoming words: words that are less predictable are generally read more slowly than predictable ones. The success of larg...
1236
Reposted by Tal Linzen
NYU Center for Data Science @nyudatascience.bsky.social · 10/09/2025
Linguistics PhD student @jacksonpetty.org finds LLMs "quiet-quit" when instructions get long, switching from reasoning to guesswork. With CDS' @tallinzen.bsky.social, @shauli.bsky.social, @lambdaviking.bsky.social, @michahu.bsky.social, and Wentao Wang. nyudatascience.medium.com/llms-switch-...
nyudatascience.medium.com
LLMs Switch to Guesswork Once Instructions Get Long
LLMs abandon reasoning for guesswork when instructions get long, new work from Linguistics PhD student Jackson Petty & CDS shows.
072
Reposted by Tal Linzen
Dr. Lucky Tran @luckytran.com · 11/07/2025
DO NOT GIVE UP! Our advocacy is working. A key Senate committee has indicated that it will reject Trump’s proposed cuts to science agencies including NASA and the NSF. Keep speaking up and calling your electeds 🗣️🗣️🗣️
Nature: US senators poised to reject Trump’s proposed massive science cuts

Committee gives first hint that policymakers might preserve, rather than slash, funding for US National Science Foundation and other agencies.
81327438
Tal Linzen @tallinzen.bsky.social · 11/07/2025
Congratulations to @linguistbrian.bsky.social for receiving this grant to study how to constrain language models to read complex sentences more like humans, and congratulations to me for getting to collaborate with him for another four years! www.umass.edu/humanities-a...
umass.edu
Brian Dillon Receives NSF Grant to Explore AI and Human Language Processing : College of Humanities & Fine Arts : UMass Amherst
Linguist Brian Dillon receives NSF grant to investigate how AI and humans differ in interpreting meaning during language comprehension.
1181
Tal Linzen @tallinzen.bsky.social · 02/07/2025
My Twitter account has been hacked :( Please don't click on any links "I" posted on that account recently!
121
Tal Linzen @tallinzen.bsky.social · 21/06/2025
I'm hiring at least one post-doc! We're interested in creating language models that process language more like humans than mainstream LLMs do, through architectural modifications and interpretability-style steering. Express interest here: docs.google.com/forms/d/e/1F...
docs.google.com
NYU LLM + cognitive science post-doc interest form
Tal Linzen's group at NYU is hiring a post-doc! We're interested in creating language models that process language more like humans than mainstream LLMs do, through architectural modifications and int...
24221
Reposted by Tal Linzen
Jackson Petty @jacksonpetty.org · 09/06/2025
How well can LLMs understand tasks with complex sets of instructions? We investigate through the lens of RELIC: REcognizing (formal) Languages In-Context, finding a significant overhang between what LLMs are able to do theoretically and how well they put this into practice.
152
Reposted by Tal Linzen
Arianna Bisazza @arianna-bis.bsky.social · 30/05/2025
Following the success story of BabyBERTa, I & many other NLPers have turned to language acquisition for inspiration. In this new paper we show that using Child-Directed Language as training data is unfortunately *not* beneficial for syntax learning, at least not in the traditional LM training regime
1246
Tal Linzen @tallinzen.bsky.social · 23/05/2025
Cross-posting the abstracts for two talks I'm giving next week! This one on formal languages for LLM pretraining and evaluation, at Apple ML Research in Copenhagen on Wednesday
1102
Tal Linzen @tallinzen.bsky.social · 12/05/2025
Updated version of our position piece on how language models can help us understand how people learn and process language, on why it's crucial to train models on cognitive plausible datasets, and on the BabyLM project that addresses this issue.
0111
Reposted by Tal Linzen
Dario Paape @dariopaape.eurosky.social · 14/03/2025
At #HSP2025, I'll present work with @tallinzen.bsky.social and @shravanvasishth.bsky.social on modeling garden-pathing in a huge benchmark dataset: hsp2025.github.io/abstracts/29.... Statistically decomposing the effect into subprocesses greatly improves predictive fit over just comparing means!
hsp2025.github.io
0112
Tal Linzen @tallinzen.bsky.social · 27/03/2025
Going to give this website another shot! What are good lists of linguistics, psycholinguistics, NLP and AI accounts?
4150
Tal Linzen @tallinzen.bsky.social · 19/11/2023
Thanks Ted for mentioning me in the same tweet as Chris! This website really is better than the other one!
040
Tal Linzen @tallinzen.bsky.social · 19/11/2023
Very little happening on here but silence is certainly better than all of the boardroom drama takes on the other website. Four different people I follow just came up with the same unfunny joke about the most recent development in the drama, apparently independently?
0100
Reposted by Tal Linzen
Jackson Petty @jacksonpetty.org · 10/11/2023
Do deep transformer LMs generalize better? In a new preprint we (Sjoerd van Steenkiste, Ishita Dasgupta, Fei Sha, Dan Garrette, & @tallinzen.bsky.social) control for parameter count to show how depth helps models on compositional generalization tasks, but diminishingly so 🧵 jacksonpetty.org/depth
jacksonpetty.org
The Impact of Depth and Width on Transformer Language Model Generalization
To process novel sentences, language models (LMs) must generalize compositionally -- combine familiar elements in new ways. What aspects of a model's structure promote compositional generalization? Fo...
1105