Sign in

William Timkey

@wtimkey.bsky.social
107 followers 193 following 30 posts

Computational psycholinguistics PhD student @NYU lingusitics | first gen!

PostsRepliesMedia
William Timkey @wtimkey.bsky.social · 10/08/2026
This result isn’t dependent on particular modeling assumptions. Using LLM surprisals, it turns out that the surprising words of garden path sentences (red) are re-read much more frequently than *equally surprising* words in everyday non-garden path sentences (blue).
110
William Timkey @wtimkey.bsky.social · 10/08/2026
However, LLMs drastically underpredicted the difficulty in cases where garden path sentence caused rereading. More specifically, LLMs’ surprisal estimates can’t explain the *rate* at which people choose to reread (bottom left), or the *amount of time* they spend rereading (bottom right).
120
William Timkey @wtimkey.bsky.social · 10/08/2026
Contrasting with prior results, the LLMs *were* able to explain the cost of processing garden path sentences, but only in cases where the participant didn’t go back and reread when the sentence got challenging (forward reading time).
120
William Timkey @wtimkey.bsky.social · 10/08/2026
We recorded eye movements of 368 participants when reading both simple and challenging sentences, and using surprisal estimates from >400 LLMs, we evaluated which eye-movement signatures of reading difficulty the prediction account can and cannot capture.
100
William Timkey @wtimkey.bsky.social · 10/08/2026
The prediction view is known to explain reading difficulty in simple non-garden path sentences, but has also been shown to drastically underpredict the difficulty readers experience in garden path sentences. (e.g. this result from Huang et al. 2024)
110
William Timkey @wtimkey.bsky.social · 10/08/2026
The answer might lie in a classic type of sentence that causes processing difficulty: “The boy fed the chicken wanted beef instead” This is a garden path sentence: a sentence that initially appears to mean one thing, but ends up meaning something different.
110
William Timkey @wtimkey.bsky.social · 10/08/2026
Why do we breeze through some sentences, but others make us slow down and reread? A popular answer is predictability: unexpected words are harder to process. In our new @pnas.org article, we used LMs as models of human prediction to ask how far this explanation can actually go🧵
1276
William Timkey @wtimkey.bsky.social · 14/11/2025
What's going on in the cases predictability can't explain? When humans reread in syntactically challenging sents, they focus on parts of the sentence that are most helpful for revising the structure (in this case, the verb) consistent with structural processing accounts. (11/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
But, LLMs drastically underpredicted the magnitude of difficulty whenever syntactically challenging parts of a sentence caused comprehenders to reread. More specifically, LLMs can’t predict the rate at which comprehenders reread, or the amount of time they spend rereading. (10/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
We recorded eye movements of 368 participants when reading syntactically challenging sentences. From these eye movements we derive multiple measures of processing difficulty, and evaluate which measures the prediction account can and cannot explain using >400 different LLMs.(8/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
One example is garden path sentences, where LLMs have been shown to only predict a small fraction of the full processing cost. (5/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
The prediction view has been shown to explain reading behavior in structurally simple sentences, but drastically underpredicts the difficulty readers experience in syntactically complex/challenging sentences. The opposite is true of the structural processing view. (4/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
New Preprint: osf.io/eq2ra Reading feels effortless, but it's actually quite complex under the hood. Most words are easy to process, but some words make us reread or linger. It turns out that LLMs can tell us about why, but only in certain cases... (1/n)
2135