Sign in

William Timkey

@wtimkey.bsky.social
107 followers 193 following 30 posts

Computational psycholinguistics PhD student @NYU lingusitics | first gen!

PostsRepliesMedia
Reposted by William Timkey
Tal Linzen @tallinzen.bsky.social · 26/08/2026
Thanks to the National Science Foundation for featuring our new @pnas.org paper, and for the continued support for basic science projects like this one!
1183
William Timkey @wtimkey.bsky.social · 10/08/2026
This article is the culmination of a huge collaborative undertaking with my brilliant co-authors: Kuan-Jung Huang, @byungdoh.bsky.social, @grushaprasad.bsky.social @sarehalli.bsky.social, @tallinzen.bsky.social and @linguistbrian.bsky.social
pnas.org
Eye movements reveal a dissociation between prediction and structural processing difficulty in language comprehension | PNAS
In the process of extracting a meaning from a text, our eyes linger much more on some words than others, and we often reread earlier portions of th...
020
William Timkey @wtimkey.bsky.social · 10/08/2026
Summary: Prediction and structural processing have dissociable influences on reading difficulty. LLMs can be a good model of when our eyes move forward, but can’t explain when, where, or why humans reread. To model rereading, we need models that go beyond next-word prediction.
120
William Timkey @wtimkey.bsky.social · 10/08/2026
This result isn’t dependent on particular modeling assumptions. Using LLM surprisals, it turns out that the surprising words of garden path sentences (red) are re-read much more frequently than *equally surprising* words in everyday non-garden path sentences (blue).
110
William Timkey @wtimkey.bsky.social · 10/08/2026
However, LLMs drastically underpredicted the difficulty in cases where garden path sentence caused rereading. More specifically, LLMs’ surprisal estimates can’t explain the *rate* at which people choose to reread (bottom left), or the *amount of time* they spend rereading (bottom right).
120
William Timkey @wtimkey.bsky.social · 10/08/2026
Contrasting with prior results, the LLMs *were* able to explain the cost of processing garden path sentences, but only in cases where the participant didn’t go back and reread when the sentence got challenging (forward reading time).
120
William Timkey @wtimkey.bsky.social · 10/08/2026
We recorded eye movements of 368 participants when reading both simple and challenging sentences, and using surprisal estimates from >400 LLMs, we evaluated which eye-movement signatures of reading difficulty the prediction account can and cannot capture.
100
William Timkey @wtimkey.bsky.social · 10/08/2026
No single theory can account for all of the data. Perhaps *both* prediction and structural (re)processing drive reading difficulty. Can we find any evidence that they dissociate, and if so, what are their behavioral signatures?
100
William Timkey @wtimkey.bsky.social · 10/08/2026
The structural (re)processing view, by contrast, can predict difficulty in garden path sentences, but cannot explain which words cause reading difficulty in simple everyday sentences.
100
William Timkey @wtimkey.bsky.social · 10/08/2026
The prediction view is known to explain reading difficulty in simple non-garden path sentences, but has also been shown to drastically underpredict the difficulty readers experience in garden path sentences. (e.g. this result from Huang et al. 2024)
110
William Timkey @wtimkey.bsky.social · 10/08/2026
(2) The structural (re)processing view: reading difficulty partly reflects the mental cost assembling the words of a sentence into a larger meaning. This process runs into costly errors in garden path sents, requiring the reader to backtrack and search for the correct meaning.
100
William Timkey @wtimkey.bsky.social · 10/08/2026
(1) The prediction view: the cost of processing each word in any sentence reflects the word's surprisal. Garden path sents. are hard because they’re unpredictable. Predicting the next word is exactly what LLMs are trained to do, so they're a great tool for evaluating this view.
110
William Timkey @wtimkey.bsky.social · 10/08/2026
We conducted a large-scale (n=368) eyetracking while reading study to test two competing views of why these sentences are difficult:
100
William Timkey @wtimkey.bsky.social · 10/08/2026
There's now an emerging consensus that predictability (surprisal, specifically) can't capture the processing of garden path sentences. We find in our new study that it's not as simple as this: surprisal does capture difficulty in some measures, but not others
110
William Timkey @wtimkey.bsky.social · 10/08/2026
The answer might lie in a classic type of sentence that causes processing difficulty: “The boy fed the chicken wanted beef instead” This is a garden path sentence: a sentence that initially appears to mean one thing, but ends up meaning something different.
110
William Timkey @wtimkey.bsky.social · 10/08/2026
Paper: www.pnas.org/doi/10.1073/... Press release: www.nyu.edu/about/news-p...
pnas.org
Eye movements reveal a dissociation between prediction and structural processing difficulty in language comprehension | PNAS
In the process of extracting a meaning from a text, our eyes linger much more on some words than others, and we often reread earlier portions of th...
110
William Timkey @wtimkey.bsky.social · 10/08/2026
Why do we breeze through some sentences, but others make us slow down and reread? A popular answer is predictability: unexpected words are harder to process. In our new @pnas.org article, we used LMs as models of human prediction to ask how far this explanation can actually go🧵
1276
William Timkey @wtimkey.bsky.social · 14/11/2025
Thanks!!
010
William Timkey @wtimkey.bsky.social · 14/11/2025
Preprint: osf.io/eq2ra This project was a huge undertaking with some fantastic co-authors: Kuan-Jung Huang, @byungdoh.bsky.social, Grusha Prasad, @sarehalli.bsky.social, @tallinzen.bsky.social and @linguistbrian.bsky.social (13/13)
osf.io
OSF
010
William Timkey @wtimkey.bsky.social · 14/11/2025
Summary: Prediction and structural processing have dissociable influences on reading behavior. LLMs are a good model of when our eyes move forward, but can’t explain when, where, or why humans reread. To model rereading, we need models that go beyond next-word prediction. (12/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
What's going on in the cases predictability can't explain? When humans reread in syntactically challenging sents, they focus on parts of the sentence that are most helpful for revising the structure (in this case, the verb) consistent with structural processing accounts. (11/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
But, LLMs drastically underpredicted the magnitude of difficulty whenever syntactically challenging parts of a sentence caused comprehenders to reread. More specifically, LLMs can’t predict the rate at which comprehenders reread, or the amount of time they spend rereading. (10/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
In contrast to prior results, the LLMs were able to explain the cost of processing syntactically challenging sentences, but only when the challenging parts of a sentence didn’t trigger re-reading (forward reading). (9/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
We recorded eye movements of 368 participants when reading syntactically challenging sentences. From these eye movements we derive multiple measures of processing difficulty, and evaluate which measures the prediction account can and cannot explain using >400 different LLMs.(8/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
This is tricky to test because we need: 1) fine-grained behavioral measures sensitive to different types of difficulty 2) a large sample of human readers 3) a diverse set of syntactically challenging stimuli 4) a diverse set of LLMs for predictability estimation. (7/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
No single theory can account for all of the data. Perhaps *both* prediction and structural processing drive reading behavior. Can we find any evidence that they dissociate, and if so, what are their behavioral signatures? (6/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
One example is garden path sentences, where LLMs have been shown to only predict a small fraction of the full processing cost. (5/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
The prediction view has been shown to explain reading behavior in structurally simple sentences, but drastically underpredicts the difficulty readers experience in syntactically complex/challenging sentences. The opposite is true of the structural processing view. (4/n)
100
William Timkey @wtimkey.bsky.social · 14/11/2025
(2) The prediction view: the cost of processing each word in a sentence can be fully reduced to the word’s contextual predictability (i.e. surprisal). Predicting the next word is exactly what LLMs are trained to do, so they’re a great tool for evaluating this view. (3/n)
101
William Timkey @wtimkey.bsky.social · 14/11/2025
We conducted a high-powered (n=368) eyetracking while reading study to test two competing views: (1) The structural processing view: eye movements reflect the cost of mentally assembling the words of a sentence into a larger meaning. (2/n)
101
William Timkey @wtimkey.bsky.social · 14/11/2025
New Preprint: osf.io/eq2ra Reading feels effortless, but it's actually quite complex under the hood. Most words are easy to process, but some words make us reread or linger. It turns out that LLMs can tell us about why, but only in certain cases... (1/n)
2135