Sign in

thomasmcgee.bsky.social

@thomasmcgee.bsky.social
25 followers 25 following 16 posts

Cognitive Neuroscience PhD Student at UCLA

PostsRepliesMedia
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Many thanks to my wonderful collaborators @ibandlank.bsky.social and Daphne Zhang! 8/8
000
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Takeaway: even in attention heads that look “syntax-specialized,” computations are not encapsulated—they are penetrable to semantic information. This aligns with evidence against syntactic encapsulation in human sentence processing. 7/8
110
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
This finding occurs even when sentences are syntactically unambiguous (no need to rely on semantics for disambiguation). Semantic influence is not optionally triggered, but fundamental to how these heads operate. 6/8
100
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Result: across LLMs, a head’s attention to its preferred dependency is lower for implausible sentences. In some cases, attention even shifts toward “lure” words that are semantically plausible but syntactically not part of the dependency. 5/8
111
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Key prediction: If these heads are encapsulated → attention to the dependency should be the same. If they are penetrable → attention should be weaker when the dependency is semantically implausible (even though it’s syntactically valid). 4/8
110
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Method: identify heads in BERT / GPT-2 / Llama-2 that prefer different dependencies (e.g., subject-verb, verb-direct object, preposition-object). Then, test their attention using minimal pairs of sentences where the preferred dependency is semantically plausible or implausible. 3/8
100
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
LLMs implement many syntactic computations. We focus on attention heads claimed to be “specialized” for specific syntactic dependencies. Why? If encapsulated syntax exists anywhere, these are the best a priori candidates—maybe they allocate attention based only on syntax? 2/8
100
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
New paper: Evidence Against Syntactic Encapsulation in Large Language Models We revisit a classic psycholinguistic question, but in LLMs: do LLMs have any syntactic computations that are isolated from meaning, or does semantics pervasively influence them? 1/8 onlinelibrary.wiley.com/doi/10.1111/...
onlinelibrary.wiley.com
Evidence Against Syntactic Encapsulation in Large Language Models
Transformer-based large language models (LLMs) have recently demonstrated exceptional performance in a variety of linguistic tasks. LLMs primarily combine information across words in a sentence using...
1225
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Many thanks to my wonderful collaborators @ibandlank.bsky.social and Yiyang Zhang! 8/8
000
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Takeaway: even in attention heads that look “syntax-specialized,” computations are not encapsulated—they are penetrable to semantic information. This aligns with evidence for lack of syntactic encapsulation in human sentence processing. 7/8
100
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
This finding occurs even when sentences are syntactically unambiguous (no need to rely on semantics for disambiguation). Semantic influence is not optionally triggered, but fundamental to how these heads operate. 6/8
100
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Result: across LLMs, a head’s attention to its preferred dependency is lower for implausible sentences. In some cases, attention even shifts toward “lure” words that are semantically plausible but syntactically not part of the dependency. 5/8
100
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Key prediction: If these heads are encapsulated → attention to the dependency should be the same. If they are penetrable → attention should be weaker when the dependency is semantically implausible (even though it’s syntactically valid). 4/8
100
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
Method: identify heads in BERT / GPT2 / Llama2 that prefer different dependencies (e.g., subject-verb, verb-direct object, preposition-object). Then, test their attention using minimal pairs of sentences where the preferred dependency is semantically plausible or implausible. 3/8
100
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 02/05/2026
LLMs implement many syntactic computations. We focus on attention heads claimed to be specialized for specific syntactic dependencies. Why? If encapsulated syntax exists anywhere, these are the best a priori candidates—maybe they allocate attention based only on syntax? 2/8
100
thomasmcgee.bsky.social @thomasmcgee.bsky.social · 07/04/2026
Reposting our 2025 ICML paper on the algorithmic evaluation of generative AI! #ICML2025
010