Sign in

Tom McCoy

@rtommccoy.bsky.social
2.2K followers 369 following 286 posts

Assistant professor at Yale Linguistics. Studying computational linguistics, cognitive science, and AI. He/him.

PostsRepliesMedia
Tom McCoy @rtommccoy.bsky.social · 28/09/2026
My sister and I saw the Lion King in mid-September, and I just realized that this was perfect timing: the pride goes before the fall
050
Tom McCoy @rtommccoy.bsky.social · 27/09/2026
A nor'easter is a storm with winds from the northeast So you might think a sou'wester would be a storm with winds from the southwest. But it's not!! It's a type of hat!! Language strikes again
An image of a sou'wester - a yellow waterproof rain hat - sourced from Wikipedia
0201
Tom McCoy @rtommccoy.bsky.social · 18/09/2026
Had a great time speaking at NYU today - there's an incredible collection of people there working on computational CogSci & AI!
0111
Reposted by Tom McCoy
Elliot Murphy @elliot-murphy.bsky.social · 15/09/2026
New work out today! The result of a 15 year project to offer genuine linking hypotheses between linguistics and neuroscience: a revised research program, and its first result - a novel neural binding operation ('Meld') capturing core properties of natural language syntax 🧵 arxiv.org/abs/2609.14384
arxiv.org
Formal Properties of Language as Constraints on Neural Dynamics
What must a neural system be capable of to implement language? Current research annotates stimuli with linguistic variables and tests which electrodes, voxels, or language-model layers predict neural ...
26915
Tom McCoy @rtommccoy.bsky.social · 07/09/2026
You're welcome, and thanks for the question - it's an important one!
010
Tom McCoy @rtommccoy.bsky.social · 07/09/2026
But that's not the case here: Once we fix a structural hypothesis ("role scheme"), our TPR approximation model is highly constrained, as it's bilinear. There's no guarantee that it can approximate anything, so when it *does* succeed, it's strong evidence that the hypothesized TPR structure is there
100
Tom McCoy @rtommccoy.bsky.social · 07/09/2026
One more point: Suppose our approximation model were a universal function approximator, like a standard neural net. Then, our approximation wouldn't mean much, since we could approximate anything, even something that didn't use the structure we hypothesized
120
Tom McCoy @rtommccoy.bsky.social · 07/09/2026
Second, once we have a TPR approximation to an LLM, we can take the LLM's representations & edit them as if they are a symbolic structure, and the edits change the LLM's behavior as intended. This editability gives more fine-grained evidence that the symbolic structure is there in the LLM
120
Tom McCoy @rtommccoy.bsky.social · 07/09/2026
Why do I feel comfortable concluding Point 3? I.e., how can we conclude "equivalence" from data that merely show "approximation"? First, the way we evaluate an approximation is by seeing if it yields the same behavior as the original. This establishes a type of equivalence (behavioral equivalence)
110
Tom McCoy @rtommccoy.bsky.social · 07/09/2026
Yeah, good q! The argument is: 1. A TPR is a symbolic structure 2. Thus, if a TPR is equivalent to the LLM rep, then the LLM rep has symbolic structure 3. Our results show that the TPR is equivalent to the LLM rep Point 3 is certainly debatable, but I do feel comfortable concluding it from our data
100
Tom McCoy @rtommccoy.bsky.social · 02/09/2026
That said, it might well be the case that it's easier for the more recent architectures to discover TPR structure than GPT-2. It could be interesting to look at the training trajectory to see if TPR structure arises faster in some models than in others!
010
Tom McCoy @rtommccoy.bsky.social · 02/09/2026
Great question! GPT-2-XL approximated about as well as the architectures in that figure. So, my intuition is that the TPR structure arises more out of top-down pressures (i.e., what the representations are used for) rather than bottom-up pressures (i.e., architectural structure).
120
Tom McCoy @rtommccoy.bsky.social · 02/09/2026
Thank you for the kind words!
010
Tom McCoy @rtommccoy.bsky.social · 02/09/2026
But I do think there could be something there - will keep thinking about it!
020
Tom McCoy @rtommccoy.bsky.social · 02/09/2026
E.g., the embedding dimensionality is a free parameter, and your choice for that dimensionality might really influence whether you get linear independence. So the linear independence route might be too dependent on properties of the analysis, rather than the model being analyzed.
130
Tom McCoy @rtommccoy.bsky.social · 02/09/2026
That said, that intuition might be hard to actually implement, because the representations of roles that we extract in our approach depend somewhat on hyperparameters.
130
Tom McCoy @rtommccoy.bsky.social · 02/09/2026
Ooh I love this idea! The TPR formalism definitely makes predictions in this direction: TPRs guarantee perfect decoding only if the embeddings of roles are linearly independent, but if they're not linearly independent, you'll get intrusions.
100
Tom McCoy @rtommccoy.bsky.social · 02/09/2026
Thanks so much, Ben!
000
Tom McCoy @rtommccoy.bsky.social · 02/09/2026
Ooh, this looks excellent - will check it out! Thank you!
020
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
That sounds like an excellent course of action!
020
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
Work done with @paulsoulos.bsky.social, @tallinzen.bsky.social, and Paul Smolensky For much more, see our paper! arxiv.org/abs/2608.29530 12/12
arxiv.org
The Emergent Symbolic Structure of Artificial Neural Networks
Modern systems in artificial intelligence (AI) somehow excel in domains for which they seem poorly suited. Intelligence has traditionally been modeled as operating over structured combinations of symb...
0301
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
These results provide a potential explanation for how neural networks can be so successful in seemingly symbolic domains, potentially reconciling longstanding symbolic conceptions of intelligence with the vector-based nature of modern AI. 11/n
1272
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
In most cases, DISCOVER also generalizes out-of-distribution to filler-role combinations that it did not encounter during training. E.g., if the word "doctor" never occurred as the subject of a sentence, it can generalize to sentences with "doctor" as the subject. 10/n
Results plot showing 7 LLMs (Gemma-3, GPT-2-XL, GPT-OSS, Llama-3.1, OLMo-2, Pythia, and Qwen-3) in two domains (encoding lists or encoding sentences). The accuracy of approximating these encodings stays strong even when there are multiple unseen role-filler pairs.
1220
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
DISCOVER enables us to edit an LLM's internal representations in targeted ways that are then appropriately reflected in changes in the output. The edits that we make would typically be challenging to do in a vector encoding, but they're straightforward w/ a symbolic encoding 9/n
Edits to a coding prompt. The original prompt says to repeat the list [Q, E, V] two times, producing [Q, E, V, Q, E, V]. We can change the number of repetitions to 3, changing the output to [Q, E, V, Q, E, V, Q, E, V]. We can change the first letter from Q to H, changing the output to [H, E, V, H, E, V]. We can change the function from repeating the list 3 times to adding a B in front of the list, changing the output to [B, Q, E, V]. We can change the variable that is the function’s input to a list equal to [J, R, M], and the output changes to [J, R, M, J, R, M].
1252
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
We refer to this method as DISCOVER (short for DISsecting COmpositionality in VEctor Representations). The official mascot of DISCOVER is the tapir (pictured below), because the word TAPIR is what results from interleaving "TPR" with "AI" 8/n
An image of a tapir, which is a mammal with a long snout.
1321
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
Given appropriate choices for the type of symbols used by the TPR, the approximation produces high accuracy in all networks we consider: The network continues to produce the right answer when fed the TPR rather than its own actual encoding. 7/n
A results plot showing the results of approximating GPT-OSS’s representations with DISCOVER. Across all 6 tasks (arithmetic, syllogisms, code execution, passivization, tense reinflection, and question formation), GPT-OSS gets a similar accuracy when fed the TPR approximation as it gets when fed its own encodings.
1220
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
In the case of LLMs, this means replacing every vector that encodes part of the input, across tokens and layers. 6/n
Top: Image of normal LLM processing. The LLM receives an input and generates one vector representation for each input token at each layer. It then generates the output one word at a time, based on these representations.
Bottom: Approximating an LLM with DISCOVER. We replace every vector that encodes part of the input with a TPR approximation, and then have the LLM generate text based on these TPR vectors.
1260
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
We analyze a network by approximating its representations w/ a TPR. We then feed the TPR approximation into the network & see if it still produces the right answer - replacing the network's entire representation-generating process with a simple, closed-form TPR equation. 5/n
Visualization of the DISCOVER process. First we train the target model, which takes in an input, encodes it as a vector E, and then decodes from E to produce the output. We then train a DISCOVER model, which aims to approximate E with a TPR encoding, E_TPR. We then feed E_TPR into the decoder of the target model to see if it produces the correct answer when given E_TPR instead of the original E.
2251
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
The TPR formalism was introduced by Smolensky (1987) at the 1st NeurIPS. Shortly after, in 1990, he hypothesized that, if the field could ever create neural networks that can process language, they might do so via implicit TPR structure. Our work tests that hypothesis! 4/n
Screenshot of text from Smolensky (1990). It reads: “In the short term, at least, our learning rules and network simulators do not seem powerful enough to make network learning of linguistic representation feasible. (2) Even if such learning is feasible at some future point, we will still need to explain how the representation is done. There are two empirical reasons to believe that such explanation will require the kind of analysis begun in this paper: explanation of the computation of real neural networks has turned out to require much analysis, as mere observation has proved woefully inadequate; the same has turned out to be true even of the self-organized connectionist networks that perform computations vastly simpler than most of natural language processing
1392
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
Our hypothesis: The networks succeed on these tasks by implicitly building symbolic structure in vector space. What makes this hypothesis testable is the Tensor Product Representation (TPR) formalism, a proposal for how symbolic structure could be realized in vectors. 3/n
The structure of a Tensor Product Representation (TPR), which realizes symbolic structure in vector space. In this example, the symbolic structure is the sentence “poets help chefs.” TPRs work by pairing each element of the structure (called a filler - here, each word is a filler) with its role (here, “subject”, “verb”, or “object”) and translating that role-filler structure into a vector.
2478
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
Paper: arxiv.org/abs/2608.29530 We consider a variety of neural networks that perform seemingly-symbolic tasks, from small-scale models trained on simple list-manipulation tasks to LLMs performing tasks in math, logic, coding, and language. 2/n
arxiv.org
The Emergent Symbolic Structure of Artificial Neural Networks
Modern systems in artificial intelligence (AI) somehow excel in domains for which they seem poorly suited. Intelligence has traditionally been modeled as operating over structured combinations of symb...
2504
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n
Overview of the paper. 
Title: The Emergent Symbolic Structure of Artificial Neural Networks
Authors: Tom McCoy, Paul Soulos, Tal Linzen, Paul Smolensky
Left: Neural networks encode information in vectors (there is then an image of a vector), yet they excel at tasks long thought to require symbolic structure (there is then an image of a symbolic representation, specifically a syntax tree). How do LLMs do it?
Right: We find that LLM representations can be closely approximated with symbolic structures. This approximation lets us edit the structure of an LLM’s output by editing the structure of its internal representations, as shown. There is then an image of two edits to LLMs. In the first one, the original input is 3 + 6 * 8, with an answer of 51. But if we swap the positions of the 3 and the 6, the output becomes 30. In the second one, the original input is a Python command repeating the list [Z, U] three times, producing [Z, U, Z, U, Z, U]. But if we edit the input in a way that adds a Q at the end of the input, the output becomes [Z, U, Q, Z, U, Q, Z, U, Q].
431888
Reposted by Tom McCoy
Sean Trott @seantrott.bsky.social · 31/08/2026
New paper out in @openmindjournal.bsky.social! "Large Language Models as Distributional Baselines for Language Tasks". With @jamichaelov.bsky.social , @camrobjones.bsky.social , Tyler Chang, and Ben Bergen. direct.mit.edu/opmi/article...
direct.mit.edu
Large Language Models as Distributional Baselines for Language Tasks
Abstract. In order to ask questions about the mechanisms underpinning human cognition, researchers must control for properties of stimuli that could confound detected effects. In experiments involving...
1126
Tom McCoy @rtommccoy.bsky.social · 27/08/2026
Come be my colleague! Yale's Wu Tsai Institute is hiring for a position in neurocognition.
070
Tom McCoy @rtommccoy.bsky.social · 20/08/2026
Indeed a valid inference!
020
Tom McCoy @rtommccoy.bsky.social · 20/08/2026
Blog post where I elaborate on this: rtmccoy.com/posts/theory...
rtmccoy.com
Theory of mind is often shortened to ToM, but I'll avoid that abbreviation. Given what my first name is, I feel self-conscious writing sentences like "ToM is very important" or "You should pay more attention to ToM."
0120
Tom McCoy @rtommccoy.bsky.social · 20/08/2026
Since many are starting grad school soon, let me re-share my One Big Tip™️ for research! Research involves many skills - collaborating, writing, presenting, etc. But many of these skills can be unified under a single overarching ability: theory of mind Blog post link in reply
Illustration of the blog post's main argument, summarized as: "Theory of Mind as a Central Skill for Researchers: Research involves many skills.If each skill is viewed separately, each one takes a long time to learn. These skills can instead be connected via theory of mind – the ability to reason about the mental states of others. This allows you to transfer your abilities across areas, making it easier to gain new skills."
28019
Tom McCoy @rtommccoy.bsky.social · 11/08/2026
But you should read the whole paper! Link: repository.ubn.ru.nl/bitstream/ha...
repository.ubn.ru.nl
000
Tom McCoy @rtommccoy.bsky.social · 11/08/2026
Quote summarizing the paper: "For if rhythm is so integral a part of our audition, then it ought to be the case that it is hard to overlook; but the most pronounced of rhythms can escape our recognition when they're reproduced in printing in an article or book."
130
Tom McCoy @rtommccoy.bsky.social · 11/08/2026
Speaking of poetry hidden in plain sight, everyone should read Anne Cutler's paper on the subject. Absolutely mind-blowing. (link and quote in reply ⬇️)
161
Tom McCoy @rtommccoy.bsky.social · 10/08/2026
I looked at the edit history, and when this sentence was added to the page, it originally came with a note. The line of text plus the note form a complete double dactyl poem! At some point an editor removed the note, leaving the more subtle half-poem that remains. 2/2
Screenshot of an old version of the Wikipedia page for “Double dactyl.”
It shows a line of text followed by a note. Taking together the text and the note gives: “Metapoetically, Roger L. Robison crafted this poem describing itself. Sonnets and haikus await his analysis; metapoetics could fill a whole shelf.” – which is a well-formed double dactyl poem!
17012
Tom McCoy @rtommccoy.bsky.social · 10/08/2026
Just found an Easter egg on Wikipedia! On the page for double dactyls (a type of poem), there's a sentence that's formatted as a normal line of text, but it's actually a stanza of double dactyl poetry! And... THE PLOT THICKENS! 1/2
Screenshot of the Wikipedia page for “Double dactyl.”The text reads:
The double dactyl is a verse form invented by Anthony Hecht and Paul Pascal in 1951.

An example by John Hollander:
Higgledy piggledy,Benjamin Harrison,Twenty-third presidentWas, and, as such,Served between Clevelands andSave for this trivialIdiosyncrasy,Didn't do much.

Metapoetically, Roger L. Robison crafted this poem describing itself:
Long-short-short, long-short-shortDactyls in dimeter,Verse form with choriambs(Masculine rhyme):One sentence (two stanzas)HexasyllabicallyChallenges poets whoDon't have the time.

There’s then an annotation noting that the text “Metapoetically, Roger L. Robison crafted this poem describing itself:” works as the first stanza of a double dactyl poem.
28718
Tom McCoy @rtommccoy.bsky.social · 14/07/2026
Amazon's "Track Package" button should really be called "Package Trackage"
090
Reposted by Tom McCoy
Tom McCoy @rtommccoy.bsky.social · 05/07/2026
Seeking two ARR emergency reviewers for papers about steering. Please let me know if you'd be able to help out!
001
Tom McCoy @rtommccoy.bsky.social · 05/07/2026
Today at #acl2026nlp: Check out the panel on "Linguistics and NLP in the LLM Era", featuring Allyson Ettinger, @futrell.bsky.social, Zoey Liu, and me! 2:00 - 3:30 in Gaslamp C&D
052
Tom McCoy @rtommccoy.bsky.social · 05/07/2026
Seeking two ARR emergency reviewers for papers about steering. Please let me know if you'd be able to help out!
001
Tom McCoy @rtommccoy.bsky.social · 05/07/2026
Very meta!
1130
Reposted by Tom McCoy
Maria Antoniak @mariaa.bsky.social · 05/07/2026
I'm looking for two emergency reviewers for #EMNLP2026 papers about cultural benchmarking and cultural reasoning. Please send me an email if you would be able to help me out in the next few days!
188
Tom McCoy @rtommccoy.bsky.social · 03/07/2026
Link for "Collocational bootstrapping: A hypothesis about the learning of subject-verb agreement in humans and neural networks": aclanthology.org/2026.conll-m...
aclanthology.org
Collocational bootstrapping: A hypothesis about the learning of subject-verb agreement in humans and neural networks
Claire Hobbs, R. Thomas McCoy. Proceedings of the 30th Conference on Computational Natural Language Learning. 2026.
030
Tom McCoy @rtommccoy.bsky.social · 03/07/2026
Link for "What Exactly do Children Receive in Language Acquisition? A Case Study on CHILDES with Automated Detection of Filler-Gap Dependencies": aclanthology.org/2026.conll-m...
aclanthology.org
What Exactly do Children Receive in Language Acquisition? A Case Study on CHILDES with Automated Detection of Filler-Gap Dependencies
Zhenghao Zhou, William Dai, Maya Viswanathan, Simon Charlow, R. Thomas McCoy, Robert Frank. Proceedings of the 30th Conference on Computational Natural Language Learning. 2026.
120