Sign in

Tom McCoy

@rtommccoy.bsky.social
2.2K followers 369 following 288 posts

Assistant professor at Yale Linguistics. Studying computational linguistics, cognitive science, and AI. He/him.

PostsRepliesMedia
Tom McCoy @rtommccoy.bsky.social · 27/09/2026
A nor'easter is a storm with winds from the northeast So you might think a sou'wester would be a storm with winds from the southwest. But it's not!! It's a type of hat!! Language strikes again
An image of a sou'wester - a yellow waterproof rain hat - sourced from Wikipedia
0201
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
In most cases, DISCOVER also generalizes out-of-distribution to filler-role combinations that it did not encounter during training. E.g., if the word "doctor" never occurred as the subject of a sentence, it can generalize to sentences with "doctor" as the subject. 10/n
Results plot showing 7 LLMs (Gemma-3, GPT-2-XL, GPT-OSS, Llama-3.1, OLMo-2, Pythia, and Qwen-3) in two domains (encoding lists or encoding sentences). The accuracy of approximating these encodings stays strong even when there are multiple unseen role-filler pairs.
1220
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
DISCOVER enables us to edit an LLM's internal representations in targeted ways that are then appropriately reflected in changes in the output. The edits that we make would typically be challenging to do in a vector encoding, but they're straightforward w/ a symbolic encoding 9/n
Edits to a coding prompt. The original prompt says to repeat the list [Q, E, V] two times, producing [Q, E, V, Q, E, V]. We can change the number of repetitions to 3, changing the output to [Q, E, V, Q, E, V, Q, E, V]. We can change the first letter from Q to H, changing the output to [H, E, V, H, E, V]. We can change the function from repeating the list 3 times to adding a B in front of the list, changing the output to [B, Q, E, V]. We can change the variable that is the function’s input to a list equal to [J, R, M], and the output changes to [J, R, M, J, R, M].
1252
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
We refer to this method as DISCOVER (short for DISsecting COmpositionality in VEctor Representations). The official mascot of DISCOVER is the tapir (pictured below), because the word TAPIR is what results from interleaving "TPR" with "AI" 8/n
An image of a tapir, which is a mammal with a long snout.
1321
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
Given appropriate choices for the type of symbols used by the TPR, the approximation produces high accuracy in all networks we consider: The network continues to produce the right answer when fed the TPR rather than its own actual encoding. 7/n
A results plot showing the results of approximating GPT-OSS’s representations with DISCOVER. Across all 6 tasks (arithmetic, syllogisms, code execution, passivization, tense reinflection, and question formation), GPT-OSS gets a similar accuracy when fed the TPR approximation as it gets when fed its own encodings.
1220
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
In the case of LLMs, this means replacing every vector that encodes part of the input, across tokens and layers. 6/n
Top: Image of normal LLM processing. The LLM receives an input and generates one vector representation for each input token at each layer. It then generates the output one word at a time, based on these representations.
Bottom: Approximating an LLM with DISCOVER. We replace every vector that encodes part of the input with a TPR approximation, and then have the LLM generate text based on these TPR vectors.
1260
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
We analyze a network by approximating its representations w/ a TPR. We then feed the TPR approximation into the network & see if it still produces the right answer - replacing the network's entire representation-generating process with a simple, closed-form TPR equation. 5/n
Visualization of the DISCOVER process. First we train the target model, which takes in an input, encodes it as a vector E, and then decodes from E to produce the output. We then train a DISCOVER model, which aims to approximate E with a TPR encoding, E_TPR. We then feed E_TPR into the decoder of the target model to see if it produces the correct answer when given E_TPR instead of the original E.
2251
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
The TPR formalism was introduced by Smolensky (1987) at the 1st NeurIPS. Shortly after, in 1990, he hypothesized that, if the field could ever create neural networks that can process language, they might do so via implicit TPR structure. Our work tests that hypothesis! 4/n
Screenshot of text from Smolensky (1990). It reads: “In the short term, at least, our learning rules and network simulators do not seem powerful enough to make network learning of linguistic representation feasible. (2) Even if such learning is feasible at some future point, we will still need to explain how the representation is done. There are two empirical reasons to believe that such explanation will require the kind of analysis begun in this paper: explanation of the computation of real neural networks has turned out to require much analysis, as mere observation has proved woefully inadequate; the same has turned out to be true even of the self-organized connectionist networks that perform computations vastly simpler than most of natural language processing
1392
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
Our hypothesis: The networks succeed on these tasks by implicitly building symbolic structure in vector space. What makes this hypothesis testable is the Tensor Product Representation (TPR) formalism, a proposal for how symbolic structure could be realized in vectors. 3/n
The structure of a Tensor Product Representation (TPR), which realizes symbolic structure in vector space. In this example, the symbolic structure is the sentence “poets help chefs.” TPRs work by pairing each element of the structure (called a filler - here, each word is a filler) with its role (here, “subject”, “verb”, or “object”) and translating that role-filler structure into a vector.
2478
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n
Overview of the paper. 
Title: The Emergent Symbolic Structure of Artificial Neural Networks
Authors: Tom McCoy, Paul Soulos, Tal Linzen, Paul Smolensky
Left: Neural networks encode information in vectors (there is then an image of a vector), yet they excel at tasks long thought to require symbolic structure (there is then an image of a symbolic representation, specifically a syntax tree). How do LLMs do it?
Right: We find that LLM representations can be closely approximated with symbolic structures. This approximation lets us edit the structure of an LLM’s output by editing the structure of its internal representations, as shown. There is then an image of two edits to LLMs. In the first one, the original input is 3 + 6 * 8, with an answer of 51. But if we swap the positions of the 3 and the 6, the output becomes 30. In the second one, the original input is a Python command repeating the list [Z, U] three times, producing [Z, U, Z, U, Z, U]. But if we edit the input in a way that adds a Q at the end of the input, the output becomes [Z, U, Q, Z, U, Q, Z, U, Q].
432088
Tom McCoy @rtommccoy.bsky.social · 20/08/2026
Since many are starting grad school soon, let me re-share my One Big Tip™️ for research! Research involves many skills - collaborating, writing, presenting, etc. But many of these skills can be unified under a single overarching ability: theory of mind Blog post link in reply
Illustration of the blog post's main argument, summarized as: "Theory of Mind as a Central Skill for Researchers: Research involves many skills.If each skill is viewed separately, each one takes a long time to learn. These skills can instead be connected via theory of mind – the ability to reason about the mental states of others. This allows you to transfer your abilities across areas, making it easier to gain new skills."
28019
Tom McCoy @rtommccoy.bsky.social · 10/08/2026
I looked at the edit history, and when this sentence was added to the page, it originally came with a note. The line of text plus the note form a complete double dactyl poem! At some point an editor removed the note, leaving the more subtle half-poem that remains. 2/2
Screenshot of an old version of the Wikipedia page for “Double dactyl.”
It shows a line of text followed by a note. Taking together the text and the note gives: “Metapoetically, Roger L. Robison crafted this poem describing itself. Sonnets and haikus await his analysis; metapoetics could fill a whole shelf.” – which is a well-formed double dactyl poem!
17012
Tom McCoy @rtommccoy.bsky.social · 10/08/2026
Just found an Easter egg on Wikipedia! On the page for double dactyls (a type of poem), there's a sentence that's formatted as a normal line of text, but it's actually a stanza of double dactyl poetry! And... THE PLOT THICKENS! 1/2
Screenshot of the Wikipedia page for “Double dactyl.”The text reads:
The double dactyl is a verse form invented by Anthony Hecht and Paul Pascal in 1951.

An example by John Hollander:
Higgledy piggledy,Benjamin Harrison,Twenty-third presidentWas, and, as such,Served between Clevelands andSave for this trivialIdiosyncrasy,Didn't do much.

Metapoetically, Roger L. Robison crafted this poem describing itself:
Long-short-short, long-short-shortDactyls in dimeter,Verse form with choriambs(Masculine rhyme):One sentence (two stanzas)HexasyllabicallyChallenges poets whoDon't have the time.

There’s then an annotation noting that the text “Metapoetically, Roger L. Robison crafted this poem describing itself:” works as the first stanza of a double dactyl poem.
28718
Tom McCoy @rtommccoy.bsky.social · 03/07/2026
Today at CoNLL (paper links in thread): 1️⃣ 2:00 - 3:30 poster session (Harbor C): Zhenghao Herbert Zhou on what children's filler-gap input looks like! 2️⃣ 4:00 - 5:30 talk session (Harbor C): Claire Hobbs on how input statistics can support the learning of abstract syntactic rules!
Schedule for some presentations at CoNLL:
2:00 - 3:30 poster session in Harbor C: "What exactly do children receive in language acquisition? A case study on CHILDES with automated detection of filler-gap dependencies"
4:00 - 5:30 talk session in Harbor C: "Collocational bootstrapping: A hypothesis about the learning of subject-verb agreement in humans and neural networks"
1101
Tom McCoy @rtommccoy.bsky.social · 08/06/2026
Wrote today's NYT crossword - enjoy!
Screenshot from the New York Times crossword page that says "The Crossword. Monday, June 8, 2026. By Tom McCoy. Edited by Will Shortz."
4311
Tom McCoy @rtommccoy.bsky.social · 28/05/2026
This is a funny title once you realize that Jane Austen wrote exactly 6 books Other listicles I want to see: - "The 11 most enjoyable months" - "The 49 greatest US states" - "The 25 most useful letters of the alphabet"
Screenshot of an article title reading "The 5 Jane Austen books everyone should read for what would be her 250th birthday"
6525
Tom McCoy @rtommccoy.bsky.social · 22/05/2026
We next analyzed child-directed language & found that the level of variability in it was close to the level that optimized neural net generalization. This is evidence that child-directed language has the statistical properties that make collocational bootstrapping effective! 9/n
A Zipfian distribution fitted to child-directed language; the empirical distribution matches the theoretical one closely
100
Tom McCoy @rtommccoy.bsky.social · 22/05/2026
We find that there is indeed a sweet spot of variability where the neural networks robustly generalize subject-verb agreement! 7/n
Plot showing neural net subject-verb agreement accuracy as a function of the variability in the training data. Accuracy is optimized (and is near 1.0) when the level of variability is medium.
110
Tom McCoy @rtommccoy.bsky.social · 22/05/2026
To test this, we train neural network language models on synthetic data across many conditions varying how predictable subject-verb pairings are We do this by sampling subject-verb pairs from Zipfian distributions that vary a parameter defining how skewed the distribution is 6/n
Zipfian distributions with various values of the alpha parameter that modulates how skewed the distribution is
110
Tom McCoy @rtommccoy.bsky.social · 22/05/2026
🤖🧠NEW PAPER🧠🤖 Children & neural networks can learn syntax from linear strings of words. How do they do it? Our hypothesis: Word co-occurrence statistics provide cues to syntax! (I.e., a new type of bootstrapping to consider!) Paper: arxiv.org/abs/2605.20529 1/n
Paper overview.
Title: "Collocational bootstrapping: A hypothesis about the learning of subject-verb agreement in humans and neural networks"
Authors: Claire Hobbs and Tom McCoy
Method: We trained many neural nets, varying how predictable a subject is given its verb. We tested them on subject-verb agreement
Findings: With the right level of predictability, neural networks robustly generalize. The predictability of child-directed language is near the neural net optimum.
Conclusion: Statistical regularities in word co-occurrence can support the learning of abstract syntactic rules
The text is accompanied by a graph showing neural-network accuracy as a function of the level of variability; the accuracy peaks at an in-between level of variability
2395
Tom McCoy @rtommccoy.bsky.social · 18/05/2026
🤖🧠 New commentary 🧠🤖 What role should large language models (LLMs) play in linguistics? I reflect on this question in a commentary now on arXiv: arxiv.org/abs/2605.10061 To appear in BBS as a commentary on @futrell.bsky.social and @kmahowald.bsky.social's excellent piece on LLMs & Linguistics!
Screenshot of the title and abstract of a paper.
Title: "Not-So-Strange Love: Language Models and Generative
Linguistic Theories are More Compatible than They Appear"
Further information: "Open Peer Commentary on “How Linguistics Learned to Stop Worrying and
Love the Language Models” by Richard Futrell and Kyle Mahowald"
Author: R. Thomas McCoy
Abstract: Futrell and Mahowald (2025) frame the success of neural language models (LMs) as supporting gradient, usage-based linguistic theories. I argue that LMs can also instantiate theories based on formal structures - the types of theories seen in the generative tradition. This argument expands the space of theories that can be tested with LMs, potentially enabling reconciliations between usage-based and generative accounts.
0295
Tom McCoy @rtommccoy.bsky.social · 17/12/2025
These products flagrantly violate the rules of English compounds! "Reese's Oreo" should be a type of Oreo, not a type of Reese's!
An image of two snack foods. At the top is Reese's Oreo, which is an Oreo-flavored brand of Reese's. At the bottom is Oreo Reese's, which is a Reese's-flavored brand of Oreo.
060
Tom McCoy @rtommccoy.bsky.social · 18/11/2025
I am partial to Laffy Taffy mainly because of this one (via www.reddit.com/r/wholesomem...)
Image of a Laffy Taffy wrapper. The joke asks "What do you deserve and is also a type of bagel?" And the answer is "Everything"
130
Tom McCoy @rtommccoy.bsky.social · 14/11/2025
🤖🧠I'll be considering applications for PhD students & postdocs to start at Yale in Fall 2026! If you are interested in the intersection of linguistics, cognitive science, & AI, I encourage you to apply! PhD link: rtmccoy.com/prospective_... Postdoc link: rtmccoy.com/prospective_...
Top: A syntax tree for the sentence "the doctor by the lawyer saw the artist".

Bottom: A continuous vector.
23713
Tom McCoy @rtommccoy.bsky.social · 30/09/2025
🤖 🧠 NEW BLOG POST 🧠 🤖 What skills do you need to be a successful researcher? The list seems long: collaborating, writing, presenting, reviewing, etc But I argue that many of these skills can be unified under a single overarching ability: theory of mind rtmccoy.com/posts/theory...
Illustration of the blog post's main argument, summarized as: "Theory of Mind as a Central Skill for Researchers: Research involves many skills.If each skill is viewed separately, each one takes a long time to learn. These skills can instead be connected via theory of mind – the ability to reason about the mental states of others. This allows you to transfer your abilities across areas, making it easier to gain new skills."
2202
Tom McCoy @rtommccoy.bsky.social · 24/08/2025
A linguistic note about David Copperfield and Demon Copperhead 🧵 [very minor spoilers for both] 1/n
On the left is the cover of David Copperfield by Charles Dickens. On the right is the cover of Demon Copperhead by Barbara Kingsolver.
180
Tom McCoy @rtommccoy.bsky.social · 15/08/2025
🤖 🧠 NEW PAPER ON COGSCI & AI 🧠 🤖 Recent neural networks capture properties long thought to require symbols: compositionality, productivity, rapid learning So what role should symbols play in theories of the mind? For our answer...read on! Paper: arxiv.org/abs/2508.05776 1/n
The top shows the title and authors of the paper: "Whither symbols in the era of advanced neural networks?" by Tom Griffiths, Brenden Lake, Tom McCoy, Ellie Pavlick, and Taylor Webb.

At the bottom is text saying "Modern neural networks display capacities traditionally believed to require symbolic systems. This motivates a re-assessment of the role of symbols in cognitive theories."

In the middle is a graphic illustrating this text by showing three capacities: compositionality, productivity, and inductive biases. For each one, there is an illustration of a neural network displaying it. For compositionality, the illustration is DALL-E 3 creating an image of a teddy bear skateboarding in Times Square. For productivity, the illustration is novel words produced by GPT-2: "IKEA-ness", "nonneotropical", "Brazilianisms", "quackdom", "Smurfverse". For inductive biases, the illustration is a graph showing that a meta-learned neural network can learn formal languages from a small number of examples.
810117
Tom McCoy @rtommccoy.bsky.social · 03/08/2025
According to Jane Austen, linguists are extraordinarily cold-hearted. (Though at least we're not as bad as mathematicians!)
Picture of a paragraph from Emma by Jane Austen. It reads: "Such an adventure as this, a fine young man and a lovely young woman thrown together in such a way, could hardly fail of suggesting certain ideas to the coldest heart and the steadiest brain. So Emma thought, at least. Could a linguist, could a grammarian, could even a mathematician have seen what she did, have witnessed their appearance together, and heard their history of it, without feeling that circumstances had been at work to make them peculiarly interesting to each other? How much more must an imaginist, like herself, be on fire with speculation and foresight? especially with such a groundwork of anticipation as her mind had already made."
070
Tom McCoy @rtommccoy.bsky.social · 20/05/2025
More dramatically, it substantially outperforms the standard neural network at learning recursion (left) and priming (right; a lower value on the y-axis shows a greater degree of priming). 13/n
Left: A plot showing recursion results for standard and prior-trained neural networks. The x-axis shows levels of recursion ranging from 0 to 10, and the y-axis shows accuracy. As the levels of recursion increase, the accuracy drops for both models, but it drops much more rapidly for the standard model than the prior-trained model.
Right: A plot showing priming results. There are 4 sub-plots, for 4 types of sentences: short plausible, long plausible, short implausible, and long implausible. In all 4 plots, the prior-trained network shows a greater degree of priming than the standard neural network.
120
Tom McCoy @rtommccoy.bsky.social · 20/05/2025
Here, its perplexity is slightly better (i.e., lower) than that of a standard neural network. 12/n
A plot showing perplexity values. Note that for perplexity, lower is better. A standard neural network achieves perplexity ranging from about 19.70 to 19.80, with a median around 19.75. A prior-trained neural network achieves perplexity ranging from about 19.63 to 19.74, with a median around 19.67. The best model from prior literature is indicated as having a perplexity of about 19.69.
110
Tom McCoy @rtommccoy.bsky.social · 20/05/2025
Even though it is a neural network, the prior-trained model can learn formal languages from small numbers of examples - far outperforming a standard neural network, and matching a Bayesian model at a fraction of the computational cost. 10/n
Plots showing results for formal languages. On the left is a line graph which has “number of training examples” as its x-axis and “F-score” as its y-axis. Three models have lines in this plot: a Bayesian model, a standard neural network, and a prior-trained neural network. The Bayesian model and prior-trained neural network perform similarly, while the standard neural network does much worse than both of them.
On the right is a table showing the amount of training time used by each approach. The Bayesian model uses from 1 minute to 7 days of training time. The neural networks (whether standard or prior-trained) use from 10 milliseconds to 2.5 minutes.
130
Tom McCoy @rtommccoy.bsky.social · 20/05/2025
Inspired by a model from Yang & @spiantado.bsky.social , the prior that we use is a distribution over formal languages (a formal language = a set of strings defined by an abstract rule). We have a neural network meta-learn by observing many formal languages sampled from this prior 8/n
Two examples of formal languages. The first example shown is AnBn, described as “n copies of A followed by n copies of B”, with some example strings from the formal language being AB, AABB, and AAABBB.
The second example shown is XXX, described as “any string X repeated three times”, with some example strings from the formal language being AAA, BABABA, and ABBABBABB
110
Tom McCoy @rtommccoy.bsky.social · 20/05/2025
In MAML, a model is exposed to many tasks. After each task, the model's weights are adjusted so that, if it were taught the same task again, it would perform better. As MAML proceeds, the model converges to a state from which it can learn any task in the distribution. 7/n
110
Tom McCoy @rtommccoy.bsky.social · 20/05/2025
Our approach (inductive bias distillation) has 3 steps: 1. Use a Bayesian model to define an inductive bias (a prior) 2. Sample learning tasks from the Bayesian model 3. Have a neural network meta-learn from these sampled tasks, to give it the Bayesian model's prior 5/n
A schematic diagram of our procedure. We start with a Bayesian model, here visualized with Bayes’ rule and some example grammatical rules that could be sampled from a Bayesian model’s prior. Then, we sample several tasks from that Bayesian model’s prior, which can serve as training data. Finally, we have a neural network meta-learn from these sampled tasks. The whole process is visualized, going from left to right as “Bayesian model’, then an arrow labeled “sampling”, then “training data”, then an arrow labeled “meta-learning”, and finally “neural network.”
130
Tom McCoy @rtommccoy.bsky.social · 20/05/2025
Neural networks have flexible representations that allow them to handle noisy natural data - as evidenced by the success of large language models. However, they notoriously require huge numbers of examples. 3/n
Left: A screenshot of ChatGPT describing itself as an AI language model developed by OpenAI. Right: A bar chart comparing the quantity of text seen by human children vs. GPT-3. The bar for GPT-3 is far higher than for humans, showing that neural networks get far more linguistic data than humans do.
110
Tom McCoy @rtommccoy.bsky.social · 20/05/2025
Bayesian models can learn from few examples because they have strong inductive biases - factors that guide generalization. But the costs of inference and the difficulty of specifying generative models can make naturalistic data a challenge. 2/n
Screenshot of a demo of Bayesian word learning (Xu & Tenenbaum 2007). After a few examples, the Bayesian learner figures out that naysayer means “horse” (rather than being more specific – “horse number 4” – or more general – “mammal”).
120
Tom McCoy @rtommccoy.bsky.social · 20/05/2025
🤖🧠 Paper out in Nature Communications! 🧠🤖 Bayesian models can learn rapidly. Neural networks can handle messy, naturalistic data. How can we combine these strengths? Our answer: Use meta-learning to distill Bayesian priors into a neural network! www.nature.com/articles/s41... 1/n
A schematic of our method. On the left are shown Bayesian inference (visualized using Bayes’ rule and a portrait of the Reverend Bayes) and neural networks (visualized as a weight matrix). Then, an arrow labeled “meta-learning” combines Bayesian inference and neural networks into a “prior-trained neural network”, described as a neural network that has the priors of a Bayesian model – visualized as the same portrait of Reverend Bayes but made out of numbers. Finally, an arrow labeled “learning” goes from the prior-trained neural network to two examples of what it can learn: formal languages (visualized with a finite-state automaton) and aspects of English syntax (visualized with a parse tree for the sentence “colorless green ideas sleep furiously”).
415444
Tom McCoy @rtommccoy.bsky.social · 07/05/2025
I constructed today's NYT crossword! This one has some personal connections, described at the WordPlay article by @samcorbin.bsky.social (contains spoilers): www.nytimes.com/2025/05/06/c... I hope you enjoy!
Screenshot of the New York Times crossword page, saying "The Crossword. Wednesday, May 7, 2025. By Tom McCoy. Edited by Will Shortz."
3211
Tom McCoy @rtommccoy.bsky.social · 02/05/2025
Made a new assignment for a class on Computational Psycholinguistics: - I trained a Transformer language model on sentences sampled from a PCFG - The students' task: Given the Transformer, try to infer the PCFG (w/ a leaderboard for who got closest) Would recommend! 1/n
On the left is a probabilistic context free grammar (PCFG). On the right is an image of the Transformer architecture. There are arrows going back and forth between the PCFG and the Transformer, showing how the assignment goes back and forth between them.
1223
Tom McCoy @rtommccoy.bsky.social · 28/01/2025
I added a new assignment to my Computational Linguistics class last semester: - Choose a linguistic phenomenon in a language other than English - Give a 3-minute presentation about that phenomenon & how it would pose a challenge for computational models Would recommend! 1/n
A list of the languages that students presented on: Ancient Greek, Arabic, American Sign Language, Bahasa Indonesia, Bengali, Cantonese, Dutch, Esperanto, French, Georgian, German, Great Andamanese, Hindi, Inuktitut, Italian, Japanese, Korean, Maltese, Mandarin, Mongolian, Nahuatl, Nepali, Romanian, Sanskrit, Spanish, Swedish, Tamil, Telugu, Tshiluba, Unangam Tunuu, Welsh, Wolof, Yiddish
2414
Tom McCoy @rtommccoy.bsky.social · 20/01/2025
Like previous models, o1-preview shows clear effects of output probability (performing better when the correct answer is a high-probability string). Interestingly, the effects don’t just show up in accuracy (top) but also in how many tokens o1 consumes to perform the task (bottom)! 3/4
Top plot: o1 scores better on several tasks when the output log probability is high than when it is low.

Bottom plot: o1 uses more tokens when the output is low probability than when it is high probability.
100
Tom McCoy @rtommccoy.bsky.social · 20/01/2025
o1 shows especially big improvements over previous models when performing rare versions of tasks (left plot). But, when the tasks are hard enough, it still does better on common task variants than rare ones (right two plots) 2/4
Plots showing OpenAI o1's performance on common and rare versions of tasks. In general, on our basic evaluations, o1 is close to 100% accuracy on both common and rare task variants. However, on harder evaluations, it shows some separation, with stronger performance on common task variants than rare ones.
100
Tom McCoy @rtommccoy.bsky.social · 20/01/2025
🔥While LLM reasoning is on people's minds... Here's a shameless plug for our work comparing o1 to previous LLMs (extending "Embers of Autoregression"): arxiv.org/abs/2410.01792 - o1 shows big improvements over GPT-4 - But qualitatively it is still sensitive to probability 1/4
A plot showing LLM performance on various algorithmic tasks. For all LLMs evaluated, including o1-preview, performance is highly influenced by the probability of the output to be produced, with lower performance on cases with low-probability outputs. The tasks being evaluated on are shift ciphers, Pig Latin, article swapping, and reversal.
1285
Tom McCoy @rtommccoy.bsky.social · 29/12/2024
An image of a table with four columns. The first three columns are passages from translations of the Odyssey (by T.E. Lawrence, Robert Fagles, and Emily Wilson). The last column is labeled "Honda" and shows a passage from the Honda Odyssey owner's manual.
090
Tom McCoy @rtommccoy.bsky.social · 04/12/2024
Excited to be visiting the Simons Institute tomorrow for a debate with Sébastien Bubeck - billed as the Sparks/Embers debate! 🔥🤖🧠 Topic: Will scaling current LLMs be sufficient to resolve major open math conjectures?
Images of two paper titles: "Sparks of Artificial General Intelligence" and "Embers of Autoregression"
3341
Tom McCoy @rtommccoy.bsky.social · 14/11/2024
🤖🧠 I'll be considering applications for postdocs & PhD students to start at Yale in Fall 2025! If you are interested in the intersection of linguistics, cognitive science, and AI, I encourage you to apply! Postdoc link: rtmccoy.com/prospective_... PhD link: rtmccoy.com/prospective_...
Top: syntax tree for the sentence "the doctor by the lawyer saw the artist"
Bottom: a continuous vector
0145
Tom McCoy @rtommccoy.bsky.social · 10/11/2024
"Worldpay" - anagram of "Wordplay"!
A credit card reader displaying the brand name "worldpay"
050