Sign in

Jackson Petty

@jacksonpetty.org
212 followers 254 following 313 posts

the passionate shepherd, to his love • ἀρετῇ • מנא הני מילי

PostsRepliesMedia
Reposted by Jackson Petty
NYU Center for Data Science @nyudatascience.bsky.social · 12/06/2026
Can a grammar book replace training data for AI translation? NYU Linguistics PhD student Jackson Petty (@jacksonpetty.org), CDS Assoc. Prof. Tal Linzen (@tallinzen.bsky.social), and CDS MS student Jaulie Goe tested this with synthetic languages. nyudatascience.medium.com/can-ai-learn...
nyudatascience.medium.com
Can AI Learn a Language from a Textbook?
Most of the world’s 7,000 languages will never have enough written text online to train an AI translation system. However, linguists have…
083
Jackson Petty @jacksonpetty.org · 13/12/2025
Zero as in “001” (the leading digits are silent)
020
Reposted by Jackson Petty
Ariel Edwards-Levy @aedwardslevy.bsky.social · 10/12/2025
"there's a new serif in town"
1005146872491
Reposted by Jackson Petty
Greg Durrett @gregdnlp.bsky.social · 02/12/2025
📢 Postdoc position 📢 I’m recruiting a postdoc for my lab at NYU! Topics include LM reasoning, creativity, limitations of scaling, AI for science, & more! Apply by Feb 1. (Different from NYU Faculty Fellows, which are also great but less connected to my lab.) Link in 🧵
22112
Jackson Petty @jacksonpetty.org · 29/10/2025
Surely a third account would help to clarify the matter
280
Jackson Petty @jacksonpetty.org · 05/10/2025
lmao were you also on the 12:17 to GC?
100
Jackson Petty @jacksonpetty.org · 05/10/2025
Ad infinitum
0297
Reposted by Jackson Petty
NYU Center for Data Science @nyudatascience.bsky.social · 10/09/2025
Linguistics PhD student @jacksonpetty.org finds LLMs "quiet-quit" when instructions get long, switching from reasoning to guesswork. With CDS' @tallinzen.bsky.social, @shauli.bsky.social, @lambdaviking.bsky.social, @michahu.bsky.social, and Wentao Wang. nyudatascience.medium.com/llms-switch-...
nyudatascience.medium.com
LLMs Switch to Guesswork Once Instructions Get Long
LLMs abandon reasoning for guesswork when instructions get long, new work from Linguistics PhD student Jackson Petty & CDS shows.
072
Jackson Petty @jacksonpetty.org · 04/07/2025
you’re telling me a star spangled this banner??
010
Reposted by Jackson Petty
Tal Linzen @tallinzen.bsky.social · 21/06/2025
I'm hiring at least one post-doc! We're interested in creating language models that process language more like humans than mainstream LLMs do, through architectural modifications and interpretability-style steering. Express interest here: docs.google.com/forms/d/e/1F...
docs.google.com
NYU LLM + cognitive science post-doc interest form
Tal Linzen's group at NYU is hiring a post-doc! We're interested in creating language models that process language more like humans than mainstream LLMs do, through architectural modifications and int...
24221
Jackson Petty @jacksonpetty.org · 22/06/2025
Kauaʻi is amazing
020
Jackson Petty @jacksonpetty.org · 09/06/2025
Thanks to my wonderful co-authors: @michahu.bsky.social, Wentao Wang, @shauli.bsky.social, @lambdaviking.bsky.social, and @tallinzen.bsky.social! Paper, dataset, and code at jacksonpetty.org/relic/
030
Jackson Petty @jacksonpetty.org · 09/06/2025
Eg: general direction following, or translation of *natural* languages based only on non-formal reference grammars. Our results here show that there is no a priori roadblock to success, but that there are overhangs between what models can do and what they actually do.
120
Jackson Petty @jacksonpetty.org · 09/06/2025
2. It’s natural to ask “well, why not just break out to tool use? Parsers can solve this task trivially.” That’s true! But I think it’s valuable to understand how formally-verifiable tasks can shed light on model behavior on tasks for which aren’t formally verifiable.
100
Jackson Petty @jacksonpetty.org · 09/06/2025
This is contrary to the view that failure means “LLMs can’t reason”—failure here is likely correctable, and hopefully will make models more robust!
100
Jackson Petty @jacksonpetty.org · 09/06/2025
Why is this important? Well, two main reasons: 1. The overhang between models’ knowledge of *how* to solve the task and their ability to follow through gives me hope that we produce models which are better at following complex instructions in-context.
100
Jackson Petty @jacksonpetty.org · 09/06/2025
So, what did we learn? 1. LLMs *do* know how to follow instructions, but they often don’t 2. The complexity of instructions and examples reliably predicts whether (current) models can solve the task 3. On hard tasks, models (and people, tbh) like to fall back to heuristics
110
Jackson Petty @jacksonpetty.org · 09/06/2025
But often models get distracted by irrelevant info, or “get lazy” and choose to rely on heuristics rather than actually verifying the instructions; we use o4-mini as an LLM judge to classify model strategies: as examples get more complex, models shift to relying on heuristics rather than rules:
110
Jackson Petty @jacksonpetty.org · 09/06/2025
So, how can LLMs succeed at this task, and why do they fail when grammars and examples get complex? Well, models in general do understand the general solution: even small models recognize they can build a CYK table or do an exhaustive top-down search of the derivation tree:
100
Jackson Petty @jacksonpetty.org · 09/06/2025
In general, we find that models tend to agree with one another on which grammars (left) and which examples (right) are hard, though again 4.1-nano and 4.1-mini pattern with each other against others. These correlations increase with complexity!
100
Jackson Petty @jacksonpetty.org · 09/06/2025
Interestingly, models’ accuracies are reflective of divergent class biases: 4.1-nano and 4.1-mini love to predict strings as being positive, while all other models have the opposite bias; these biases also change with example complexity!
100
Jackson Petty @jacksonpetty.org · 09/06/2025
What do we find? All models struggle on complex instruction sets (grammars) and tasks (strings); the best reasoning models are better than the rest, but still approach ~chance accuracy when grammars (top) have ~500 rules, or when strings (bottom) have >25 symbols.
100
Jackson Petty @jacksonpetty.org · 09/06/2025
We release the static dataset used in our evals as RELIC-500, where the grammar complexity is capped at 500 rules.
100
Jackson Petty @jacksonpetty.org · 09/06/2025
We introduce RELIC as an LLM evaluation: 1. generate a CFG of a given complexity; 2. sample positive (parses) and negative (doesn’t parse) strings from the grammar’s terminal symbols; 3. prompt the LLM with a (grammar, sample) pair and ask it to classify if the grammar generates the given string
100
Jackson Petty @jacksonpetty.org · 09/06/2025
As an analogue for instruction sets and tasks, formal grammars have some really nice properties: they can be made arbitrarily complex, we can sample new ones easily (avoiding problems with dataset contamination), and we can verify a model’s accuracy using formal tools (parsers).
100
Jackson Petty @jacksonpetty.org · 09/06/2025
LLMs are increasingly used to solve tasks “zero-shot,” with only a specification of the task given in a prompt. To evaluate LLMs on increasingly complex instructions, we turn to a classic problem in computer science and linguistics: recognizing if a formal grammar generates a given string.
110
Jackson Petty @jacksonpetty.org · 09/06/2025
Code, dataset, and paper at jacksonpetty.org/relic/
100
Jackson Petty @jacksonpetty.org · 09/06/2025
How well can LLMs understand tasks with complex sets of instructions? We investigate through the lens of RELIC: REcognizing (formal) Languages In-Context, finding a significant overhang between what LLMs are able to do theoretically and how well they put this into practice.
152
Jackson Petty @jacksonpetty.org · 23/05/2025
Such a shame that Apple doesn’t have much cash on hand for such expenditures
010
Jackson Petty @jacksonpetty.org · 11/05/2025
(not that this would _replace_ the scraped data in the near or medium term, but it probably would curry favor with public sentiment)
040
Jackson Petty @jacksonpetty.org · 11/05/2025
I’m surprised that some AI lab hasn’t tried to get some good PR by throwing money at artists / writers / etc to create private training distributions. Apple isn’t really leading in AI but this is the kind of thing I would have expected them to do.
120
Jackson Petty @jacksonpetty.org · 05/03/2025
"deserve" is probably the wrong framing, though perhaps the necessary one given public opinion on transit in the US. But a better calculus would be to weigh the opportunity cost of _not_ having intra- or inter-city transit: decreased income & property tax revenues, increased infrastructure $...
010
Jackson Petty @jacksonpetty.org · 27/02/2025
of all sad words of tongue or pen, the saddest are these: `! Extra }, or forgotten \endgroup.`
110
Jackson Petty @jacksonpetty.org · 20/02/2025
[guy who only watched the pilot episode of Caprica voice] Nice!
010
Jackson Petty @jacksonpetty.org · 20/02/2025
big “Aside from that, Mr. Mund, how was the museum?” energy
120
Reposted by Jackson Petty
Kyle Mahowald @kmahowald.bsky.social · 04/02/2025
Looking forward to speaking tomorrow (Tues am) in this Simons workshop in Berkeley simons.berkeley.edu/workshops/ll.... Will talk about some empirical work and also share some takes from this recent preprint from me and @futrell.bsky.social arxiv.org/abs/2501.17047
simons.berkeley.edu
LLMs, Cognitive Science, Linguistics, and Neuroscience
At a conceptual level, LLMs profoundly change the landscape for theories of human language, of the brain and computation, and of the nature of human intelligence. In linguistics, they provide a new wa...
0102
Jackson Petty @jacksonpetty.org · 19/01/2025
@foodnoms.com Can I submit a feature request on bsky? It would be amazing if we could tie Goals to bodyweight, as read in from Apple Health data. Many recommendations for daily protein intake, for instance, are given in grams/kg of bodyweight.
110
Jackson Petty @jacksonpetty.org · 15/01/2025
have you tried finding true contentment and fulfillment in life?
010
Reposted by Jackson Petty
Jackson Petty @jacksonpetty.org · 02/01/2025
Eighth night // Dedication of the House (Shimen Frug)
001
Jackson Petty @jacksonpetty.org · 02/01/2025
Eighth night // Dedication of the House (Shimen Frug)
001
Reposted by Jackson Petty
Jackson Petty @jacksonpetty.org · 01/01/2025
Seventh night // Dedication of the House (Shimen Frug)
101
Jackson Petty @jacksonpetty.org · 01/01/2025
Seventh night // Dedication of the House (Shimen Frug)
101
Jackson Petty @jacksonpetty.org · 31/12/2024
Sixth Night
110
Jackson Petty @jacksonpetty.org · 30/12/2024
Much of the recent advancement in “reasoning” comes in domains with formal verification (math, physics, etc) where it’s relatively easy to measure performance. (Heck we even saw this in the spread btw physics and bio on GPQA.) “good fiction” doesn’t work like that.
010
Jackson Petty @jacksonpetty.org · 30/12/2024
Total Maccabean Victory
120
Jackson Petty @jacksonpetty.org · 29/12/2024
unclear if it can really be adjudicated without trump on the ballot, and all times he’s been have been…odd. And this was the last round with him on it. Plus it seems like his appeal is really focused around him personally, not appeals to the right in general.
110
Jackson Petty @jacksonpetty.org · 29/12/2024
say more
000
Jackson Petty @jacksonpetty.org · 29/12/2024
good take
010
Jackson Petty @jacksonpetty.org · 29/12/2024
effective electorally? maybe but i think the jury is still out on this one (maybe with a mistrial incoming)
110
Jackson Petty @jacksonpetty.org · 29/12/2024
No one who speaks Afrikaans could be a bad clown
010