Sign in

Jackson Petty

@jacksonpetty.org
212 followers 254 following 313 posts

the passionate shepherd, to his love • ἀρετῇ • מנא הני מילי

PostsRepliesMedia
Jackson Petty @jacksonpetty.org · 05/10/2025
Ad infinitum
0297
Jackson Petty @jacksonpetty.org · 09/06/2025
So, what did we learn? 1. LLMs *do* know how to follow instructions, but they often don’t 2. The complexity of instructions and examples reliably predicts whether (current) models can solve the task 3. On hard tasks, models (and people, tbh) like to fall back to heuristics
110
Jackson Petty @jacksonpetty.org · 09/06/2025
But often models get distracted by irrelevant info, or “get lazy” and choose to rely on heuristics rather than actually verifying the instructions; we use o4-mini as an LLM judge to classify model strategies: as examples get more complex, models shift to relying on heuristics rather than rules:
110
Jackson Petty @jacksonpetty.org · 09/06/2025
So, how can LLMs succeed at this task, and why do they fail when grammars and examples get complex? Well, models in general do understand the general solution: even small models recognize they can build a CYK table or do an exhaustive top-down search of the derivation tree:
100
Jackson Petty @jacksonpetty.org · 09/06/2025
In general, we find that models tend to agree with one another on which grammars (left) and which examples (right) are hard, though again 4.1-nano and 4.1-mini pattern with each other against others. These correlations increase with complexity!
100
Jackson Petty @jacksonpetty.org · 09/06/2025
Interestingly, models’ accuracies are reflective of divergent class biases: 4.1-nano and 4.1-mini love to predict strings as being positive, while all other models have the opposite bias; these biases also change with example complexity!
100
Jackson Petty @jacksonpetty.org · 09/06/2025
What do we find? All models struggle on complex instruction sets (grammars) and tasks (strings); the best reasoning models are better than the rest, but still approach ~chance accuracy when grammars (top) have ~500 rules, or when strings (bottom) have >25 symbols.
100
Jackson Petty @jacksonpetty.org · 09/06/2025
LLMs are increasingly used to solve tasks “zero-shot,” with only a specification of the task given in a prompt. To evaluate LLMs on increasingly complex instructions, we turn to a classic problem in computer science and linguistics: recognizing if a formal grammar generates a given string.
110
Jackson Petty @jacksonpetty.org · 09/06/2025
How well can LLMs understand tasks with complex sets of instructions? We investigate through the lens of RELIC: REcognizing (formal) Languages In-Context, finding a significant overhang between what LLMs are able to do theoretically and how well they put this into practice.
152
Jackson Petty @jacksonpetty.org · 02/01/2025
Eighth night // Dedication of the House (Shimen Frug)
001
Jackson Petty @jacksonpetty.org · 01/01/2025
Seventh night // Dedication of the House (Shimen Frug)
101
Jackson Petty @jacksonpetty.org · 31/12/2024
Sixth Night
110
Jackson Petty @jacksonpetty.org · 30/12/2024
Total Maccabean Victory
120
Jackson Petty @jacksonpetty.org · 29/12/2024
No one who speaks Afrikaans could be a bad clown
010
Jackson Petty @jacksonpetty.org · 29/12/2024
100
Jackson Petty @jacksonpetty.org · 29/12/2024
Happy Hanukkah
111
Jackson Petty @jacksonpetty.org · 28/12/2024
shabbat shalom
101
Jackson Petty @jacksonpetty.org · 27/12/2024
בימים ההם בזמן הזה
101
Jackson Petty @jacksonpetty.org · 26/12/2024
Happy Hanukkah
130
Jackson Petty @jacksonpetty.org · 25/12/2024
010
Jackson Petty @jacksonpetty.org · 21/12/2024
I think 2m43s is a new personal record
010
Jackson Petty @jacksonpetty.org · 19/12/2024
the plan to flip the house in 2026? we get all bluesky-curious voters on the platform, fill it with fake verified accounts, and heaven-ban them from ever interacting with the actual politicians
020
Jackson Petty @jacksonpetty.org · 03/12/2024
Of all sad words of tongue or pen, the saddest are these:
130
Jackson Petty @jacksonpetty.org · 28/11/2024
big “Gashleycrumb Tinies” energy
010
Jackson Petty @jacksonpetty.org · 28/11/2024
girlfriend tax implies the existence of the girlfriend laugher curve, which is the theoretical relationship between how bad my jokes are and how much she laughs at them
000
Jackson Petty @jacksonpetty.org · 28/11/2024
1) what
010
Jackson Petty @jacksonpetty.org · 21/11/2024
so apparently the Ancient Greek verb for “to hurt” is just “aaaaaooooow”
1194
Jackson Petty @jacksonpetty.org · 18/11/2024
In case it’s of use to other academics, I put together an overleaf project which shows how to use different styles for in-text citations and bibliographies (eg, author-year in text vs numeric in bib). Works in BibLaTeX and hackily in natbib. Link: www.overleaf.com/read/tkxbfty...
120
Jackson Petty @jacksonpetty.org · 18/11/2024
Pingali and Bilardi (2015) just get me
161
Jackson Petty @jacksonpetty.org · 17/11/2024
some pictures of the surrounding roses
100
Jackson Petty @jacksonpetty.org · 16/11/2024
Scenes from Miami
010
Jackson Petty @jacksonpetty.org · 16/11/2024
100
Jackson Petty @jacksonpetty.org · 14/11/2024
140
Jackson Petty @jacksonpetty.org · 14/11/2024
020
Jackson Petty @jacksonpetty.org · 14/11/2024
November in New Haven
A red-tailed hawk sitting atop a post across from SSSMacchiato at Crêpes Coupette
000
Jackson Petty @jacksonpetty.org · 10/11/2024
Are you coming to #EMNLP2024? Come say hi! I’ll be presenting this work at BlackboxNLP on Friday, and will be around all week to chat! DM to meet up, let’s grab lunch!
070
Jackson Petty @jacksonpetty.org · 13/11/2023
Even though BERTs are good at Type 1 generalization, they're much worse at Type 2 generalization: when we introduce novel verbs, BERTs fail to make correct generalizations about their argument structure!
100
Jackson Petty @jacksonpetty.org · 13/11/2023
But humans also have a stronger competency: generalization between related contexts, like learning how “passives” or “clefts” work independent of specific nouns or verbs. We call this “Type 2” knowledge.
100
Jackson Petty @jacksonpetty.org · 13/11/2023
BERTs do well at this task, echoing previous work from @najoung.bsky.social & Paul Smolensky, as well as our previous paper (Petty, Wilson, & Frank 2022):
100
Jackson Petty @jacksonpetty.org · 13/11/2023
We show that the BERT family of LMs is quite good at Type 1 generalization by introducing novel nouns and fine-tuning LMs on sentences which use them in only active contexts and testing them on novel contexts!
100
Jackson Petty @jacksonpetty.org · 13/11/2023
But humans know much more than this–we can infer that if a word appears in context C then it can appear in a related context C’. If you know that “ball” can be the object of “kick,” then you know it can also be the subject of passive “be kicked.” We call this “Type 1” knowledge.
100
Jackson Petty @jacksonpetty.org · 13/11/2023
What do LLMs learn about language? Well, their explicit task is to learn a conditional probability distribution over words. We call this “Type 0” knowledge, and LLMs are quite good at it!
100
Jackson Petty @jacksonpetty.org · 10/11/2023
An addendum: in language modeling, how come our deepest models aren't the best? We couple depth with _only_ the feedforward size, keeping the attention mechanism unchanged across depths. This means the feed-forward block becomes a lossy transformation when d_ff < d_model!
000
Jackson Petty @jacksonpetty.org · 10/11/2023
Nope! We can correct for the LM performance by choosing our fine-tuning checkpoints to have equal perplexity across depths. Even when we do this, deeper models still generalize better!
100
Jackson Petty @jacksonpetty.org · 10/11/2023
But our main interest is in compositional generalization. We fine-tune our pretrained models on four different compositional generalization dataset (COGS, COGS-vf, GeoQuery, and English Passivization). Here too, depth helps but the marginal utility diminishes rapidly!
110
Jackson Petty @jacksonpetty.org · 10/11/2023
To understand how depth matters, we examine two things: first, we pretrain our models as causal LMs and look at perplexity. We find that deeper models (mostly!) are better, but that most of the benefit comes in having just a few layers!
100
Jackson Petty @jacksonpetty.org · 10/11/2023
Our approach is to control for model size, making deeper models narrow and shallow models wide by coupling the # of layers to the size of the feed-forward dimension! We start with three off-the-shelf models (41M, 134M, 374M parameters) and create deep and shallow variants.
100
Jackson Petty @jacksonpetty.org · 27/06/2023
Thanks Google
010
Jackson Petty @jacksonpetty.org · 12/05/2023
I do not mean to pry…
000
Jackson Petty @jacksonpetty.org · 11/05/2023
Just need the right truck
000