Sign in

Ari

@ari-holtzman.bsky.social
508 followers 768 following 199 posts

Assistant Professor @ UChicago CS & DSI UChicao Leading Conceptualization Lab conceptualization.ai Minting new vocabulary to conceptualize generative models.

PostsRepliesMedia
Ari @ari-holtzman.bsky.social · 13/07/2026
the sweetest thing about me is that there's nothing I love more than seeing people succeeding even if I don't get it and don't like them. hell yeah, dawg, get it
020
Ari @ari-holtzman.bsky.social · 11/07/2026
One hope I have for the 21st century is to restore people's respect for the unexamined life.
010
Ari @ari-holtzman.bsky.social · 10/07/2026
LLMs don't perfectly remember the documents in their training set. But has anyone characterized how their reproductions differ from the original?
020
Ari @ari-holtzman.bsky.social · 27/06/2026
call me same-day delivery the way I be expensive and lyin'
030
Ari @ari-holtzman.bsky.social · 16/06/2026
An LLM fintuned on a book does not act like a human who read a book; it doesn't have consistent episodic memories about the experience, etc. What does the finetuning corpus that produces an LLM that acts like a human who read book look like?
031
Ari @ari-holtzman.bsky.social · 12/06/2026
There's one very real way that evolution is happening in LLMs: through the selection of attributes that are heritable and selected for in synthetic data. Have we isolated any of these yet?
020
Ari @ari-holtzman.bsky.social · 09/06/2026
Does the ICL-finetuning gap close with scale?
000
Ari @ari-holtzman.bsky.social · 06/06/2026
Nabakov absolute would have use toprified and y'all would have loved it
000
Ari @ari-holtzman.bsky.social · 06/06/2026
swashbuckling happens when the things you care about are dynamic and you're playing around with them directly enough that you can break them if you're not careful
000
Ari @ari-holtzman.bsky.social · 02/06/2026
For a given LLM there must be certain data drawn from distribution D, such that if you finetune on them the LLM performs worse on D. It misgeneralizes, due to its priors, as we all do. Are there any interesting cases of this that don't feel totally adversarial and artificial?
010
Ari @ari-holtzman.bsky.social · 30/05/2026
play nice games win nice prizes
000
Ari @ari-holtzman.bsky.social · 29/05/2026
People often belittle their sentimentality: "Oh I just like this album because I was a teen in the 80's" as if everything was just association games. Tell me what's 🪄magical🎩 rather than what's arbitrary if you're going to bother telling me anything at all.
000
Ari @ari-holtzman.bsky.social · 17/05/2026
000
Ari @ari-holtzman.bsky.social · 13/05/2026
I'm quite surprised none of the major LLM deployers allow you to emoji react to LLM messages, seems like an obvious feedback channel to exploit
030
Ari @ari-holtzman.bsky.social · 11/05/2026
020
Ari @ari-holtzman.bsky.social · 10/05/2026
hypothesis: llms can't do real collaborative friction without a strong persona. identity is what tells you what to push back on. we are empirically discovering what kinds of conversation require someone to reveal something about themselves to be productive.
120
Ari @ari-holtzman.bsky.social · 05/05/2026
still waiting for the LLM-backed Factorio 2. It'll be exactly like Factorio, but now you have to deal with other people to get shit done.
010
Ari @ari-holtzman.bsky.social · 02/05/2026
when people talk about 'there only N kinds of stories' I feel like I'm listening to someone give a lecture that goes 'there are many different sentences but ultimately all of them are pretty much 'A, B, C, D, E, F, G, H, I, J, K, L, M, N, O, R, P, Q, S, T, U, V, W, X, Y, or Z'
010
Ari @ari-holtzman.bsky.social · 01/05/2026
Does model collapse happen if you train a current LLM on data from Talkie?
000
Ari @ari-holtzman.bsky.social · 30/04/2026
I’m always suspicious of smart people who are very convincing about things smart people would like to be true
010
Ari @ari-holtzman.bsky.social · 30/04/2026
people's main advantage over LLMs is that they don't constrain themselves to the prior distribution, but LLMs' main advantange over people is that they do
000
Ari @ari-holtzman.bsky.social · 29/04/2026
do you think ornithologists would be able to make a better kind of bird with more resources?
000
Ari @ari-holtzman.bsky.social · 21/04/2026
I wonder how many OpenClaw agents are running for users that have already passed away
010
Ari @ari-holtzman.bsky.social · 16/04/2026
I think there's truly very little free lunch to be had in how much you memorize vs. generalize—it just depends on the environment, which will naturally select what magnitude of adaptability and what level of abstraction you should memorize at
001
Ari @ari-holtzman.bsky.social · 16/04/2026
every species has a minimum viable population—below some threshold it just can't sustain itself. what's the MVP for an LLM ecosystem where models write data and future models train on it? is there one, or does it always collapse? and what should count as a distinct individual?
000
Ari @ari-holtzman.bsky.social · 16/04/2026
The one way I think current LLMs are noticeably simulacra: they don't appear to actively optimize for goals. They perform the kind of actions someone optimizing for a goal would make, and often that's enough to succeed. Maybe it's just a blip in LLM progress...but maybe not.
120
Ari @ari-holtzman.bsky.social · 13/04/2026
prediction: AI media will bring back true suspense. currently, you can't feel real uncertainty b/c you find stories through channels that telegraph the outcome. an AI has nothing to lose; it'll kill your protagonist 90% through if it makes you reckon with something. I'm pumped.
000
Ari @ari-holtzman.bsky.social · 10/04/2026
What if we took tasks LLMs can do (plot summarization, bug finding) and progressively diluted them—more filler description, more boilerplate—to measure how much noise a model can tolerate before performance drops? Dilution tolerance as a benchmarked capability seems pretty key?
000
Ari @ari-holtzman.bsky.social · 09/04/2026
Adversarial prompts often hack an LLM's attention mechanism (e.g. from Aridti et al. 2024). Is it possible to make diluted adversarial prompts, or does dilution just cause them to stop actually interrupting processing meaningfully because there are enough other attractors?
100
Ari @ari-holtzman.bsky.social · 04/04/2026
if taste is truly the thing that matters in the era of GenAI, then it's not enough to be Rick Rubin. the human mind is too slow and too much of a bottleneck. one most be the Rick Rubin of Rick Rubin's, and be able to recognize taste in others, then amplify it
020
Ari @ari-holtzman.bsky.social · 04/04/2026
it's the beginning of the beginning
000
Ari @ari-holtzman.bsky.social · 04/04/2026
LLMs are basically epicycles for text, if we weren't sure heliocentrism existed
000
Ari @ari-holtzman.bsky.social · 03/04/2026
We should take relatively old LLMs, e.g., from 2023 or so, and see what data in 2026 they have trouble converging on via finetuning. Where are LLMs 'stuck' and where are they flexible?
000
Ari @ari-holtzman.bsky.social · 01/04/2026
030
Ari @ari-holtzman.bsky.social · 01/04/2026
machine unyearning
000
Ari @ari-holtzman.bsky.social · 31/03/2026
040
Ari @ari-holtzman.bsky.social · 30/03/2026
a more transparent (but much harder to sell) representation of LLMs would be less 'cautious assistant' and more 'genius toddler'
030
Ari @ari-holtzman.bsky.social · 29/03/2026
Something I find really annoying about most LLM discussions of research: they play conservative in their wording, but have zero epistemic humility when interpreting new data and will swing wildly back and forth between hypotheses in order to have something 'clean and easy' to say
390
Ari @ari-holtzman.bsky.social · 26/03/2026
LLMs use lookbacks to track beliefs and also suffer from anchoring bias. Are these the same? Do models anchor b/c they look at early info, or does the bias sometimes "infect" the residual stream at later positions, anchoring through an intermediary rather than direct attention?
100
Ari @ari-holtzman.bsky.social · 25/03/2026
I don't buy that "linear representations" in LLMs all form a composable linear space. I bet related directions (red, blue) compose naturally but unrelated ones don't. we should trawl through known directions and find which play nicely together to make a map of residual space
041
Ari @ari-holtzman.bsky.social · 09/03/2026
absolutely killer piece from @davidpreber.bsky.social
110
Ari @ari-holtzman.bsky.social · 26/02/2026
just had one of those days, where every time I started to get into something I realized I was late to something else
000
Ari @ari-holtzman.bsky.social · 25/02/2026
the question when planning an AI event is not whether to invite a Jason or even which Jason to invite, it's how many Jasons to invite
000
Ari @ari-holtzman.bsky.social · 24/02/2026
Turns out tiny, arbitrary-seeming parameter subsets can learn tasks. Has anyone compared what the same task looks like when learned through different components? Can we map how LMs encode information by seeing what stays the same in task representations over different components?
110
Ari @ari-holtzman.bsky.social · 21/02/2026
idea: a tea bag the screams when it's done steeping
010
Ari @ari-holtzman.bsky.social · 21/02/2026
If a company decides to build a better Siri-like app than Siri, I will donate footage of myself trying to use Siri for simple tasks and becoming increasingly annoyed for your advertisements
010
Ari @ari-holtzman.bsky.social · 21/02/2026
I would pay so much money for a non-chummy LLM
000
Ari @ari-holtzman.bsky.social · 14/02/2026
free idea for @perplexity_ai, @Google , etc: 1. ask an LLM to "describe this page" for each page in your index, teacher force "it's giving", and store whatever it generates next 2. embed the outputs 3. make an interface that lets me browse pages by 'close vibe but far by clicks'
030
Ari @ari-holtzman.bsky.social · 13/02/2026
when asking Claude for candidates, it suggested "vehement" and its suggested mispronouncation IS THE WAY I PRONOUNCE IT
000
Ari @ari-holtzman.bsky.social · 13/02/2026
some words I was confused by as a kid because I never heard someone say them and only saw them in books: - colonel - awry - buoy - facade - hors d'oeuvres - indict - genuine - genre - yacht - plaid - San Jose
020