Sign in

Aaron Roth

@aaroth.bsky.social
4.4K followers 392 following 476 posts

Professor at Penn, Amazon Scholar at AWS. Interested in machine learning, uncertainty quantification, game theory, privacy, fairness, and most of the intersections therein

PostsRepliesMedia
Reposted by Aaron Roth
Wyatt Walls @wwalls.bsky.social · 30/09/2026
Anthropic system prompt change that shows how careful you need to be with words: Opus 4.1: "Claude never curses unless the human asks for it" Opus 4.5: "Claude never curses unless the person asks Claude to curse"
626421
Aaron Roth @aaroth.bsky.social · 26/09/2026
There are two papers we decided to submit to journals rather than ICLR to avoid the anticipated mess of this years review process. (But last time we submitted something to a Journal - GEB - the review process took so long that the journal collapsed before we ever got a decision)
2140
Aaron Roth @aaroth.bsky.social · 25/09/2026
(Re: training, the capabilities I'm most impressed by seem not to come from next-token training. So the training story is now complicated. What remains straightforward next-token-prediction is just the inference architecture)
000
Aaron Roth @aaroth.bsky.social · 25/09/2026
My point is mundane; that any distribution over sequences can be factored into conditional distributions on the next element in the sequence. As a result, the fact that LLMs produce outputs factored in this way provides no evidence of its capabilities one way or another, we need to look elsewhere.
110
Aaron Roth @aaroth.bsky.social · 25/09/2026
I feel like you've misunderstood what I'm trying to say.
001
Aaron Roth @aaroth.bsky.social · 25/09/2026
Who said anything about consciousness? I also have little explanation for why llms work so well, and I don't think anyone does at this point. Clearly not all systems are equivalent.
110
Aaron Roth @aaroth.bsky.social · 25/09/2026
The interesting question a decade ago about representing output distributions as next token distributions was computational: maybe this would introduce computational difficulties that would make it hard to learn useful behaviors. It turned out it didn't (at least given what was spent)!
020
Aaron Roth @aaroth.bsky.social · 25/09/2026
In this sense you, or any other system that can output text can be written as a "next token predictor". This isn't making any statement about whether brains function similarly to LLMs or not at an internal level, its just a boring syntactic fact about probability distributions.
250
Aaron Roth @aaroth.bsky.social · 25/09/2026
An annoying part of the twitter/bluesky discussion of whether LLMs are "only" "next token predictor"/stochastic parrots is that "next token predictor" is contentless - all mappings from inputs to output strings can be factored into a sequence of "next token" distributions.
160
Aaron Roth @aaroth.bsky.social · 24/09/2026
AI tools are very useful for mathematical research. But at present, producing good work requires a labor intensive step that I have been calling "deslopping" - making the thing readable. Please do not skip this step; even if the theorem is great an unreadable paper is worthless.
4273
Aaron Roth @aaroth.bsky.social · 15/09/2026
I don't have an opinion about these papers specifically because I haven't read them. But I do think that acceptance criteria will have to and will change. And they will have to focus much more on exposition and clarity than they used to.
021
Aaron Roth @aaroth.bsky.social · 15/09/2026
Our paper is here: arxiv.org/abs/2609.158... comments welcome! This is joint work with the wonderful @sikatasengupta.bsky.social @ncollina.bsky.social and @surbhigoel.bsky.social
050
Aaron Roth @aaroth.bsky.social · 15/09/2026
But despite being weaker it is not guaranteed to be satisfied. Still, we find some preliminary evidence of non-trivial safety/completeness tradeoffs among real reward models, and these effects come from coallitional rather than individual alignment.
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
We call this condition "coalitional alignment" -- it is very similar to the market alignment condition that drives our earlier work: arxiv.org/abs/2509.15090 It is a substantially weaker condition than individual alignment (That the reviewer agents have your utility exactly)
arxiv.org
Emergent Alignment via Competition
Aligning AI systems with human values remains a fundamental challenge, but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We study a strategic sett...
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
What about strategic agents? Let both the driver agent and the reviewers be fully strategic and optimizing for their long run discounted payoff. Under the same non-negative span condition, every Nash equilibrium of the game is safe for the Principal--i.e. no worse than baseline.
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
What about long running agents? We study an MDP model where utilities depend on action and state and actions change state. We again get a characterization The performance difference identity lets us lift the one-shot conic-hull characterization (on Q values) to the full MDP.
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
There is a remarkably clean characterization of when you can. We start with a simple one-decision model. A system is safe if and only if the Principal's utility is in the non-negative span of the reviewer's utilities. The forward direction is elementary; backwards is LP duality.
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
The difficulty is reviewer agents are misaligned - their utility functions differ from yours. Can you still guarantee that the system is safe, even with an arbitrary driver agent? Safe means that principal utility is guaranteed to be no worse than following the baseline policy.
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
We study this and model both you (the principal) and the reviewer agent(s) as having utility functions you are trying to maximize. The driver agent repeatedly proposes actions. The reviewer agents compare them to a baseline and vote to approve or deny based on perceived utility.
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
Codex and Claude Code have a neat auto approve feature where 1) it doesn't ask for permissions, but 2) you get to feel safe. Well --- do you? The premise is it is asking an agent for permission. But if you don't trust the driver agent, why should you trust the review agent?
2255
Aaron Roth @aaroth.bsky.social · 15/09/2026
A world without open problems Here are some that fell today: K-server: arxiv.org/abs/2609.15979 Matroid Secretary: arxiv.org/abs/2609.145... Matrix Spencer: arxiv.org/abs/2609.15025 (Well Matrix Spencer was maybe also a few weeks ago, but who's counting? arxiv.org/abs/2608.28816 )
arxiv.org
The $k$-server conjecture is true
The $k$-server conjecture states that a deterministic online algorithm can achieve competitive ratio $k$ on every metric space. We prove the conjecture. Specifically, we show that the work function al...
3316
Aaron Roth @aaroth.bsky.social · 14/09/2026
i.e. that the laws of computation bind human cognition just as much as they bind silicon. This is merely a conjecture --- maybe its not true and we have some non-computational secret sauce. But now making this argument requires ignoring evidence in silico (and in browser tab)
030
Aaron Roth @aaroth.bsky.social · 14/09/2026
Applying this kind of argument to more general intelligence would have been wrong a decade ago but now requires ignoring the obvious success of AI and particularly the reasoning models from the last few years. A decade ago we could have said it denies computationalism:
120
Aaron Roth @aaroth.bsky.social · 14/09/2026
But of course this has been no obstacle to solving almost every real-world statistical learning problem of this sort. The model of worst-case complexity is elegant but turns out to be a poor guide for what is possible (or even easy and reliable) on naturally occuring problems.
280
Aaron Roth @aaroth.bsky.social · 14/09/2026
For example, the simplest problem in machine learning - learning a linear classifier to minimize classification error --is NP hard even to approximate (to distinguish the ability to get 51% accuracy from 99% accuracy). epubs.siam.org/doi/10.1137/... and thus everything more general too.
epubs.siam.org
Hardness of Learning Halfspaces with Noise | SIAM Journal on Computing
Learning an unknown halfspace (also called a perceptron) from labeled examples is one of the classic problems in machine learning. In the noise-free case, when a halfspace consistent with all the trai...
130
Aaron Roth @aaroth.bsky.social · 14/09/2026
I sometimes see people trying to use complexity theory to argue that building "true" AI is impossible. I find this unreasonably annoying. It requires ignoring what is in front of your face and it ignores that worst-case complexity has been an awful guide in machine learning.
1220
Reposted by Aaron Roth
Amazon Science @amazon.science · 10/09/2026
Years of iterating against the same benchmarks should, by textbook logic, produce overfitting. It largely doesn't. New research explains why: strategies that generalize can be expressed in too compact a form to allow memorization, while the ones that overfit don't survive a compression.
amazon.science
Why don’t machine learning research agents overfit?
New research indicates that AI agents learn compressible models of data, which don’t have enough space to enable memorization.
094
Aaron Roth @aaroth.bsky.social · 10/09/2026
A blog post on some neat work with @zstevenwu.bsky.social and Martin Bertran: www.amazon.science/blog/why-don...
amazon.science
Why don’t machine learning research agents overfit?
New research indicates that AI agents learn compressible models of data, which don’t have enough space to enable memorization.
063
Reposted by Aaron Roth
Clément Canonne @ccanonne.github.io · 08/09/2026
Well, this Navier-Stokes affair did blow up in finite time
843862
Aaron Roth @aaroth.bsky.social · 25/08/2026
A new semester, and the first lecture is in the books for my class on the "Mathematical Foundations of AI Alignment". aaroth.github.io/cis-7000-ai-... What does that mean? Good question. We have about a semester in which to figure it out.
aaroth.github.io
CIS 7000 — Mathematical Foundations of AI Alignment
Course topics and reading list.
0323
Reposted by Aaron Roth
Foundations of Responsible Computing @forcconf.bsky.social · 17/08/2026
Recordings from FORC 2026 are now available! Please check them out. Also, subscribe to FORC's new YouTube channel while you're at it! www.youtube.com/playlist?lis...
youtube.com
FORC 2026 - YouTube
Talk recordings from FORC 2026
056
Aaron Roth @aaroth.bsky.social · 14/08/2026
Ya you get stronger regularity lemmas from multicalibration. But its important not to use these in the online setting since (even marginal) calibration can't be obtained at sqrt{T} rates in online adversarial settings, whereas multiaccuracy can be.
020
Aaron Roth @aaroth.bsky.social · 14/08/2026
Putting them together works pretty well! This is joint work with the excellent Georgy Noarov. The paper is here: arxiv.org/abs/2608.13554 and because we live in a magical new world, the code is here: github.com/aaroth/defen...
arxiv.org
Defensive Boosting for Online Probabilistic Forecasting
We study online probabilistic forecasting of binary outcomes chosen by an adaptive adversary. Given an online learning algorithm for a weak hypothesis class $H$, we would like to efficiently obtain tw...
000
Aaron Roth @aaroth.bsky.social · 14/08/2026
In the end we're putting a bunch of existing pieces together. The fact that multiaccuracy implies weights for a hard core distribution was noticed by @sicapu.bsky.social , Dwork, and Vadhan. Lots of good prior work gave online adversarial algorithms for multiaccuracy and related notions.
100
Aaron Roth @aaroth.bsky.social · 14/08/2026
Online learning tricks let us make this guarantee adaptive - i.e. we get the boosting guarantee on every sub-interval on which the weak learning condition holds, even if it doesn't on others. This works through an adaptive version of a hard core distribution which seems like an interesting object.
100
Aaron Roth @aaroth.bsky.social · 14/08/2026
So we get two incomparable boosting guarantees in one. And the algorithm works well, matching or beating the best comparator on real and synthetic datasets. Since we don't maintain an ensemble we run ~ 50x faster than these competitor algorithms as well.
100
Aaron Roth @aaroth.bsky.social · 14/08/2026
The weak learning assumption is that there is no such distribution - so if it holds, by contrapositive, our error was low. Adding one more orthogonality condition also lets us compete with any predictor in the span of the weak class - the usual gradient boosting guarantee.
100
Aaron Roth @aaroth.bsky.social · 14/08/2026
We make probabilistic forecasts that are multiaccurate with respect to the weak class---i.e. with residuals orthogonal to the weak class. If we have high error, then ex-post transcript reweighting by our own residual error gives a smooth distribution on which no weak learner has non-trivial edge.
100
Aaron Roth @aaroth.bsky.social · 14/08/2026
We have a new online boosting algorithm which is very efficient and effective. Unlike prior algorithms which maintain many weak learners and ensemble them, we operationalize the "dual view" of boosting. We don't maintain an ensemble. We try to construct an online hard core distribution.
162
Reposted by Aaron Roth
Gautam Kamath @gautamkamath.com · 10/08/2026
I'm pleased to share our #ICML2026 tutorial on machine unlearning! Presented by Vinith Suriyakumar and myself, it includes a full set of videos recorded and posted to YouTube! Please check it out: unlearning-tutorial.github.io
1187
Aaron Roth @aaroth.bsky.social · 07/08/2026
People used to be able to impress and intimidate reviewers with complicated proofs. This will change. In the age of AI inscrutable proofs are cheap. It is understandable proofs that are valuable. Opaque complexity is now it is a sign of laziness or lack of insight.
1638
Reposted by Aaron Roth
Clément Canonne @ccanonne.github.io · 01/08/2026
"We're all worried," as what it means to do research (in my field, Theoretical CS) seems to be shifting, and shifting fast. What to do? Senior researchers must lead by example, knowing that not everything will pan out. What I'm suggesting below may not work everywhere, but here's my own advice: 1/
620351
Aaron Roth @aaroth.bsky.social · 01/08/2026
So we are clearly moving to a world that will be without open problems, but this doesn't mean math is going away. Interesting work will correspond to discovering new questions and coherent theoretical frameworks. Well defined agreed-to-be-interesting problems won't last long.
090
Aaron Roth @aaroth.bsky.social · 01/08/2026
Could it be that once AI models exhaust the overhang of existing problems we care about, no new such problems emerge? There will be plenty of true statements with difficult proofs, but we may stop caring about (any of) them.
130
Aaron Roth @aaroth.bsky.social · 01/08/2026
But as we move to a world in which well defined "open problems" are solved by AI, will we find new problems that we as a community care about? We don't care about most true statements, even if they are difficult to prove. The process of agreeing on interesting problems is social.
160
Aaron Roth @aaroth.bsky.social · 01/08/2026
Right now we have "problem overhang" - lots of problems we as a community are interested in because smart and charismatic people thought about them and convinced us that these problems are important. So we are happy/interested to see them solved by AI.
1100
Reposted by Aaron Roth
Aaron Roth @aaroth.bsky.social · 25/07/2026
It was fun giving this tutorial. If you missed it, check out calibration-tutorial.github.io where all of our materials are available, including slides, hundreds of pages of lecture notes, an interactive demo, and an annotated reading list.
0106
Aaron Roth @aaroth.bsky.social · 25/07/2026
It was fun giving this tutorial. If you missed it, check out calibration-tutorial.github.io where all of our materials are available, including slides, hundreds of pages of lecture notes, an interactive demo, and an annotated reading list.
0106
Reposted by Aaron Roth
Gautam Kamath @gautamkamath.com · 24/07/2026
I've been referring people to this perspective every day for the last couple weeks. As we work together to determine new norms, we should not tolerate pure AI or otherwise bad writing. Tell your friends if they fall into this trap.
1172
Reposted by Aaron Roth
Ira Globus-Harris @iraglobusharris.bsky.social · 06/07/2026
Reminder, this is in a few hours!! j o i n u s
021