Sign in

Aaron Roth

@aaroth.bsky.social
4.4K followers 392 following 477 posts

Professor at Penn, Amazon Scholar at AWS. Interested in machine learning, uncertainty quantification, game theory, privacy, fairness, and most of the intersections therein

PostsRepliesMedia
Aaron Roth @aaroth.bsky.social · 15/09/2026
Our paper is here: arxiv.org/abs/2609.158... comments welcome! This is joint work with the wonderful @sikatasengupta.bsky.social @ncollina.bsky.social and @surbhigoel.bsky.social
050
Aaron Roth @aaroth.bsky.social · 15/09/2026
But despite being weaker it is not guaranteed to be satisfied. Still, we find some preliminary evidence of non-trivial safety/completeness tradeoffs among real reward models, and these effects come from coallitional rather than individual alignment.
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
What about long running agents? We study an MDP model where utilities depend on action and state and actions change state. We again get a characterization The performance difference identity lets us lift the one-shot conic-hull characterization (on Q values) to the full MDP.
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
There is a remarkably clean characterization of when you can. We start with a simple one-decision model. A system is safe if and only if the Principal's utility is in the non-negative span of the reviewer's utilities. The forward direction is elementary; backwards is LP duality.
110
Aaron Roth @aaroth.bsky.social · 15/09/2026
Codex and Claude Code have a neat auto approve feature where 1) it doesn't ask for permissions, but 2) you get to feel safe. Well --- do you? The premise is it is asking an agent for permission. But if you don't trust the driver agent, why should you trust the review agent?
2255
Aaron Roth @aaroth.bsky.social · 14/08/2026
So we get two incomparable boosting guarantees in one. And the algorithm works well, matching or beating the best comparator on real and synthetic datasets. Since we don't maintain an ensemble we run ~ 50x faster than these competitor algorithms as well.
100
Aaron Roth @aaroth.bsky.social · 14/08/2026
We have a new online boosting algorithm which is very efficient and effective. Unlike prior algorithms which maintain many weak learners and ensemble them, we operationalize the "dual view" of boosting. We don't maintain an ensemble. We try to construct an online hard core distribution.
162
Aaron Roth @aaroth.bsky.social · 25/07/2026
It was fun giving this tutorial. If you missed it, check out calibration-tutorial.github.io where all of our materials are available, including slides, hundreds of pages of lecture notes, an interactive demo, and an annotated reading list.
0106
Aaron Roth @aaroth.bsky.social · 19/06/2026
So do this. You get a randomized predictor whose feature conditional variance scales smoothly with the frequency of x. And once you have this you have avoided the obstruction in the red/blue example above, and fixing the randomness works!
000
Aaron Roth @aaroth.bsky.social · 19/06/2026
For a long time we didn't know if test-time randomization was needed for sample-optimal multicalibration and omniprediction. It's not. arxiv.org/abs/2606.20557
1112
Aaron Roth @aaroth.bsky.social · 10/06/2026
Modern LLMs are incredibly good compression algorithms, which can shed light on why autonomous data science agents don't overfit as much as you might think. arxiv.org/abs/2606.11045
1235
Aaron Roth @aaroth.bsky.social · 20/05/2026
A clearly hallucinated citation! NeurIPS 2026 decisions aren't out yet. But wait --- the hallucination is also present in the bibtex entries from openreview openreview.net/forum?id=fAj... and Google Scholar scholar.googleusercontent.com/scholar.bib?...
020
Aaron Roth @aaroth.bsky.social · 12/05/2026
This is joint work with the great Zhiming Huang, Jamie Morgenstern, and Claire Jie Zhang.
110
Aaron Roth @aaroth.bsky.social · 12/05/2026
Recently we showed that the minimax optimal rate for multicalibration is T^{2/3}. But that doesn't mean you have to do that badly on all instances. We give an algorithm that can adapt to easy instances and get better rates while still being minimax optimal in the worst case. arxiv.org/abs/2605.09273
1121
Aaron Roth @aaroth.bsky.social · 24/04/2026
Paper here: arxiv.org/abs/2604.21923 Joint work with the great @ncollina.bsky.social Jiuyao Lu, and George Noarov.
110
Aaron Roth @aaroth.bsky.social · 24/04/2026
How many samples do you need from an unknown distribution in order to train a model with multicalibration error at most epsilon? Answer: 1/epsilon^3 samples is both necessary and sufficient.
1200
Aaron Roth @aaroth.bsky.social · 13/03/2026
Alpha_0 joke from Dogman
020
Aaron Roth @aaroth.bsky.social · 12/03/2026
The paper is here: arxiv.org/abs/2602.23360 and is joint work with Eric Eaton, @surbhigoel.bsky.social, @marcelhussing.bsky.social, @mkearnsphilly.bsky.social, @sikatasengupta.bsky.social and @optimistsinc.bsky.social
061
Aaron Roth @aaroth.bsky.social · 12/03/2026
No matter what the Bayes error is, no matter how complicated the problem, the learning curve is monotonically decreasing and bounded above and below. So no matter what its shape is, it can't avoid "flatness" for a long stretch. Whenever you get flatness you get agreement.
100
Aaron Roth @aaroth.bsky.social · 12/03/2026
This lets us control out-of-the-box disagreement by the "local learning curve" --- how much can you decrease error by doubling the complexity of your model class? If the answer is not much, then you get out of the box agreement. The magic is that this is guaranteed to happen.
100
Aaron Roth @aaroth.bsky.social · 12/03/2026
The more interesting case: most model classes are not convex. But we can parameterize a family of model classes by a measure of complexity: size, depth, etc. Usually, the average of two models of complexity n is a model in the same class but with higher complexity, say 2n.
100
Aaron Roth @aaroth.bsky.social · 12/03/2026
But we can think about it the other way around. What if I imagined the ensemble of your model and mine. If this imagined ensemble would not substantially decrease error, then it must be that your model and mine already mostly agree. When does averaging not reduce error?
100
Aaron Roth @aaroth.bsky.social · 12/03/2026
The "Ambiguity Decomposition" is a fact that comes from ensembling theory. It says that if I average two models, the error of the ensemble is equal to the average of the model errors, less the "disagreement" of the two models. The goal is usually low error for the ensemble.
100
Aaron Roth @aaroth.bsky.social · 12/03/2026
Neural networks are highly non-convex, so approximate error minimizers need not look anything like each other in parameter space. But we show that nevertheless (for many model sizes) approximate error minimizers must closely agree in function/prediction space despite this!
1107
Aaron Roth @aaroth.bsky.social · 09/01/2026
The paper is here: arxiv.org/abs/2601.05245 Its joint work with @ncollina.bsky.social, Jiuyao Lu, and George Noarov. Natalie and George are on the job market --- check them out. www.seas.upenn.edu/~ncollina/ noarov.com
142
Aaron Roth @aaroth.bsky.social · 09/01/2026
Excited about a new paper! Multicalibration turns out to be strictly harder than marginal calibration. We prove tight Omega(T^{2/3}) lower bounds for online multicalibration, separating it from online marginal calibration for which better rates were recently discovered.
1225
Aaron Roth @aaroth.bsky.social · 30/10/2025
How should you use forecasts f:X->R^d to make decisions? It depends what properties they have. If they are fully calibrated (E[y | f(x) = p] = p), then you should be maximally aggressive and act as if they are correct --- i.e. play argmax_a E_{o ~ f(x)}[u(a,o)]. On the other hand
1162
Aaron Roth @aaroth.bsky.social · 24/09/2025
The opportunities and risks of the entry of LLMs into mathematical research in one screenshot. I think it is clear that LLMs will make trained researchers more effective. But they will also lead to a flood of bad/wrong papers, and I'm not sure we have the tools to deal with this.
4296
Aaron Roth @aaroth.bsky.social · 19/09/2025
The paper is here: arxiv.org/abs/2509.15090 and is joint work with the excellent @ncollina.bsky.social , @surbhigoel.bsky.social , Emily Ryu, and Mirah Shi.
071
Aaron Roth @aaroth.bsky.social · 19/09/2025
I'm excited about this line of work. We get clean results in a stylized setting, and there is much to do to bring these kinds of ideas closer to practice. But I think that ideas from market and mechanism design should have lots to say about the practical alignment problem too.
150
Aaron Roth @aaroth.bsky.social · 19/09/2025
We conduct another simple experiment in which we explicitly compute equilibria amongst differently misaligned sets of agents. The results validate our theory --- user utility can sometimes match our worst case bound (so its tight), but is often much better.
100
Aaron Roth @aaroth.bsky.social · 19/09/2025
We give simple experiments (where LLM personas are generated with prompt variation) demonstrating that representing user utility functions somewhere in the convex hull of LLM utility functions is a much easier target than finding a single well aligned LLM utility function.
100
Aaron Roth @aaroth.bsky.social · 19/09/2025
Even if all of the AI providers are very badly aligned, if it is possible to approximate the user's utility by any non-negative linear combination of their utilities, then the user does as well as they would with a perfectly aligned AI. Alignment emerges from competition.
120
Aaron Roth @aaroth.bsky.social · 19/09/2025
We give several mathematical models of AI competition in which the answer is yes, provided that the user's utility function lies in the convex hull of the AI utility functions. Under this condition, all equilibria of the game between AI providers leads to high user utility.
100
Aaron Roth @aaroth.bsky.social · 19/09/2025
Aligning an AI with human preferences might be hard. But there is more than one AI out there, and users can choose which to use. Can we get the benefits of a fully aligned AI without solving the alignment problem? In a new paper we study a setting in which the answer is yes.
1274
Aaron Roth @aaroth.bsky.social · 24/07/2025
The paper is here: arxiv.org/abs/2507.09683 This is joint work with @mkearnsphilly.bsky.social and Emily Ryu, from a fun visit in June.
131
Aaron Roth @aaroth.bsky.social · 24/07/2025
Suppose we have a regression problem that many people want to solve, but information is distributed across agents --- different agents see different subsets of features. And the agents are embedded in a network (DAG). When one makes predictions, they are observed by its children.
1151
Aaron Roth @aaroth.bsky.social · 12/06/2025
Agentic LLM tooling like Windsurf is amazing. But you can already tell from the interface (you can pick from two dozen models from half a dozen different model providers, seamlessly) that AI is going to be a commodity. Not at all clear that LLM developers will get the surplus.
130
Aaron Roth @aaroth.bsky.social · 09/04/2025
The paper is here: arxiv.org/abs/2504.06075 and is joint work with the excellent @ncollina.bsky.social , @iraglobusharris.bsky.social , @surbhigoel.bsky.social , Varun Gupta, and Mirah Shi!
062
Aaron Roth @aaroth.bsky.social · 09/04/2025
Just as our last paper generalized Aumann's agreement theorem, this paper tractably generalizes "information aggregation" theorems for Bayesian reasoners. Our results lift back to the classic Bayesian setting and give the first distribution free information aggregation theorems.
150
Aaron Roth @aaroth.bsky.social · 09/04/2025
For any function classes H(A), H(B), H(J) satisfying this weak learning condition, we show how two parties can collaborate to be as accurate at H(J), by only needing to solve a small number of squared error regression problems on their own data over H(A) and H(B) respectively.
120
Aaron Roth @aaroth.bsky.social · 09/04/2025
H(A) and H(B) are weak learners wrt H(J) if, whenever learning over H(J) can improve on constant prediction, learning over H(A) or H(B) can as well. It turns out linear functions over X(A) and linear functions over X(B) are weak learners wrt linear functions over the joint features.
120
Aaron Roth @aaroth.bsky.social · 09/04/2025
In our new paper, what we show is that if both parties further have no swap regret (are "multicalibrated") with respect to function classes H(A), H(B) defined on their own features, then we can give accuracy guarantees with respect to a joint function class H(J).
140
Aaron Roth @aaroth.bsky.social · 09/04/2025
Suppose you and I both have different features about the same instance. Maybe I have CT scans and you have physician notes. We'd like to collaborate to make predictions that are more accurate than possible from either feature set alone, while only having to train on our own data.
2437
Aaron Roth @aaroth.bsky.social · 27/03/2025
Its interesting to see OpenAI slowly tighten the guardrails around the Studio Ghibli meme.
032
Aaron Roth @aaroth.bsky.social · 24/03/2025
My first attempt at writing a Philosophy paper with Alex. Forthcoming in Philosophy of Science! We study the reference class problem which is about the indeterminacy of "individual probabilities" from data. What is the chance that Alice will die in the next 12 months? philsci-archive.pitt.edu/23589/
4203
Aaron Roth @aaroth.bsky.social · 01/03/2025
A breakthrough in swap theory.
030
Aaron Roth @aaroth.bsky.social · 25/02/2025
Alex Tolbert is running a stellar conference next week at Emory that I am bummed to be missing out on. The speaker lineup is especially remarkable- spanning theoretical computer science to machine learning to law to philosophy. You should go and enjoy it for me. 39893947.hs-sites.com/aiethicsconf...
0123
Aaron Roth @aaroth.bsky.social · 18/02/2025
Can you solve group-conditional online conformal prediction with a no-regret learning algorithm? Not with vanilla regret, but -yes- with swap regret. And algorithms from the follow-the-regularized leader family (notably online gradient descent) work really well for other reasons.
1202
Aaron Roth @aaroth.bsky.social · 14/02/2025
Also this is an incredibly wide confidence interval given the price.
170