Aaron Roth @aaroth.bsky.social · 26/09/2026There are two papers we decided to submit to journals rather than ICLR to avoid the anticipated mess of this years review process. (But last time we submitted something to a Journal - GEB - the review process took so long that the journal collapsed before we ever got a decision) 2140
Aaron Roth @aaroth.bsky.social · 25/09/2026An annoying part of the twitter/bluesky discussion of whether LLMs are "only" "next token predictor"/stochastic parrots is that "next token predictor" is contentless - all mappings from inputs to output strings can be factored into a sequence of "next token" distributions. 160
Aaron Roth @aaroth.bsky.social · 24/09/2026AI tools are very useful for mathematical research. But at present, producing good work requires a labor intensive step that I have been calling "deslopping" - making the thing readable. Please do not skip this step; even if the theorem is great an unreadable paper is worthless. 4273
Aaron Roth @aaroth.bsky.social · 15/09/2026Codex and Claude Code have a neat auto approve feature where 1) it doesn't ask for permissions, but 2) you get to feel safe. Well --- do you? The premise is it is asking an agent for permission. But if you don't trust the driver agent, why should you trust the review agent? 2255
Aaron Roth @aaroth.bsky.social · 15/09/2026A world without open problems Here are some that fell today: K-server: arxiv.org/abs/2609.15979 Matroid Secretary: arxiv.org/abs/2609.145... Matrix Spencer: arxiv.org/abs/2609.15025 (Well Matrix Spencer was maybe also a few weeks ago, but who's counting? arxiv.org/abs/2608.28816 )arxiv.orgThe $k$-server conjecture is trueThe $k$-server conjecture states that a deterministic online algorithm can achieve competitive ratio $k$ on every metric space. We prove the conjecture. Specifically, we show that the work function al... 3316
Aaron Roth @aaroth.bsky.social · 14/09/2026I sometimes see people trying to use complexity theory to argue that building "true" AI is impossible. I find this unreasonably annoying. It requires ignoring what is in front of your face and it ignores that worst-case complexity has been an awful guide in machine learning. 1220
Reposted by Aaron RothAmazon Science @amazon.science · 10/09/2026Years of iterating against the same benchmarks should, by textbook logic, produce overfitting. It largely doesn't. New research explains why: strategies that generalize can be expressed in too compact a form to allow memorization, while the ones that overfit don't survive a compression.amazon.scienceWhy don’t machine learning research agents overfit?New research indicates that AI agents learn compressible models of data, which don’t have enough space to enable memorization. 094
Aaron Roth @aaroth.bsky.social · 10/09/2026A blog post on some neat work with @zstevenwu.bsky.social and Martin Bertran: www.amazon.science/blog/why-don...amazon.scienceWhy don’t machine learning research agents overfit?New research indicates that AI agents learn compressible models of data, which don’t have enough space to enable memorization. 063
Reposted by Aaron RothClément Canonne @ccanonne.github.io · 08/09/2026Well, this Navier-Stokes affair did blow up in finite time 843762
Aaron Roth @aaroth.bsky.social · 25/08/2026A new semester, and the first lecture is in the books for my class on the "Mathematical Foundations of AI Alignment". aaroth.github.io/cis-7000-ai-... What does that mean? Good question. We have about a semester in which to figure it out.aaroth.github.ioCIS 7000 — Mathematical Foundations of AI AlignmentCourse topics and reading list. 0323
Reposted by Aaron RothFoundations of Responsible Computing @forcconf.bsky.social · 17/08/2026Recordings from FORC 2026 are now available! Please check them out. Also, subscribe to FORC's new YouTube channel while you're at it! www.youtube.com/playlist?lis...youtube.comFORC 2026 - YouTubeTalk recordings from FORC 2026 056
Aaron Roth @aaroth.bsky.social · 14/08/2026We have a new online boosting algorithm which is very efficient and effective. Unlike prior algorithms which maintain many weak learners and ensemble them, we operationalize the "dual view" of boosting. We don't maintain an ensemble. We try to construct an online hard core distribution. 162
Reposted by Aaron RothGautam Kamath @gautamkamath.com · 10/08/2026I'm pleased to share our #ICML2026 tutorial on machine unlearning! Presented by Vinith Suriyakumar and myself, it includes a full set of videos recorded and posted to YouTube! Please check it out: unlearning-tutorial.github.io 1187
Aaron Roth @aaroth.bsky.social · 07/08/2026People used to be able to impress and intimidate reviewers with complicated proofs. This will change. In the age of AI inscrutable proofs are cheap. It is understandable proofs that are valuable. Opaque complexity is now it is a sign of laziness or lack of insight. 1638
Reposted by Aaron RothClément Canonne @ccanonne.github.io · 01/08/2026"We're all worried," as what it means to do research (in my field, Theoretical CS) seems to be shifting, and shifting fast. What to do? Senior researchers must lead by example, knowing that not everything will pan out. What I'm suggesting below may not work everywhere, but here's my own advice: 1/ 620351
Aaron Roth @aaroth.bsky.social · 01/08/2026Right now we have "problem overhang" - lots of problems we as a community are interested in because smart and charismatic people thought about them and convinced us that these problems are important. So we are happy/interested to see them solved by AI. 1100
Reposted by Aaron RothAaron Roth @aaroth.bsky.social · 25/07/2026It was fun giving this tutorial. If you missed it, check out calibration-tutorial.github.io where all of our materials are available, including slides, hundreds of pages of lecture notes, an interactive demo, and an annotated reading list. 0106
Aaron Roth @aaroth.bsky.social · 25/07/2026It was fun giving this tutorial. If you missed it, check out calibration-tutorial.github.io where all of our materials are available, including slides, hundreds of pages of lecture notes, an interactive demo, and an annotated reading list. 0106
Reposted by Aaron RothGautam Kamath @gautamkamath.com · 24/07/2026I've been referring people to this perspective every day for the last couple weeks. As we work together to determine new norms, we should not tolerate pure AI or otherwise bad writing. Tell your friends if they fall into this trap. 1172
Reposted by Aaron RothIra Globus-Harris @iraglobusharris.bsky.social · 06/07/2026Reminder, this is in a few hours!! j o i n u s 021
Reposted by Aaron RothGautam Kamath @gautamkamath.com · 02/07/2026Looking forward to being at #ICML2026 in Seoul next week! Unfortunately I'll be there for only 60 hours, but 2.5 of those will be at our machine unlearning tutorial (co-presented with Vinith Suriyakumar)! Monday July 6 at 9 AM -- don't miss out! 182
Reposted by Aaron RothAaron Roth @aaroth.bsky.social · 03/07/2026On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.calibration-tutorial.github.ioCalibration, Decisions, and Collaboration in Learning | ICML 2026An ICML 2026 tutorial on making probabilistic predictions trustworthy for downstream decision-making and collaboration. 133711
Reposted by Aaron RothIra Globus-Harris @iraglobusharris.bsky.social · 03/07/2026Are you at ICML next week? Feel like your decision-making for which sessions to attend might not be risk minimizing? Don't incur (swap) regret and come to my, @aaroth.bsky.social, and @ncollina.bsky.social's tutorial Monday on multicalibration, decision-making, and collaborative learning! 22011
Aaron Roth @aaroth.bsky.social · 03/07/2026On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.calibration-tutorial.github.ioCalibration, Decisions, and Collaboration in Learning | ICML 2026An ICML 2026 tutorial on making probabilistic predictions trustworthy for downstream decision-making and collaboration. 133711
Aaron Roth @aaroth.bsky.social · 01/07/2026AI is getting good at math. What are our jobs as researchers now that we have proof machines? The raw proofs that come from LLMs are difficult to understand, even if correct. So its now easy to quickly write many badly written papers that nevertheless contain correct proofs of interesting theorems. 2222
Aaron Roth @aaroth.bsky.social · 22/06/2026Interested to see how this goes. Now that the cost of generating a paper-like-object has dropped so low, publication venues are going to start having to impose costs on submissions of various sorts. We'll need experiments to figure out the best way to do this without disrupting science. 130
Reposted by Aaron RothAaron Roth @aaroth.bsky.social · 19/06/2026For a long time we didn't know if test-time randomization was needed for sample-optimal multicalibration and omniprediction. It's not. arxiv.org/abs/2606.20557 1112
Aaron Roth @aaroth.bsky.social · 19/06/2026For a long time we didn't know if test-time randomization was needed for sample-optimal multicalibration and omniprediction. It's not. arxiv.org/abs/2606.20557 1112
Aaron Roth @aaroth.bsky.social · 10/06/2026Modern LLMs are incredibly good compression algorithms, which can shed light on why autonomous data science agents don't overfit as much as you might think. arxiv.org/abs/2606.11045 1235
Reposted by Aaron RothGautam Kamath @gautamkamath.com · 28/05/2026In the last 48h: - Jr researcher asked me wheter to use AI in making talks - Saw two talks, with AI {slop, enhanced} slides Collected my thoughts and wrote a post. Tl;dr: don't steal your own thinking, don't remove *you* from your talks. Also, give a &#@% about your talks. 24913
Aaron Roth @aaroth.bsky.social · 20/05/2026A clearly hallucinated citation! NeurIPS 2026 decisions aren't out yet. But wait --- the hallucination is also present in the bibtex entries from openreview openreview.net/forum?id=fAj... and Google Scholar scholar.googleusercontent.com/scholar.bib?... 020
Aaron Roth @aaroth.bsky.social · 12/05/2026Recently we showed that the minimax optimal rate for multicalibration is T^{2/3}. But that doesn't mean you have to do that badly on all instances. We give an algorithm that can adapt to easy instances and get better rates while still being minimax optimal in the worst case. arxiv.org/abs/2605.09273 1121
Aaron Roth @aaroth.bsky.social · 05/05/2026I'm giving this talk at the MIT CS theory seminar tomorrow. Stop by if you are around! 040
Aaron Roth @aaroth.bsky.social · 27/04/2026We updated our paper --- and solved the open problem highlighted in the old version. Now our lower bound construction has only polylog(1/eps) many groups instead of poly(1/eps) many groups. The construction is also simplified. 061
Reposted by Aaron RothMarcel Hussing @marcelhussing.bsky.social · 26/04/2026Why do all LLMs predict 27 as their favorite number? There may be a principled explanation. Learn more at Agents in the Wild at #ICLR2026. @ericeaton.bsky.social, me, @surbhigoel.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, @optimistsinc.bsky.social 2122
Aaron Roth @aaroth.bsky.social · 24/04/2026How many samples do you need from an unknown distribution in order to train a model with multicalibration error at most epsilon? Answer: 1/epsilon^3 samples is both necessary and sufficient. 1200
Reposted by Aaron RothThe Warren Center for Network & Data Sciences @warrencenter.bsky.social · 21/04/2026April is #AIMonthAtPenn! On 4/24, Warren Center faculty affiliate Aaron Roth will give the George H. Heilmeier Faculty Award Lecture in Amy Gutmann Hall. More information and registration here: ai.upenn.edu/heilmei... 011
Aaron Roth @aaroth.bsky.social · 16/04/2026I've recently been getting invitations to talk about how to use AI tools to assist with TCS research. Its something I've been doing a lot, but don't have structured thoughts about how to explain process. But I'm going to try -- first such talk is tomorrow: t.co/wlHPBzXzDmt.cohttps://www.cics.umass.edu/events/research-ai-era-seminar-aaron-roth 072
Aaron Roth @aaroth.bsky.social · 13/04/2026AI Agents like Codex are very good at figuring out taxes, including obscure local ones that Intuit doesn't bother with (looking at you, Philadelphia local taxes). Businesses that provide financial/legal services that involve reasoning through dense but public documentation are in trouble. 050
Reposted by Aaron RothICML Conference @icmlconf.bsky.social · 02/04/2026Announcing the #ICML2026 tutorials! All ten tutorials will be presented the first day of the conference, Monday July 6. Read the blog post for more details on the selection process! blog.icml.cc/2026/04/02/a... 0138
Reposted by Aaron RothStephanie Tuerk @stephanietuerk.net · 13/03/2026So many interesting things here. (N.b. I get to think ab interfaces for this this all day long at work :)) One thing I find interesting here is how similar the real work of science is to that of the humanities, both of which are centered around human judgement ab what is relevant and interesting. 051
Aaron Roth @aaroth.bsky.social · 12/03/2026Very cool work. Empirical science has many researcher-degrees-of-freedom which makes it hard to interpret specific studies --- these are only a single trajectory through the data analysis multiverse. Human researchers are opaque. But with agents you can explore the whole space! 030
Aaron Roth @aaroth.bsky.social · 12/03/2026Neural networks are highly non-convex, so approximate error minimizers need not look anything like each other in parameter space. But we show that nevertheless (for many model sizes) approximate error minimizers must closely agree in function/prediction space despite this! 1107
Reposted by Aaron RothAaron Roth @aaroth.bsky.social · 09/03/2026Michael @mkearnsphilly.bsky.social ) and I wrote a blog post about our experiences using AI for research, and our thoughts on what these developments will mean for research, publication, and education: www.amazon.science/blog/how-ai-...amazon.scienceHow AI is changing the nature of mathematical researchWhat machine learning theorists learned using AI agents to generate proofs — and what comes next. 13013
Reposted by Aaron RothAmazon Science @amazon.science · 09/03/2026AI is bringing a sea change in scientific research methodology, training, and peer review. Amazon Scholars and Penn professors @mkearnsphilly.bsky.social and @aaroth.bsky.social on what agentic AI tools mean for the next generation of researchers.amazon.scienceHow AI is changing the nature of mathematical researchWhat machine learning theorists learned using AI agents to generate proofs — and what comes next. 052
Aaron Roth @aaroth.bsky.social · 09/03/2026Michael @mkearnsphilly.bsky.social ) and I wrote a blog post about our experiences using AI for research, and our thoughts on what these developments will mean for research, publication, and education: www.amazon.science/blog/how-ai-...amazon.scienceHow AI is changing the nature of mathematical researchWhat machine learning theorists learned using AI agents to generate proofs — and what comes next. 13013
Aaron Roth @aaroth.bsky.social · 04/03/2026Which is the better model for math? GPT 5.2, or GPT 5.3 codex (either one on high reasoning)? 120
Reposted by Aaron RothGautam Kamath @gautamkamath.com · 24/02/2026Sometimes you gotta split the difference. From Aaron Roth's (@aaroth.bsky.social) plenary talk at #ALT2026 0184