Sign in

Dylan Foster 🐢

@djfoster.bsky.social
2.5K followers 838 following 113 posts

Incoming professor in EECS and Statistics at UC Berkeley. Principal Researcher @ Microsoft Research NE/NYC. Previously @ MIT, Cornell. AI + RL Foundations. RL Theory Lecture Notes: arxiv.org/abs/2312.16730 dylanfoster.net

PostsRepliesMedia
Reposted by Dylan Foster 🐢
RL Theory Virtual Seminars @rl-theory.bsky.social · 08/06/2026
Tomorrow, Zak will talk about his new deep RL method for hard exploration problems. Join us! The talk will be hosted by Csaba.
052
Reposted by Dylan Foster 🐢
Clément Canonne @ccanonne.github.io · 04/06/2026
Huge congratulations to Ilias Diakonikolas, Gautam Kamath, Daniel Kane, Jerry Li, Ankur Moitra, and Alistair Stewart on being awarded the Gödel prize for their breakthrough work on algorithmic robustness! www.sigact.org/prizes/g%C3%...
sigact.org
ACM SIGACT - Gödel Prize
1678
Reposted by Dylan Foster 🐢
Chris Paxton @cpaxton.bsky.social · 16/12/2025
New work in why action chunking is so important for robot control (it helps fight compounding error) arxiv.org/abs/2507.09061
1304
Reposted by Dylan Foster 🐢
Nathan Lambert @natolambert.bsky.social · 07/12/2025
Building Olmo 3 Think Foundations of Reasoning in Language Models @ NeurIPS 2025 Today 13:45 - 14:30
1171
Reposted by Dylan Foster 🐢
let-all.com @let-all.com · 26/11/2025
At #NeurIPS2025? Join us for a Social on Wednesday at 7 PM, featuring a fireside chat with Jon Kleinberg and mentoring tables. Ft. mentors @djfoster.bsky.social @surbhigoel.bsky.social @aifi.bsky.social @gautamkamath.com and more!
0144
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(12/12) Paper link: www.arxiv.org/abs/2510.15020 Many interesting questions: What are the minimal conditions for RL success (maybe a more semantic notion of coverage)? Do other algos/interventions secretly have coverage-based interpretations?
010
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(12/12) Another algo example: We give some tournament-style methods for selecting checkpoints (motivated by coverage) that can beat cross-entropy based selection. See figure, where coverage-based selection finds model w/ best BoN performance.
110
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(10/12) Perhaps the coolest part: the coverage perspective seems to have non-trivial algo design consequences: 1. Vanilla SGD can have poor coverage; gradient normalization (a-la Adam) fixes this 2. Test-time training (SGD on tokens during generation) provably improves coverage
110
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(9/12) We show this through two lenses: 1. New generalization analysis for next-token prediction/maximum likelihood for general function classes, in the vein of stat. learning. 2. SGD in the one-pass regime for overparameterized autoregressive linear models
100
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(8/12) Our analysis is based on a phenomenon we call the coverage principle, where next-token prediction implicitly optimizes toward a model with good coverage. Neat consequence: Coverage generalizes faster as we go further into tail (corresponding to more test-time compute)
100
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(7/12) Example (see figure): - Cross-entropy decreases throughout training. - Coverage improves to a point, but begins to drop as the model learns a spurious shortcut. - BoN performance follows trend of coverage, not CE (increasing initially, dropping as shortcut is learned).
110
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(6/12) Our main result: 1) Next-token prediction (more generally, MLE) provably learns a model w/ good coverage, inheriting the training corpus’ coverage over tasks of interest. 2) Coverage generalizes *faster* than cross-entropy itself, and consequently can be more predictive.
100
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(5/12) Example: For tasks that are well-represented in pre-training distribution, low Cov_N is necessary and sufficient for Best-of-N to succeed. Empirically, some notion of coverage also seems necessary for current RL methods (eg arxiv.org/abs/2504.13837)
100
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(4/12) Intuition: Cross-entropy (KL) tracks the mean of the log density ratio, while coverage profile tracks the *tail* or CDF via N. N controls how far into tail we go, and roughly corresponds to test-time compute/RL effort. For downstream performance, the tail is what matters.
100
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(3/12) Our answer is a quantity we call the *coverage profile*, which captures the extent to which the pre-trained model "covers" the pre-training distribution: Cov_N = P_{π}[π(y|x)/\hat{π}(y|x) ≥ N], (π is pre-training dist, \hat{π} is pre-trained model, N is a param)
100
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(2/12) We wanted to take a first step toward developing some theoretical understanding of the relationship btw next-token prediction and downstream perf (particularly w/ test-time scaling). Concretely, what metrics of the pre-trained model can we link to downstream success?
100
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(1/12) Many have observed starting that from a better next‑token predictor (low cross-entropy) does not always give better downstream post-training/test-time perf. In some cases, CE can even be anti‑correlated with BoN success (see fig or arxiv.org/abs/2502.07154). Why is this?
100
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
Led by my amazing intern Fan Chen, with awesome team Audrey Huang (@ahahaudrey.bsky.social), Noah Golowich, Sadhika Malladi (@sadhika.bsky.social), Adam Block, Jordan Ash, and Akshay Krishnamurthy. Paper: www.arxiv.org/abs/2510.15020 Thread below.
110
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
The coverage principle: How pre-training enables post-training New preprint where we look at the mechanisms through which next-token prediction produces models that succeed at downstream tasks. The answer involves a metric we call the "coverage profile", not cross-entropy.
1181
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025
(10/12) Perhaps the coolest part: the coverage perspective seems to have non-trivial algo design consequences: 1. Vanilla SGD can have poor coverage; gradient normalization (a-la Adam) fixes this 2. Test-time training (SGD on tokens during generation) provably improves coverage
000
Dylan Foster 🐢 @djfoster.bsky.social · 22/10/2025
Sorry about that! There was an issue where the application briefly closed early. Today is the deadline and it should be open again now.
010
Reposted by Dylan Foster 🐢
Aviad Rubinstein @aviad-rubinstein.bsky.social · 13/10/2025
The new call for Motwani postdocs application is now open! academicjobsonline.org/ajo/jobs/30865 BTW- Not quite ready for a postdoc? We updated the TCS Masters programs spreadsheet: www.cs.princeton.edu/~smattw/mast... Any career stage and in the (SF) Bay Area? Save the date for TOCA-SV on 11/7!
academicjobsonline.org
Stanford University, Computer Science/Theory Lab/Stanford University
Job #AJO30865, Postdoc in Theoretical Computer Science at Stanford, Computer Science/Theory Lab/Stanford University, Stanford University, Stanford, California, US
0148
Dylan Foster 🐢 @djfoster.bsky.social · 12/10/2025
Lots of interesting directions here! We think there is a lot more to do building on the connection to the discrete sampling/TCS literature and algos from this space, as well as moving beyond autoregressive generation.
000
Dylan Foster 🐢 @djfoster.bsky.social · 12/10/2025
Empirically, we have only tried this w/ small-ish scale so far, but find consistently that VGB outperforms textbook algos on either (1) accuracy; or (2) diversity when compute-normalized. Ex: for Dyck language, VGB escapes the accuracy-diversity frontier for baselines algos.
100
Dylan Foster 🐢 @djfoster.bsky.social · 12/10/2025
Main guarantee: - As long as you have exact/verifiable outcome rewards, always converges to optimal distribution. - Runtime depends on process verifier quality, gracefully degrading as quality gets worse.
100
Dylan Foster 🐢 @djfoster.bsky.social · 12/10/2025
VGB generalizes the Sinclair–Jerrum '89 random walk (people.eecs.berkeley.edu/~sinclair/ap...) from TCS (used to prove equivalence of apx. counting & sampling for self-reducible problems), linking test-time RL/alignment with discrete sampling theory. We are super excited about this connection.
100
Dylan Foster 🐢 @djfoster.bsky.social · 12/10/2025
We give a new algo, Value-Guided Backtracking (VGB), where the idea is to view autoregressive generation as a random walk on the tree of partial outputs, and add a *stochastic backtracking* step—occasionally erasing tokens in a principled way—to counter error amplification.
100
Dylan Foster 🐢 @djfoster.bsky.social · 12/10/2025
Test-time guidance with learned process verifiers has untapped potential to enhance language model reasoning, but one of the issues with getting this to actually work is that small verifier mistakes are amplified by textbook algos (e.g., block-wise BoN), w/ errors compounding as length increases.
100
Dylan Foster 🐢 @djfoster.bsky.social · 12/10/2025
With amazing team: Dhruv Rohatgi, Abhishek Shetty, Donya Saless, Yuchen Li, Ankur Moitra, and Andrej Risteski (andrejristeski.bsky.social). Short thread below.
andrejristeski.bsky.social
Andrej Risteski (@andrejristeski.bsky.social)
Machine learning researcher. Professor in ML department at CMU.
100
Dylan Foster 🐢 @djfoster.bsky.social · 12/10/2025
Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking. A totally new framework based on ~backtracking~ for using process verifiers to guide inference, w/ connections to approximate counting/sampling in theoretical CS. Paper: www.arxiv.org/abs/2510.03149
130
Dylan Foster 🐢 @djfoster.bsky.social · 02/10/2025
MSR NYC is hiring spring and summer interns in AI/ML/RL! Apply here: jobs.careers.microsoft.com/global/en/jo...
microsoft.com
Microsoft Research Lab - New York City - Microsoft Research
Apply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML.
0207
Reposted by Dylan Foster 🐢
Miro Dudik @mdudik.bsky.social · 18/09/2025
🚨Microsoft Research NYC is hiring🚨 We're hiring postdocs and senior researchers in AI/ML broadly, and in specific areas like test-time scaling and science of DL. Postdoc applications due Oct 22, 2025. Senior researcher applications considered on a rolling basis. Links to apply: aka.ms/msrnyc-jobs
aka.ms
Microsoft Research Lab - New York City - Microsoft Research
Apply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML.
0187
Dylan Foster 🐢 @djfoster.bsky.social · 12/09/2025
For details, see the following links: Empirical ML/AI: jobs.careers.microsoft.com/global/en/jo... Theoretical ML/AI: jobs.careers.microsoft.com/global/en/jo...
jobs.careers.microsoft.com
Search Jobs | Microsoft Careers
011
Dylan Foster 🐢 @djfoster.bsky.social · 12/09/2025
Microsoft Research New York City (www.microsoft.com/en-us/resear...) is seeking applicants for multiple Postdoctoral Researcher positions in ML/AI! These are positions for up to 2 years, starting in July 2026. Application deadline: October 22, 2025
microsoft.com
Microsoft Research Lab - New York City - Microsoft Research
Apply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML.
184
Dylan Foster 🐢 @djfoster.bsky.social · 27/08/2025
Quick reminder: The deadline for our workshop on Foundations of Reasoning in Language Models (FoRLM) at NeurIPS 2025 is next Wednesday, Sept 3!
000
Dylan Foster 🐢 @djfoster.bsky.social · 11/08/2025
Help us understand how reasoning emerges, where it fails, and how it can be systematically improved. Website (& CFP/instructions): reasoning-workshop.github.io Submission link: openreview.net/group?id=Neu... See you in San Diego!
reasoning-workshop.github.io
FoRLM @ NeurIPS'25
NeurIPS 2025 Workshop -- San Diego, California, USA
000
Dylan Foster 🐢 @djfoster.bsky.social · 11/08/2025
Led by Audrey Huang (ahahaudrey.bsky.social), with co-organizers Adam Block, Sadhika Malladi (sadhika.bsky.social), Will Merrill (lambdaviking.bsky.social), Pavel Izmailov, Akshay Krishnamurthy (akshaykr.bsky.social), Tatsunori Hashimoto, and myself.
ahaaudrey.bsky.social
Bluesky
100
Dylan Foster 🐢 @djfoster.bsky.social · 11/08/2025
Featuring amazing speakers Alekh Agarwal, Yejin Choi (yejinchoinka.bsky.social), Michael Hahn (m-hahn.bsky.social), and Nathan Lambert (natolambert.bsky.social).
yejinchoinka.bsky.social
yejinchoinka.bsky.social
100
Dylan Foster 🐢 @djfoster.bsky.social · 11/08/2025
Announcing the first workshop on Foundations of Language Model Reasoning (FoRLM) at NeurIPS 2025! 📝Soliciting abstracts that advance foundational understanding of reasoning in language models, from theoretical analyses to rigorous empirical studies. 📆 Deadline: Sept 3, 2025
1103
Dylan Foster 🐢 @djfoster.bsky.social · 15/07/2025
For those at ICML, Audrey will be presenting this paper at the 4:30pm poster session this afternoon! West Exhibition Hall B2-B3 W-1009
030
Reposted by Dylan Foster 🐢
Gautam Kamath @gautamkamath.com · 30/06/2025
ICML's election for their board of directors has begun. I've thrown my hat in the ring. Please consider voting for Gautam Kamath. I have experience with the governance of TMLR, COLT, and ALT, and I think I've demonstrated myself as a consciencious and engaged community member.
0295
Reposted by Dylan Foster 🐢
Tom Silver @tomssilver.bsky.social · 29/06/2025
This week's #PaperILike is "The Power of Resets in Online Reinforcement Learning" (Mhammedi et al., 2024). If you're doing RL in sim, why not use the sim to its full potential? Reset to any state! (gym.Env.reset() is not all we need.) PDF: arxiv.org/abs/2404.15417
arxiv.org
The Power of Resets in Online Reinforcement Learning
Simulators are a pervasive tool in reinforcement learning, but most existing algorithms cannot efficiently exploit simulator access -- particularly in high-dimensional domains that require general fun...
052
Reposted by Dylan Foster 🐢
let-all.com @let-all.com · 24/06/2025
📣Join us at COLT 2025 in Lyon for a community event! 📅When: Mon, June 30 | 16:00 CET What: Fireside chat w/ Peter Bartlett & Vitaly Feldman on communicating a research agenda, followed by mentorship roundtable to practice elevator pitches & mingle w/ COLT community! let-all.com/colt25.html
0156
Reposted by Dylan Foster 🐢
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 21/06/2025
Hiring a postdoc to scale up and deploy RL-based planning onto some self-driving cars! We'll be building on arxiv.org/abs/2502.03349 and learn what the limits and challenges of RL planning are. Shoot me a message if interested and help spread the word please! Full posting to come in a bit.
arxiv.org
Robust Autonomy Emerges from Self-Play
Self-play has powered breakthroughs in two-player and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic drivi...
36025
Reposted by Dylan Foster 🐢
Jason Hartline @jasonhartline.bsky.social · 09/06/2025
At the IDEAL annual meeting and saw this paper presented. Basically: reducing length of chain of thought LLM computations by deleting intermediate computations, more like classical functional programming where only function call and return values are important. arxiv.org/abs/2503.14337
arxiv.org
PENCIL: Long Thoughts with Short Memory
While recent works (e.g. o1, DeepSeek R1) have demonstrated great promise of using long Chain-of-Thought (CoT) to improve reasoning capabilities of language models, scaling it up during test-time is c...
031
Reposted by Dylan Foster 🐢
Clément Canonne @ccanonne.github.io · 04/06/2025
RADEMACHER CHAOS 🤘
031
Dylan Foster 🐢 @djfoster.bsky.social · 26/05/2025
Link: sites.google.com/view/rltheor...
sites.google.com
RL theory seminars - Next Seminar
May 27th 2025, 6 pm UTC Speaker: Dhruv Rohatgi (MIT) Title: Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification Pape...
000
Dylan Foster 🐢 @djfoster.bsky.social · 26/05/2025
Dhruv Rohatgi will be giving a lecture on our recent work on comp-stat tradeoffs in next-token prediction at the RL Theory virtual seminar series (rl-theory.bsky.social) tomorrow at 2pm EST! Should be a fun talk---come check it out!!
1105
Reposted by Dylan Foster 🐢
RL Theory Virtual Seminars @rl-theory.bsky.social · 20/05/2025
Later today, Sikata and Marcel will talk about their recent work on oracle-efficient RL with ensembles. Join us!
054
Dylan Foster 🐢 @djfoster.bsky.social · 19/05/2025
The abstract submission deadline for FoPt has been extended to the 21st of May (11:59pm UTC). Submission website: openreview.net/group?id=lea...
041