Reposted by Dylan Foster 🐢RL Theory Virtual Seminars @rl-theory.bsky.social · 08/06/2026Tomorrow, Zak will talk about his new deep RL method for hard exploration problems. Join us! The talk will be hosted by Csaba. 052
Reposted by Dylan Foster 🐢Clément Canonne @ccanonne.github.io · 04/06/2026Huge congratulations to Ilias Diakonikolas, Gautam Kamath, Daniel Kane, Jerry Li, Ankur Moitra, and Alistair Stewart on being awarded the Gödel prize for their breakthrough work on algorithmic robustness! www.sigact.org/prizes/g%C3%...sigact.orgACM SIGACT - Gödel Prize 1678
Reposted by Dylan Foster 🐢Chris Paxton @cpaxton.bsky.social · 16/12/2025New work in why action chunking is so important for robot control (it helps fight compounding error) arxiv.org/abs/2507.09061 1304
Reposted by Dylan Foster 🐢Nathan Lambert @natolambert.bsky.social · 07/12/2025Building Olmo 3 Think Foundations of Reasoning in Language Models @ NeurIPS 2025 Today 13:45 - 14:30 1171
Reposted by Dylan Foster 🐢let-all.com @let-all.com · 26/11/2025At #NeurIPS2025? Join us for a Social on Wednesday at 7 PM, featuring a fireside chat with Jon Kleinberg and mentoring tables. Ft. mentors @djfoster.bsky.social @surbhigoel.bsky.social @aifi.bsky.social @gautamkamath.com and more! 0144
Dylan Foster 🐢 @djfoster.bsky.social · 25/10/2025The coverage principle: How pre-training enables post-training New preprint where we look at the mechanisms through which next-token prediction produces models that succeed at downstream tasks. The answer involves a metric we call the "coverage profile", not cross-entropy. 1181
Reposted by Dylan Foster 🐢Aviad Rubinstein @aviad-rubinstein.bsky.social · 13/10/2025The new call for Motwani postdocs application is now open! academicjobsonline.org/ajo/jobs/30865 BTW- Not quite ready for a postdoc? We updated the TCS Masters programs spreadsheet: www.cs.princeton.edu/~smattw/mast... Any career stage and in the (SF) Bay Area? Save the date for TOCA-SV on 11/7!academicjobsonline.orgStanford University, Computer Science/Theory Lab/Stanford University Job #AJO30865, Postdoc in Theoretical Computer Science at Stanford, Computer Science/Theory Lab/Stanford University, Stanford University, Stanford, California, US 0148
Dylan Foster 🐢 @djfoster.bsky.social · 12/10/2025Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking. A totally new framework based on ~backtracking~ for using process verifiers to guide inference, w/ connections to approximate counting/sampling in theoretical CS. Paper: www.arxiv.org/abs/2510.03149 130
Dylan Foster 🐢 @djfoster.bsky.social · 02/10/2025MSR NYC is hiring spring and summer interns in AI/ML/RL! Apply here: jobs.careers.microsoft.com/global/en/jo...microsoft.comMicrosoft Research Lab - New York City - Microsoft ResearchApply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML. 0207
Reposted by Dylan Foster 🐢Miro Dudik @mdudik.bsky.social · 18/09/2025🚨Microsoft Research NYC is hiring🚨 We're hiring postdocs and senior researchers in AI/ML broadly, and in specific areas like test-time scaling and science of DL. Postdoc applications due Oct 22, 2025. Senior researcher applications considered on a rolling basis. Links to apply: aka.ms/msrnyc-jobsaka.msMicrosoft Research Lab - New York City - Microsoft ResearchApply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML. 0187
Dylan Foster 🐢 @djfoster.bsky.social · 12/09/2025Microsoft Research New York City (www.microsoft.com/en-us/resear...) is seeking applicants for multiple Postdoctoral Researcher positions in ML/AI! These are positions for up to 2 years, starting in July 2026. Application deadline: October 22, 2025microsoft.comMicrosoft Research Lab - New York City - Microsoft ResearchApply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML. 184
Dylan Foster 🐢 @djfoster.bsky.social · 27/08/2025Quick reminder: The deadline for our workshop on Foundations of Reasoning in Language Models (FoRLM) at NeurIPS 2025 is next Wednesday, Sept 3! 000
Dylan Foster 🐢 @djfoster.bsky.social · 11/08/2025Announcing the first workshop on Foundations of Language Model Reasoning (FoRLM) at NeurIPS 2025! 📝Soliciting abstracts that advance foundational understanding of reasoning in language models, from theoretical analyses to rigorous empirical studies. 📆 Deadline: Sept 3, 2025 1103
Dylan Foster 🐢 @djfoster.bsky.social · 15/07/2025For those at ICML, Audrey will be presenting this paper at the 4:30pm poster session this afternoon! West Exhibition Hall B2-B3 W-1009 030
Reposted by Dylan Foster 🐢Gautam Kamath @gautamkamath.com · 30/06/2025ICML's election for their board of directors has begun. I've thrown my hat in the ring. Please consider voting for Gautam Kamath. I have experience with the governance of TMLR, COLT, and ALT, and I think I've demonstrated myself as a consciencious and engaged community member. 0295
Reposted by Dylan Foster 🐢Tom Silver @tomssilver.bsky.social · 29/06/2025This week's #PaperILike is "The Power of Resets in Online Reinforcement Learning" (Mhammedi et al., 2024). If you're doing RL in sim, why not use the sim to its full potential? Reset to any state! (gym.Env.reset() is not all we need.) PDF: arxiv.org/abs/2404.15417arxiv.orgThe Power of Resets in Online Reinforcement LearningSimulators are a pervasive tool in reinforcement learning, but most existing algorithms cannot efficiently exploit simulator access -- particularly in high-dimensional domains that require general fun... 052
Reposted by Dylan Foster 🐢let-all.com @let-all.com · 24/06/2025📣Join us at COLT 2025 in Lyon for a community event! 📅When: Mon, June 30 | 16:00 CET What: Fireside chat w/ Peter Bartlett & Vitaly Feldman on communicating a research agenda, followed by mentorship roundtable to practice elevator pitches & mingle w/ COLT community! let-all.com/colt25.html 0156
Reposted by Dylan Foster 🐢Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 21/06/2025Hiring a postdoc to scale up and deploy RL-based planning onto some self-driving cars! We'll be building on arxiv.org/abs/2502.03349 and learn what the limits and challenges of RL planning are. Shoot me a message if interested and help spread the word please! Full posting to come in a bit.arxiv.orgRobust Autonomy Emerges from Self-PlaySelf-play has powered breakthroughs in two-player and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic drivi... 36025
Reposted by Dylan Foster 🐢Jason Hartline @jasonhartline.bsky.social · 09/06/2025At the IDEAL annual meeting and saw this paper presented. Basically: reducing length of chain of thought LLM computations by deleting intermediate computations, more like classical functional programming where only function call and return values are important. arxiv.org/abs/2503.14337arxiv.orgPENCIL: Long Thoughts with Short MemoryWhile recent works (e.g. o1, DeepSeek R1) have demonstrated great promise of using long Chain-of-Thought (CoT) to improve reasoning capabilities of language models, scaling it up during test-time is c... 031
Dylan Foster 🐢 @djfoster.bsky.social · 26/05/2025Dhruv Rohatgi will be giving a lecture on our recent work on comp-stat tradeoffs in next-token prediction at the RL Theory virtual seminar series (rl-theory.bsky.social) tomorrow at 2pm EST! Should be a fun talk---come check it out!! 1105
Reposted by Dylan Foster 🐢RL Theory Virtual Seminars @rl-theory.bsky.social · 20/05/2025Later today, Sikata and Marcel will talk about their recent work on oracle-efficient RL with ensembles. Join us! 054
Dylan Foster 🐢 @djfoster.bsky.social · 19/05/2025The abstract submission deadline for FoPt has been extended to the 21st of May (11:59pm UTC). Submission website: openreview.net/group?id=lea... 041
Reposted by Dylan Foster 🐢Dylan Foster 🐢 @djfoster.bsky.social · 09/05/2025Announcing the first workshop on Foundations of Post-Training (FoPT) at COLT 2025! 📝 Soliciting abstracts/posters exploring theoretical & practical aspects of post-training and RL with language models! 🗓️ Deadline: May 19, 2025 1165
Dylan Foster 🐢 @djfoster.bsky.social · 09/05/2025Announcing the first workshop on Foundations of Post-Training (FoPT) at COLT 2025! 📝 Soliciting abstracts/posters exploring theoretical & practical aspects of post-training and RL with language models! 🗓️ Deadline: May 19, 2025 1165
Dylan Foster 🐢 @djfoster.bsky.social · 03/05/2025Is Best-of-N really the best we can do for language model inference? New paper (appearing at ICML) led by the amazing Audrey Huang (ahahaudrey.bsky.social) with Adam Block, Qinghua Liu, Nan Jiang, and Akshay Krishnamurthy (akshaykr.bsky.social). 1/11 1215
Reposted by Dylan Foster 🐢RL Theory Virtual Seminars @rl-theory.bsky.social · 16/04/2025Last seminars before the summer break: 04/29: Max Simchowitz (CMU) 05/06: Jeongyeol Kwon (Univ. of Widsconsin-Madison) 05/20: Sikata Sengupta & Marcel Hussing (Univ. of Pennsylvania) 05/27: Dhruv Rohatgi (MIT) 06/03: David Janz (Univ. of Oxford) 06/10: Nneka Okolo (MIT) 0145
Reposted by Dylan Foster 🐢Carlo Sferrazza @carlosferrazza.bsky.social · 17/04/2025What is the place of exploration in today's AI landscape and in which settings can exploration algorithms address current open challenges? Join us to discuss this at our exciting workshop at @icmlconf.bsky.social 2025: EXAIT! exait-workshop.github.io #ICML2025 193
Dylan Foster 🐢 @djfoster.bsky.social · 27/03/2025Reinforcement learning has led to amazing breakthroughs in reasoning (e.g., R1), but can it discover truly new behaviors not already present in the base model? A new paper with Zak Mhammedi and Dhruv Rohatgi: The Computational Role of the Base Model in Exploration arxiv.org/abs/2503.07453 14313
Reposted by Dylan Foster 🐢RL Theory Virtual Seminars @rl-theory.bsky.social · 24/03/2025Join us tomorrow to attend Vlad's presentation! Related to the seminar from last week, but this time in the offline setting. Tuesday March 25, 6 PM UTC. 031
Reposted by Dylan Foster 🐢augustychen.bsky.social @augustychen.bsky.social · 10/03/2025Excited to share new paper: Efficiently Escaping Saddle Points under Generalized Smoothness via Self-Bounding Regularity Link: arxiv.org/abs/2503.04712 Work with Karthik Sridharan and two great undergrads at Cornell, Daniel Yiming Cao and Benjamin Tang 1/8 121
Reposted by Dylan Foster 🐢arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 11/03/2025Dylan J. Foster, Zakaria Mhammedi, Dhruv Rohatgi: Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration arxiv.org/abs/2503.07453 arxiv.org/pdf/2503.07453 arxiv.org/html/2503.07453 155
Reposted by Dylan Foster 🐢let-all.com @let-all.com · 10/03/2025We have a new blog post on reinforcement learning theory for language model post training! By Akshay Krishnamurthy (@akshaykr.bsky.social) and Audrey Huang (@ahahaudrey.bsky.social)! www.let-all.com/blog/2025/03... 0113
Reposted by Dylan Foster 🐢Michele Guindani @mguindani.bsky.social · 08/03/2025Congratulations to @lestermackey.bsky.social for receiving the 2025 COPSS Award! 🎉👏 Lester is currently the Chair of the Section on Bayesian Statistical Sciences (SBSS) of the American Statistical Association. 2184
Dylan Foster 🐢 @djfoster.bsky.social · 23/02/2025Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier arxiv.org/abs/2502.12465 New paper (another fun internship project!) with Dhruv Rohatgi, Adam Block, Audrey Huang (ahahaudrey.bsky.social), and Akshay Krishnamurthy (akshaykr.bsky.social). 1/11 1272
Dylan Foster 🐢 @djfoster.bsky.social · 20/02/2025What are the minimal supervised learning primitives required to perform RL efficiently? New paper led by my amazing intern Dhruv Rohatgi: Necessary and Sufficient Oracles: Toward a Computational Taxonomy for Reinforcement Learning arxiv.org/abs/2502.08632 1/ 1254
Reposted by Dylan Foster 🐢arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 19/02/2025Dhruv Rohatgi, Adam Block, Audrey Huang, Akshay Krishnamurthy, Dylan J. Foster: Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning und... arxiv.org/abs/2502.12465 arxiv.org/pdf/2502.12465 arxiv.org/html/2502.12465 061
Reposted by Dylan Foster 🐢Mark Riedl @markriedl.bsky.social · 08/02/2025What kind of policy optimization is this? 3182
Reposted by Dylan Foster 🐢Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 24/01/2025One of the reasons LLMs make RL fun again is: 1) They make efficient exploration and low sample complexity genuinely matter 2) They give us good base policies so less time flailing around 5495
Reposted by Dylan Foster 🐢Maxim Raginsky @mraginsky.bsky.social · 22/01/2025It’s finally out — and I got to blurb it! 2949
Reposted by Dylan Foster 🐢Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 20/01/2025Finally finally finally some scaling curves for imitation learning in the large-scale-data regime: arxiv.org/abs/2411.04434arxiv.orgScaling Laws for Pre-training Agents and World ModelsThe performance of embodied agents has been shown to improve by increasing model parameters, dataset size, and compute. This has been demonstrated in domains from robotics to video games, when generat... 2548
Reposted by Dylan Foster 🐢Miro Dudik @mdudik.bsky.social · 16/01/2025📣My team at Microsoft Research New York is hiring a senior researcher in AI, both broadly in AI/ML, and in some specific areas including science of deep learning and modular transfer learning. Apply by February 7, 2025 on the link below. jobs.careers.microsoft.com/global/en/jo...jobs.careers.microsoft.comSearch Jobs | Microsoft Careers 0238
Reposted by Dylan Foster 🐢RL Theory Virtual Seminars @rl-theory.bsky.social · 05/01/2025Happy new year! Upcoming: 01/07: Jean Tarbouriech 01/14: Adrienne Tuynman 01/21: Victor Boone 02/18: Itai Sufaro 02/25: Jongmin Lee 03/04: Dhruv Rohatgi 03/11: David Cheikhi 03/18: Zakaria Mhammedi 03/25: Vlad Tkachuk 0101
Reposted by Dylan Foster 🐢Csaba Szepesvari @skiandsolve.bsky.social · 19/12/2024If you are into ML theory (RL or not) with a proven track record, and you are interested in an industry research position, PM me. Feel free to spread the word. 27531
Dylan Foster 🐢 @djfoster.bsky.social · 14/12/2024Given a high-quality verifier, language model accuracy can be improved by scaling inference-time compute (e.g., w/ repeated sampling). When can we expect similar gains without an external verifier? New paper: Self-Improvement in Language Models: The Sharpening Mechanism arxiv.org/abs/2412.01951 3416
Dylan Foster 🐢 @djfoster.bsky.social · 13/12/2024Preaching THE POWER OF RESETS with Zak! Poster 6402 in the West ballroom 0120
Reposted by Dylan Foster 🐢Microsoft Research @msftresearch.bsky.social · 13/12/2024In their 2024 NeurIPS paper on RL under latent dynamics, researchers examine whether existing algorithms designed for simple RL problems can be used to solve more complex RL problems. Dylan Foster discusses the modular approach his team explored. www.microsoft.com/en-us/resear... 061
Dylan Foster 🐢 @djfoster.bsky.social · 10/12/2024www.microsoft.com/en-us/resear... Did this short MSR podcast about our paper arxiv.org/abs/2410.17904 on modular approaches to RL with latent dynamics!microsoft.comAbstracts: NeurIPS 2024 with Dylan Foster - Microsoft ResearchIn their 2024 NeurIPS paper on RL under latent dynamics, researchers examine whether existing algorithms designed for simple RL problems can be used to solve more complex RL problems. Dylan Foster dis... 2120
Dylan Foster 🐢 @djfoster.bsky.social · 08/12/2024Heading to neurips! Stop by and say hi if you're around. 0371