Steven Wu @zstevenwu.bsky.social · 12/03/2026There's growing evidence that LLMs can p-hack. But p-hacking also points to something bigger: a data science multiverse of defensible analytical choices. We wrote a paper (arxiv.org/abs/2602.18710) on using LLM agents to map this multiverse systematically. 🧵 1113
Reposted by Steven WuGokul Swamy @gokul.dev · 06/03/2025I was lucky enough to be invited give a talk on our new paper on the value of RL in fine-tuning at Cornell last week! Because of my poor time management skills, the talk isn't as polished as I'd like, but I think the "vibes" are accurate enough to share: youtu.be/E4b3cSirpsg.youtu.beAll Roads Lead to Likelihood: The Value of RL in Fine-TuningYouTube video by Gokul Swamy 0153
Reposted by Steven WuGokul Swamy @gokul.dev · 04/03/20251.5 yrs ago, we set out to answer a seemingly simple question: what are we *actually* getting out of RL in fine-tuning? I'm thrilled to share a pearl we found on the deepest dive of my PhD: the value of RL in RLHF seems to come from *generation-verification gaps*. Get ready to 🤿: 15911
Reposted by Steven WuMarc Lanctot @sharky6000.bsky.social · 21/11/2024@gswamy.bsky.social et al propose SPO which builds a game from a preferences, solving for the minimax winner. Handles non-Markovian, intransitive, and stochastic preferences. Nice empirical eval ranging from small demonstrative domains to huge RL domain (Mujoco). arxiv.org/abs/2401.04056 2/3.arxiv.orgA Minimaximalist Approach to Reinforcement Learning from Human FeedbackWe present Self-Play Preference Optimization (SPO), an algorithm for reinforcement learning from human feedback. Our approach is minimalist in that it does not require training a reward model nor unst... 2172
Reposted by Steven WuMarc Lanctot @sharky6000.bsky.social · 21/11/2024I have become a fan of the game-theoretic approaches to RLHF, so here are two more papers in that category! (with one more tomorrow 😅) 1. Self-Play Preference Optimization (SPO). 2. Direct Nash Optimization (DNO). 🧵 1/3. 2739