Sign in

Gokul Swamy

@gokul.dev
4.4K followers 417 following 109 posts

PhD student at @cmurobotics.bsky.social working on efficient algorithms for interactive learning (e.g. imitation / RL / RLHF). no model is an island. prefers email. gokul.dev. on the job market!

PostsRepliesMedia
Gokul Swamy @gokul.dev · 26/11/2025
Couldn't think of a better excuse than #neurips2025 to visit my home town of San Diego! I'll be around 12/3-12/6 -- shoot me a DM / email if you'd like to catch up!
010
Reposted by Gokul Swamy
Natalie Collina @ncollina.bsky.social · 22/10/2025
Our paper on algorithmic collusion was featured in a Quanta article! www.quantamagazine.org/the-game-the...
quantamagazine.org
The Game Theory of How Algorithms Can Drive Up Prices | Quanta Magazine
Recent findings reveal that even simple pricing algorithms can make things more expensive.
2309
Gokul Swamy @gokul.dev · 22/10/2025
Woooaaaah :O
010
Gokul Swamy @gokul.dev · 22/10/2025
Just discovered this lovely talk from my favorite professor from undergrad, who is still as inspiring as I remember him: www.youtube.com/watch?v=XLZ0...
youtube.com
“Raising Our Sights (A Long Rant From an Accidental Engineer)”, Scott Shenker
YouTube video by Plamadiso - Platforms, Markets & Digital Society
000
Gokul Swamy @gokul.dev · 13/10/2025
Please apply or help out!
010
Gokul Swamy @gokul.dev · 28/09/2025
Late, but arxiv.org/abs/0804.2996 is *incredible*, so many good lines (e.g., "This comes close to being an accusation of a false claim of priority for a false discovery of an untrue fact, which would be a rare triple-negative in the history of intellectual property disputes.").
arxiv.org
The Epic Story of Maximum Likelihood
At a superficial level, the idea of maximum likelihood must be prehistoric: early hunters and gatherers may not have used the words ``method of maximum likelihood'' to describe their choice of where a...
0113
Gokul Swamy @gokul.dev · 23/08/2025
I've been really enjoying the new Ninajirachi album -- it's very Boiler Room-core :)
110
Gokul Swamy @gokul.dev · 23/08/2025
Thanks for the shout-out and I hope the lectures were at least somewhat understandable! Yeah, once things settle down a bit for me, I'd like to more deeply understand the connection between Rust's structural estimation and IRL as I conceive of it.
120
Gokul Swamy @gokul.dev · 15/07/2025
We therefore advocate for caution when making or evaluating claims about LLM reasoning and beyond with GRPO and PPO, ideally using algorithms like RLoo or REBEL instead. Check out our blog post for links to our code and W&B logs if you'd like to reproduce our experiments.
010
Gokul Swamy @gokul.dev · 15/07/2025
While this worked out for the better on some seeds, it doesn't have to in general. After all, an algorithm that behaves unexpectedly *well* in one setting can perform unexpectedly *poorly* in another, perhaps more important, setting.
110
Gokul Swamy @gokul.dev · 15/07/2025
We see similar results on a didactic bandit problem -- i.e. a problem that has nothing to do with LLMs or reasoning! This implies that PPO / GRPO are fundamentally *not* following the true policy gradient.
110
Gokul Swamy @gokul.dev · 15/07/2025
We find that RLoo (an unbiased estimate of the vanilla PG) and REBEL (a regression-based approximation of online mirror descent) preserve performance as expected. In contrast, algorithms like PPO / GRPO that include heuristics (e.g. clipping) show a marked and unexpected change in performance.
110
Gokul Swamy @gokul.dev · 15/07/2025
So, with a truly random reward function, all policies look equally good. Thus, the *true* policy gradient is zero, as the initial policy is optimal by construction. So, we'd expect performance to flatline. We use random rewards as a *diagnostic task* to compare different RL algs.
110
Gokul Swamy @gokul.dev · 15/07/2025
Lead by Owen Oertell & Wenhao Zhan, joint w/ Steven Wu, Kiante Brantley, Jason Lee, and Wen Sun. If a project has got Wen, Owen, Wenhao, and Qwen on it, you know it's gotta be good 😛.
110
Gokul Swamy @gokul.dev · 15/07/2025
Recent work has seemed somewhat magical: how can RL with *random* rewards make LLMs reason? We pull back the curtain on these claims and find out this unexpected behavior hinges on the inclusion of certain *heuristics* in the RL algorithm. Our blog post: tinyurl.com/heuristics-c...
tinyurl.com
Heuristics Considered Harmful: RL With Random Rewards Should Not Make LLMs Reason | Notion
Owen Oertell*, Wenhao Zhao*, Gokul Swamy, Zhiwei Steven Wu, Kiante Brantley, Jason Lee, Wen Sun
162
Reposted by Gokul Swamy
quokkka.bsky.social @quokkka.bsky.social · 20/06/2025
very nice lectures, watch them from time to time
061
Reposted by Gokul Swamy
Marc Lanctot @sharky6000.bsky.social · 20/06/2025
Want to learn about online learning, game solving, RL, imitation learning with applications to robotics, and RLHF with applications to language modeling? Check out this course! 👍
061
Gokul Swamy @gokul.dev · 20/06/2025
While I can't promise everything will be crystal-clear after going though the lectures (especially because of my handwriting :p), I hope that if nothing else, you can tell how beautiful we all find these ideas. If that feeling comes across, I'll feel like I have succeeded! :)
130
Gokul Swamy @gokul.dev · 20/06/2025
The second was being able to teach this course with my amazing advisors, Drew Bagnell and Steven Wu -- the folks I learned all of this stuff from. Fun fact: because of parking fees, Drew actually *paid* to lecture. And I'm always grateful to ZSW for pushing me out of the nest.
130
Gokul Swamy @gokul.dev · 20/06/2025
Two other things made this course particularly special. The first was the students and their *incredible* questions -- there were so many times where I was like wow, it took me *YEARS* before I realized that was the right question to be asking.
140
Gokul Swamy @gokul.dev · 20/06/2025
We also had wonderful guest lectures from Yuda Song on hybrid RL (youtu.be/1B2XGXQ2hfA), Sanjiban Choudhury on scaling imitation (youtu.be/KnXSeTuCgFI), and Wen Sun on RLHF algorithms (youtu.be/qdkBZJywi_4).
youtu.be
Algorithmic Foundations of Interactive Learning SP25: Lecture 17
YouTube video by Gokul Swamy
160
Gokul Swamy @gokul.dev · 20/06/2025
My favorite lectures to give were on the value of interaction in imitation / RLHF! youtu.be/uESAXg-CXFs, youtu.be/N8-Nh_iTmps, youtu.be/qHvB30J5gyo, youtu.be/ZzFjoH47GIg. It took 5 years, but I finally have an answer at least I find compelling :p.
youtu.be
Algorithmic Foundations of Interactive Learning SP25: Lecture 19
YouTube video by Gokul Swamy
160
Gokul Swamy @gokul.dev · 20/06/2025
To do so, we worked backwards from things like ChatGPT and RMA and "backed out" a "dependency graph". We then did a "forward pass" over the semester, going from online learning, to game solving, to core RL, to imitation learning / robot learning, to RLHF / LLM fine-tuning.
150
Gokul Swamy @gokul.dev · 20/06/2025
I think in a field as fast-paced as machine learning, a good course gives students a conceptual framework for understanding new developments quickly + what is actually "new" vs. the classical algorithms. We also wanted to explain *when* scale isn't "all you need."
160
Gokul Swamy @gokul.dev · 20/06/2025
You can access all the content here: Course Website: interactive-learning-algos.github.io Lecture Playlist: youtube.com/playlist?lis... Scribe Notes "Book": interactive-learning-algos.github.io/assets/pdfs/.... Homeworks / class competition material are also public!
interactive-learning-algos.github.io
Home
Website for AFIL course.
1141
Gokul Swamy @gokul.dev · 20/06/2025
It was a dream come true to teach the course I wish existed at the start of my PhD. We built up the algorithmic foundations of modern-day RL, imitation learning, and RLHF, going deeper than the usual "grab bag of tricks". All 25 lectures + 150 pages of notes are now public!
34710
Gokul Swamy @gokul.dev · 13/06/2025
Shortcut models enable scaling offline RL, both at train-time at test-time! We beat so many other algorithms on so many tasks we had to stick most of the results in the appendix 😅. Very proud of @nico-espinosa-dice.bsky.social for spearheading this project, check out his thread!
011
Gokul Swamy @gokul.dev · 27/04/2025
Boston friends: I'll be in the Cambridge area for the next few days, shoot me a message if you'd like to catch up :).
010
Gokul Swamy @gokul.dev · 22/04/2025
I won't be at #ICLR2025 myself this time around but please go talk to lead authors Nico, Zhaolin, and Runzhe about their bleeding-edge algorithms for imitation learning and RLHF!
030
Gokul Swamy @gokul.dev · 07/04/2025
As always, I'm incredibly grateful to Wen / Sanjiban for letting me borrow their excellent students for a bit to work on my harebrained schemes. Full paper at arxiv.org/abs/2503.13162. [17/n, n=17]
arxiv.org
Efficient Imitation under Misspecification
We consider the problem of imitation learning under misspecification: settings where the learner is fundamentally unable to replicate expert behavior everywhere. This is often true in practice due to ...
000
Gokul Swamy @gokul.dev · 07/04/2025
I am also incredibly grateful to @nico-espinosa-dice.bsky.social for sticking with me through the many iterations it took us to figure out what the "right questions" were to ask here. Nico basically rewrote the whole paper / did a new set of experiments *after* it had been accepted 😱. [16/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
There's a bunch of other fancy theory in the paper but at the highest level, one reason I am proud of this paper is that I finally understand how we can close embodiment / sensory gaps without needing a queryable expert/ global RL, which should making scaling IL easier. [15/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
We then show that this sort of suboptimal, offline data can help significantly with speeding up local search on challenging maze-based exploration problems where the learner needs to act meaningfully different from the expert (our method in orange). [14/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
Among our contributions, we give a precise (and rather aesthetically pleasingly typeset if I don't say so myself) theoretical condition under which the use of suboptimal data to help with figuring out *where* to locally search from is helps with policy performance! [13/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
Intuitively, we likely want to explore from a "wider" set of states that covers policies we can actually choose, not just the unrealizable "tightrope-walking" expert. In practice, we can use suboptimal / failure / play data to "broaden" where we perform local search. [12/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
Second, we may be unable to "follow-through" the same way the expert does (e.g. insert a peg even when its close to the hole due to no haptic feedback). Together, this begs the question of *where* we should actually perform local search from. [11/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
First, we might be unable to actually reach the expert states in the first place (e.g. a humanoid might not be able to do the first half of a backflip). [10/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
However, in the misspecified setting where we can't imitate the expert in the first place, it's much less clear that locally exploring from *expert* states on "tightrope" like problems is the right thing to do. More explicitly, this is basically for two reasons. [9/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
(Side note: in practice, it's often best to do this *inside* a learned world model -- you can't reset a humanoid halfway through a backflip. Check out gokul.dev/hyper/ for the relevant theory, x.com/YilinWu11/st... for a particularly cool VLM-based instantiation!) [8/n]
gokul.dev
Hybrid Inverse Reinforcement Learning
We present a general theoretical framework for designing efficient algorithms for interactive imitation learning. We leverage this framework to derive novel inverse reinforcement learning ...
100
Gokul Swamy @gokul.dev · 07/04/2025
Some of my prior work (gokul.dev/filter/) showed that "local search" -- exploring just from expert states -- is "all you need" in the well-specified setting -- i.e. when you can *perfectly* imitate the expert. [7/n]
gokul.dev
Inverse Reinforcement Learning without Reinforcement Learning
Inverse reinforcement learning approaches reduce the problem of imitation learning to repeatedly solving a computationally expensive reinforcement learning problem. We prove and empirically validate t...
100
Gokul Swamy @gokul.dev · 07/04/2025
Ideally, we'd avoid needing to query the expert in the loop (i.e. DAgger) or repeatedly solve a global search / RL problem (i.e. inverse RL) as both hard to scale. [6/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
The downstream affect of misspecification is that the policy you're learning must make *some* mistakes. The question then becomes how can we *efficiently* avoid the most *costly* of these mistakes (i.e. those that lead to "compounding errors"). [5/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
We also introduce misspecification when we distill actions from a privileged expert (e.g. the sort of RMA-style thing that underlies many of the most impressive recent results in quadruped locomotion) -- here, misspecification comes from partial observability. [4/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
Ok, so, why should you care about "misspecification in imitation" other than the fact it rhymes? Well, for problems like motion retargeting for humanoids, one needs to account for the fact that most people don't look like UniTree G1 / NEO, making naive action transfer hard. [3/n]
100
Gokul Swamy @gokul.dev · 07/04/2025
Check out arxiv.org/abs/2503.13162 for the full paper! This is a truly outstanding first paper of a PhD for @nico-espinosa-dice.bsky.social. It is also my first (and hopefully not last) last author paper :p. Joint work w/ Sanjiban and Wen. [2/n]
arxiv.org
Efficient Imitation under Misspecification
We consider the problem of imitation learning under misspecification: settings where the learner is fundamentally unable to replicate expert behavior everywhere. This is often true in practice due to ...
100
Gokul Swamy @gokul.dev · 07/04/2025
I think of misspecification (e.g. embodiment / sensory gaps) as the fundamental reason behavioral cloning isn't "all you need" for imitation as matching actions matching outcomes. Introducing @nico-espinosa-dice.bsky.social's #ICLR2025 paper proving that "local search" *is* all you need! [1/n]
140
Gokul Swamy @gokul.dev · 04/04/2025
(I should also say that I think writing about the good times in a way that doesn't come across as generic is a fundamentally harder skill imo -- "All happy families are alike; each unhappy family is unhappy in its own way" as Tolstoy opens Anna Karenina.)
010
Gokul Swamy @gokul.dev · 04/04/2025
Honestly, way less impressed with it than her last album. I feel like writing incisively about happy moments is a skill she hasn't totally mastered yet. I did like "Best Guess" though!
100
Gokul Swamy @gokul.dev · 06/03/2025
I was lucky enough to be invited give a talk on our new paper on the value of RL in fine-tuning at Cornell last week! Because of my poor time management skills, the talk isn't as polished as I'd like, but I think the "vibes" are accurate enough to share: youtu.be/E4b3cSirpsg.
youtu.be
All Roads Lead to Likelihood: The Value of RL in Fine-Tuning
YouTube video by Gokul Swamy
0153
Gokul Swamy @gokul.dev · 04/03/2025
This project is easily the hardest thing I've worked on. It's also the project I'm proudest of. I am very, very grateful to have advisors like Drew & Steven who let me prioritize deep thought even in today's research climate. Check out arxiv.org/abs/2503.01067 for more. [17/17]
arxiv.org
All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning
From a first-principles perspective, it may seem odd that the strongest results in foundation model fine-tuning (FT) are achieved via a relatively complex, two-stage training procedure. Specifically, ...
0100