Sign in

Daphne Cornelisse

@daphne-cornelisse.bsky.social
388 followers 54 following 21 posts

Multi-agent simulation & RL | www.daphne-cornelisse.com

PostsRepliesMedia
Reposted by Daphne Cornelisse
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 26/08/2026
I am begging you, stop using LLMs to write your papers it is horrible to read and causes an instant DNF.
71349
Daphne Cornelisse @daphne-cornelisse.bsky.social · 25/08/2026
btw, if you're working on molecular dynamics, drug design, or a related area and are curious about RL and high-performance simulation, please reach out! I'm looking for challenging problems in this space and would love to connect with people who bring domain expertise.
030
Daphne Cornelisse @daphne-cornelisse.bsky.social · 21/08/2026
What are the biggest bottlenecks to applying reinforcement learning and high-performance simulation in biology? In a new series, I’ll explore this question one domain at a time, starting with weakly electric fish. daphnecornelisse.substack.com/p/training-a...
0304
Reposted by Daphne Cornelisse
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 19/06/2026
The answer turns out to be that it's even less data than we thought we needed. You should use as much data as you can get your hands on but maybe, and this is still to be figured out, it's possible to identify norms/conventions/irrationalities from a little data once the base skills are there
2131
Reposted by Daphne Cornelisse
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 19/06/2026
This is part of a puzzle we've been trying to figure out for a while now: what is really the need for human data? We know we'd like our cars to not crash and obey some fairly simple constraints, so could we flip the script and mostly do RL with a little human fine-tuning on top?
1141
Reposted by Daphne Cornelisse
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 19/06/2026
New Paper: arxiv.org/abs/2606.19370 Self-play yields capabilities but requires frustrating cost-function tuning. Surprisingly, just 30 minutes of demonstration data produces much more human-like driving policies! Led by @daphne-cornelisse.bsky.social Website: spiced-self-play.com
110518
Daphne Cornelisse @daphne-cornelisse.bsky.social · 20/02/2026
Regularized self-play RL in grounded simulation effectively adapts driving policies to completely new cities. 🗽 -> 🗼 Really enjoyed collaborating on this work, led by Zilin and Saeed! Check out Zilin's post below for a great summary 🧵: x.com/nirhso/statu... 📄: arxiv.org/abs/2602.15891
0223
Daphne Cornelisse @daphne-cornelisse.bsky.social · 08/02/2026
The most important finding from this analysis! See the post for more details
061
Daphne Cornelisse @daphne-cornelisse.bsky.social · 30/12/2025
What if you could train agents on a 𝗱𝗲𝗰𝗮𝗱𝗲 of driving experience in 𝘂𝗻𝗱𝗲𝗿 𝗮𝗻 𝗵𝗼𝘂𝗿, on a single GPU? Excited to share 𝙋𝙪𝙛𝙛𝙚𝙧𝘿𝙧𝙞𝙫𝙚 2.0: A fast, friendly driving simulator with RL training via PufferLib at 𝟯𝟬𝟬𝗞 𝘀𝘁𝗲𝗽𝘀/𝘀𝗲𝗰 🐡 + 🚗 youtu.be/LfQ324R-cbE?...
youtu.be
PufferDrive 2.0 release
YouTube video by Daphne Cornelisse
35210
Reposted by Daphne Cornelisse
Mark Ho @markkho.bsky.social · 13/11/2025
Excited to share a new preprint, accepted as a spotlight at #NeurIPS2025! Humans are imperfect decision-makers, and autonomous systems should understand how we deviate from idealized rationality Our paper aims to address this! 👀🧠✨ arxiv.org/abs/2510.25951 a 🧵⤵️
arxiv.org
Estimating cognitive biases with attention-aware inverse planning
People's goal-directed behaviors are influenced by their cognitive biases, and autonomous systems that interact with people should be aware of this. For example, people's attention to objects in their...
16614
Daphne Cornelisse @daphne-cornelisse.bsky.social · 13/10/2025
Rapid RL experimentation is great. But how do you catch silent errors before they slip by? In this post, I share tools and habits that help me move quickly from idea to result without sacrificing reliability.
open.substack.com
How to catch subtle RL bugs before they catch you
Tools and habits for reliable, fast RL experimentation and development
0425
Reposted by Daphne Cornelisse
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 08/09/2025
The single biggest epistemic challenge in the internet era is remaining calibrated about what "normal" people think while the internet throws up an infinite wall of crazy. Thousands of people sharing an absurd opinion on the internet tells you very little!
812911
Daphne Cornelisse @daphne-cornelisse.bsky.social · 19/04/2025
Overnight runs are the overnight oats of research — prep, forget, and rewarding by morning
0144
Reposted by Daphne Cornelisse
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 12/03/2025
Building a "human-level" simulated driver that zero-shot generalizes to many benchmarks: a fun interview with @natolambert.bsky.social www.youtube.com/watch?v=2Q66...
youtube.com
Self-play for Self-driving and where Scaling Reinforcement Learning is Heading with Eugene Vinitsky
YouTube video by Interconnects AI
0183
Daphne Cornelisse @daphne-cornelisse.bsky.social · 28/02/2025
Sim agents are key for developing autonomous systems for safety-critical systems, like self-driving cars. We're open-sourcing sim agents that achieve a 99.8% success rate with < 0.8% failures on the Waymo Dataset. These agents are built through scaling self-play.
3335
Daphne Cornelisse @daphne-cornelisse.bsky.social · 20/02/2025
GPUDrive got accepted to ICLR 2025! With that, we release GPUDrive v0.4.0! 🚨 You can now install the repo and run your first fast PPO experiment in under 10 minutes. I’m honestly so excited about the new opportunities and research the sim makes possible. 🚀 1/2
2454
Reposted by Daphne Cornelisse
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 18/02/2025
A large group of us (spearheaded by Denizalp Goktas) have put out a position paper on paths towards foundation models for strategic decision-making. Language models still lack these capabilities so we'll need to build them: hal.science/hal-04925309...
2337