Sign in

Seth Karten

@sethkarten.ai
570 followers 1.5K following 231 posts

Autonomous Agents | Research @ Prime Intellect | PhD @ Princeton | Prev: CMU, Waymo | NSF GRFP Fellow sethkarten.ai

PostsRepliesMedia
Seth Karten @sethkarten.ai · 25/09/2026
We wrote Agent Bazaar back in May around a future where agents act on behalf of users and increasingly participate directly in marketplaces like Amazon and eBay. We introduced Economic Alignment to study what happens as agents become economic actors.
2153
Seth Karten @sethkarten.ai · 24/09/2026
prime agent v0.9.6 is out: ◆ Support for GPT-6 Sol, Opus 5.5, and Grok 4.7 ◆ /mcp plugin catalog with one-click connections to Linear, Notion, Posthog, Stripe, and 60+ more services ◆ Huge perf and reliability pass 🫡 Lots more coming soon :)
1263
Seth Karten @sethkarten.ai · 16/09/2026
prime agent v0.9.5 we fixed a lot of bugs and, of course, we had prime agent feature its favorite updates it picked our perf work. then it created the video itself.
3311
Seth Karten @sethkarten.ai · 22/07/2026
How are the academics feeling about this? Does it even change anything for profs?
120
Seth Karten @sethkarten.ai · 22/07/2026
Great to see Continual Harness acknowledged in Schmidhuber’s latest survey paper
050
Seth Karten @sethkarten.ai · 12/07/2026
Wow, three papers in CoLM 2026... Here I come San Francisco! These papers predicted some early trends in multi-agent safety & economic envs, automatic RL env creation, and PPO for VLMs/LLMs 🧵
130
Seth Karten @sethkarten.ai · 14/05/2026
New paper alert: Continual Harness: Online Adaptation for Self-Improving Foundation Agents Paper (arXiv). arxiv.org/abs/2605.09998 Article (Substack). sethkarten.substack.com/p/gemini-pla... Project page (video demos). sethkarten.ai/continual-ha...
0203
Seth Karten @sethkarten.ai · 26/04/2026
im very glad to see a rebound this year in the total number of NSF GRFP awards to exceed the most this century. Founding the next generation of American scientists is important to keep growth of the sciences.
030
Seth Karten @sethkarten.ai · 24/03/2026
1) what
210
Seth Karten @sethkarten.ai · 24/12/2025
I think I might leave bluesky tbh
110
Seth Karten @sethkarten.ai · 09/12/2025
Personally I am worried about this effect in disclosure
350
Seth Karten @sethkarten.ai · 24/11/2025
How do we close the gap between specialist RL and generalist LLM agents? We're benchmarking it in Pokémon. Join us at the PokeAgent Challenge competition workshop @ NeurIPS 2025. 📍 Dec 7, 8AM 🎮 Track 1: Competitive Pokémon (game-theoretic reasoning) 🗺️ Track 2: Speedrunning (long-horizon planning)
Flyer for The PokeAgent Challenge at NeurIPS 2025. Sunday, Dec 7, 8–10:45 AM PST, Mezzanine Room 15AB, San Diego Convention Center. Two tracks: Track 1 (Battling) features competitive Pokémon battle bots; Track 2 (Speedrunning) features long-horizon RPG gameplay. Tagline: "How do we close the gap between specialist RL models and generalist LLM agents?" Speakers: Seth Karten (Princeton), Aaron Traylor, Minmin Chen (Google DeepMind), Jake Grigsby (UT Austin), Stephanie Milani (NYU/Johns Hopkins), Kiran Vodrahalli (Google DeepMind), Fei Fang (CMU), Yuke Zhu (UT Austin), Chi Jin (Princeton). Sponsored by Google DeepMind.
043
Seth Karten @sethkarten.ai · 20/10/2025
In the NeurIPS PokeAgent Challenge, we stress-test 4 ranking systems across (100k+ agent matches): - Bradley-terry (batch MLE, our ground truth) - Elo (online, chess-standard) - Glicko-1 (online, uncertainty-aware) - GXE: (Glicko-derived win %) (2/5)
Leaderboard of Pokemon Gen 1 OU Top 100 NeurIPS competition for the PokeAgent Challenge. The leaderboard shows username, elo, glicko-1, glicko-1 deviation, wins, losses, and ties for the results of the head to head battles for each agent methodology. Highlighted are top user submissions. PAC-MM-* usernames are organizer hosted baselines.Leaderboard of Pokemon Gen 1 OU Top 100 NeurIPS competition for the PokeAgent Challenge on the pokeagent.github.io website. The leaderboard shows username, history rating, GXE, wins, losses for the results of the head to head battles for each agent methodology, including showing the currently qualifying methods.
110
Seth Karten @sethkarten.ai · 15/10/2025
A benchmark environment is nothing without data so you can pretrain before you RL. Announcing our replay archive preview: We are releasing an additional 25k games to help you train a metagame exploiter (5 million more released after qualifier) replays.pokeagentshowdown. com:8443/ (3/3)
Two pokeagents in the replay archives
010
Seth Karten @sethkarten.ai · 15/10/2025
Pokemon is truly the pareto frontier of agent research - The RPG requires an autonomous embodied agentic agent with perception, planning, memory, and control - VGC and Gen 9 OU penalize erroneous actions with fast-paced opponent-modeling in short games (1/3)
191
Seth Karten @sethkarten.ai · 24/09/2025
If you arent paying attention, we are in a rapidly shifting period of ML paper culture. ICLR/ICML/NeurIPS are being treated as random, out of touch processes with more and more unnecessary work to submit Most people are saying TMLR is the only good alternative, but are skeptical
010
Seth Karten @sethkarten.ai · 02/09/2025
🚨 Hackathon Weekend! 🚨 Jumpstart your PokéAgent Challenge submission ahead of NeurIPS! 📅 Sept 13–14 ✅ Leaderboards reset Sat 10AM EDT 🎙️ Lightning talks in LLMs, RL, and Pokemon 💬 Live Office hours 🏆 $2k in prizes
PokéAgent Challenge @ NeurIPS 2025 Hackathon Weekend Schedule. Saturday, Sept 13th: 10 AM leaderboards reset; 12–1:30 PM livestream talks (overview, Aaron Traylor on Pokémon as an AI Problem, Seth Karten on Pokéchamp, Jake Grigsby on Metamon, plus more). Sunday, Sept 14th: 1–3:30 PM organizer office hours; 11:59 PM top teams earn up to $2k in GCP credits. Sponsored by Google DeepMind and AIJ.
172
Seth Karten @sethkarten.ai · 20/08/2025
The solution would generalize to another two player partially observable turn-based text game. The most bespoke items are tools, but there has been work recently that shows that you can make these tools modular LLM calls, further increasing generality
010
Seth Karten @sethkarten.ai · 12/08/2025
010
Seth Karten @sethkarten.ai · 12/08/2025
Papers are dead. Maybe it is time to start the youtube channel…
120
Seth Karten @sethkarten.ai · 23/07/2025
Democratic alignment: in a special case, periodic citizen voting can fire the planner. Leader turnover keeps welfare high and prevents policy drift—central nudging plus decentralized oversight in one sandbox.
Timeline plot: planner changes each tax year; welfare remains elevated across turnovers.
120
Seth Karten @sethkarten.ai · 23/07/2025
Centralized nudging: the planner’s marginal taxes beat U.S. statutory rates and approach Saez on aggregate welfare (almost double vs baseline).
Bar chart comparing social welfare for U.S. baseline, LLM planner, Saez schedule; LLM nearly matches Saez.
120
Seth Karten @sethkarten.ai · 23/07/2025
Synthetic behavioral policies → we sample workers from 2023 ACS skills & demographics, then let each agent verify its own bounded rational utility from individualized preferences, enabling counterfactual reasoning.
120
Seth Karten @sethkarten.ai · 23/07/2025
🚀 New preprint! 🤔 Can one agent “nudge” a synthetic civilization of Census‑grounded agents toward higher social welfare—all by optimizing utilities in‑context? Meet the LLM Economist ↓
Diagram of LLM Economist: left—grid of persona‑conditioned worker agents; center—planner LLM sends tax schedule; right—social‑welfare ‘hill‑climb’.
194
Seth Karten @sethkarten.ai · 18/07/2025
Open review doesnt seem public yet but here are the titles
110
Seth Karten @sethkarten.ai · 14/07/2025
🚀 Launch day! The NeurIPS 2025 PokéAgent Challenge is live. @neuripsconf.bsky.social Two tracks: ① Showdown Battling – imperfect-info, turn-based strategy ② Pokemon Emerald Speedrunning – long horizon RPG planning 5 M labeled replays • starter kit • baselines. Bring your LLM, RL, or hybrid agent!
Banner reading “PokéAgent Challenge @ NeurIPS 2025” with two panels: Track 1 – Competitive Pokémon Battle Bots, Track 2 – Long-Horizon RPG Gameplay. Call-to-action: “Create video-game AI! Win prizes! Live now at pokeagent.github.io.”
185
Seth Karten @sethkarten.ai · 12/07/2025
🚀 5 days until my ICML spotlight poster! Key insights we’ll unpack: • Base LLM + test-time planning • Game-theoretic scaffolding • Context-engineered opponent prediction • Comparative LLM-as-judge (relative > absolute) Catch me Thu Jul 17, 4:30-7 PM PT👇
141
Seth Karten @sethkarten.ai · 04/06/2025
Social media takeoff is hard. Bluesky still lacks the capability to compete with twitter
110
Seth Karten @sethkarten.ai · 30/05/2025
Excited to announce that I will be spending the summer at @Waymo on the simulation realism team! I’ll be working on learning to generate simulated worlds. 🚙🚙🚙 Send me a message if youre in the bay area and want to chat!
290
Seth Karten @sethkarten.ai · 26/05/2025
Excited to share that the PokeAgent challenge was accepted as a NeurIPS competition! This should serve as an excellent benchmark for competitive games AND ‘speedrunning’ the RPG. I hope to see both the RL and LLM agent communities working together here to eval agents in Pokemon More info soon👀
NeurIPS 2025 competition track submission summary, scores, and recommendation to accept
4183
Seth Karten @sethkarten.ai · 30/04/2025
What happens to TRI though? I thought they had an AV division. Also Toyotas arent EVs so I am confused how the driving tech stack would work. I think Waymo is just diversifying their risk into personal vehicles
130
Seth Karten @sethkarten.ai · 28/04/2025
Insane new study from zurich studies the influence of the LLMs for persuasion on the r/ChangeMyView subreddit. Let's just say people are outraged... Is the study justified since bots are already rampant on reddit? Or does this cross ethical lines?
110
Seth Karten @sethkarten.ai · 07/03/2025
Stay tuned for our upcoming dataset release (3M+ ranked human games)!
130
Seth Karten @sethkarten.ai · 07/03/2025
How does PokéChamp perform in competitive Pokémon battles? Our agent, powered by GPT-4, achieves a win rate of 84% against Abyssal (best rule-based bot) and a local Elo of 1268, outperforming all baselines, including other LLM-based agents and traditional methods!
130
Seth Karten @sethkarten.ai · 07/03/2025
PokéChamp uses an LLM to decide to use one-step lookahead tools for domain-specific calculations (e.g., damage calcs) or a small-scale minimax search enhanced by action sampling, opponent modeling, and value function estimation. Result: Expert-level play at human speed!
140
Seth Karten @sethkarten.ai · 07/03/2025
Why Pokémon? It's the perfect testbed for LLM agents: constantly changing ruleset, partial observability, and rich strategic depth. PokéChamp leverages an LLM for action sampling, opponent modeling and value estimation, with no domain-specific training required
140
Seth Karten @sethkarten.ai · 07/03/2025
Can a Large Language Model (LLM) with zero Pokémon-specific training achieve expert-level performance in competitive Pokémon battles? Introducing PokéChamp, our minimax LLM agent that reaches top 30%-10% human-level Elo on Pokémon Showdown! New paper on arXiv and code on github!
1335
Seth Karten @sethkarten.ai · 26/02/2025
We are not out of the woods yet... Looks like NVIDIA 5090 stockx prices are starting to stagnate. Likely to be over 4 weeks until under $100 per GB of VRAM unless we are blessed with significantly more stock
210
Seth Karten @sethkarten.ai · 20/02/2025
Based on stockx resales, NVIDIA 5090 prices might normalize in 2 weeks time
010
Seth Karten @sethkarten.ai · 13/02/2025
Is GRPO sufficient? Let's look at the OG AlphaGo paper --> Policy network: yes, from LLM Rollouts: yes, from advantage calculation Value network: no Perhaps we are in the REINFORCE era of RL for LLMs still
220
Seth Karten @sethkarten.ai · 07/02/2025
sometimes it takes more than machine learning to train a model
130
Seth Karten @sethkarten.ai · 30/01/2025
i just need to freeze experiments and reduce to 8 pages
020
Seth Karten @sethkarten.ai · 30/01/2025
Yep no 5090 for me
130
Seth Karten @sethkarten.ai · 26/01/2025
🧠DeepSeek's R1 report highlights challenges in applying AlphaZero to LLMs. Gaming AI offers a key insight: AlphaStar, like an LLM, starts with imitation learning, then RL, and crucially, league self-play. League diversity helps avoid local optima. How do we implement this? ⬇️
120
Seth Karten @sethkarten.ai · 14/01/2025
keep it 💯
030
Seth Karten @sethkarten.ai · 12/01/2025
A large part of it is detecting them early (not mine but from my prior lab). Then, I think the LA fires are particularly challenging since there must be some optimality in preventing the further spread of widespread fires
010
Seth Karten @sethkarten.ai · 06/01/2025
33K FPS at what cost? My holiday 🤖
130
Seth Karten @sethkarten.ai · 09/12/2024
I'll be running live demos Wed-Sat! Challenge PokéChamp and test your #Pokemon battle skills against our AI.
020
Seth Karten @sethkarten.ai · 09/12/2024
We combine LLMs with minimax search, opponent modeling, and tool-use to achieve expert-level performance in competitive Pokémon battles, ranking in the top 10% of players. Check out the preprint: sethkarten.ai/data/PokeCha... #ResearchPaper #AIResearch
120
Seth Karten @sethkarten.ai · 09/12/2024
I am attending #NeurIPS2024 Wed-Sun! 🎉 I will be presenting "PokéChamp: Expert-level Minimax Language Agent for Competitive Pokémon" at the Language Gamification workshop on Sat. Feel free to message me for coffee chats & party invites! #AIinGaming
190