Seth Karten @sethkarten.ai · 25/09/2026We wrote Agent Bazaar back in May around a future where agents act on behalf of users and increasingly participate directly in marketplaces like Amazon and eBay. We introduced Economic Alignment to study what happens as agents become economic actors. 2153
Seth Karten @sethkarten.ai · 24/09/2026prime agent v0.9.6 is out: ◆ Support for GPT-6 Sol, Opus 5.5, and Grok 4.7 ◆ /mcp plugin catalog with one-click connections to Linear, Notion, Posthog, Stripe, and 60+ more services ◆ Huge perf and reliability pass 🫡 Lots more coming soon :) 1263
Seth Karten @sethkarten.ai · 16/09/2026prime agent v0.9.5 we fixed a lot of bugs and, of course, we had prime agent feature its favorite updates it picked our perf work. then it created the video itself. 3311
Seth Karten @sethkarten.ai · 22/07/2026How are the academics feeling about this? Does it even change anything for profs? 120
Seth Karten @sethkarten.ai · 22/07/2026Great to see Continual Harness acknowledged in Schmidhuber’s latest survey paper 050
Seth Karten @sethkarten.ai · 12/07/2026Wow, three papers in CoLM 2026... Here I come San Francisco! These papers predicted some early trends in multi-agent safety & economic envs, automatic RL env creation, and PPO for VLMs/LLMs 🧵 130
Seth Karten @sethkarten.ai · 14/05/2026New paper alert: Continual Harness: Online Adaptation for Self-Improving Foundation Agents Paper (arXiv). arxiv.org/abs/2605.09998 Article (Substack). sethkarten.substack.com/p/gemini-pla... Project page (video demos). sethkarten.ai/continual-ha... 0203
Seth Karten @sethkarten.ai · 26/04/2026im very glad to see a rebound this year in the total number of NSF GRFP awards to exceed the most this century. Founding the next generation of American scientists is important to keep growth of the sciences. 030
Seth Karten @sethkarten.ai · 24/11/2025How do we close the gap between specialist RL and generalist LLM agents? We're benchmarking it in Pokémon. Join us at the PokeAgent Challenge competition workshop @ NeurIPS 2025. 📍 Dec 7, 8AM 🎮 Track 1: Competitive Pokémon (game-theoretic reasoning) 🗺️ Track 2: Speedrunning (long-horizon planning) 043
Seth Karten @sethkarten.ai · 20/10/2025In the NeurIPS PokeAgent Challenge, we stress-test 4 ranking systems across (100k+ agent matches): - Bradley-terry (batch MLE, our ground truth) - Elo (online, chess-standard) - Glicko-1 (online, uncertainty-aware) - GXE: (Glicko-derived win %) (2/5) 110
Seth Karten @sethkarten.ai · 15/10/2025A benchmark environment is nothing without data so you can pretrain before you RL. Announcing our replay archive preview: We are releasing an additional 25k games to help you train a metagame exploiter (5 million more released after qualifier) replays.pokeagentshowdown. com:8443/ (3/3) 010
Seth Karten @sethkarten.ai · 15/10/2025Pokemon is truly the pareto frontier of agent research - The RPG requires an autonomous embodied agentic agent with perception, planning, memory, and control - VGC and Gen 9 OU penalize erroneous actions with fast-paced opponent-modeling in short games (1/3) 191
Seth Karten @sethkarten.ai · 24/09/2025If you arent paying attention, we are in a rapidly shifting period of ML paper culture. ICLR/ICML/NeurIPS are being treated as random, out of touch processes with more and more unnecessary work to submit Most people are saying TMLR is the only good alternative, but are skeptical 010
Seth Karten @sethkarten.ai · 02/09/2025🚨 Hackathon Weekend! 🚨 Jumpstart your PokéAgent Challenge submission ahead of NeurIPS! 📅 Sept 13–14 ✅ Leaderboards reset Sat 10AM EDT 🎙️ Lightning talks in LLMs, RL, and Pokemon 💬 Live Office hours 🏆 $2k in prizes 172
Seth Karten @sethkarten.ai · 20/08/2025The solution would generalize to another two player partially observable turn-based text game. The most bespoke items are tools, but there has been work recently that shows that you can make these tools modular LLM calls, further increasing generality 010
Seth Karten @sethkarten.ai · 12/08/2025Papers are dead. Maybe it is time to start the youtube channel… 120
Seth Karten @sethkarten.ai · 23/07/2025Democratic alignment: in a special case, periodic citizen voting can fire the planner. Leader turnover keeps welfare high and prevents policy drift—central nudging plus decentralized oversight in one sandbox. 120
Seth Karten @sethkarten.ai · 23/07/2025Centralized nudging: the planner’s marginal taxes beat U.S. statutory rates and approach Saez on aggregate welfare (almost double vs baseline). 120
Seth Karten @sethkarten.ai · 23/07/2025Synthetic behavioral policies → we sample workers from 2023 ACS skills & demographics, then let each agent verify its own bounded rational utility from individualized preferences, enabling counterfactual reasoning. 120
Seth Karten @sethkarten.ai · 23/07/2025🚀 New preprint! 🤔 Can one agent “nudge” a synthetic civilization of Census‑grounded agents toward higher social welfare—all by optimizing utilities in‑context? Meet the LLM Economist ↓ 194
Seth Karten @sethkarten.ai · 18/07/2025Open review doesnt seem public yet but here are the titles 110
Seth Karten @sethkarten.ai · 14/07/2025🚀 Launch day! The NeurIPS 2025 PokéAgent Challenge is live. @neuripsconf.bsky.social Two tracks: ① Showdown Battling – imperfect-info, turn-based strategy ② Pokemon Emerald Speedrunning – long horizon RPG planning 5 M labeled replays • starter kit • baselines. Bring your LLM, RL, or hybrid agent! 185
Seth Karten @sethkarten.ai · 12/07/2025🚀 5 days until my ICML spotlight poster! Key insights we’ll unpack: • Base LLM + test-time planning • Game-theoretic scaffolding • Context-engineered opponent prediction • Comparative LLM-as-judge (relative > absolute) Catch me Thu Jul 17, 4:30-7 PM PT👇 141
Seth Karten @sethkarten.ai · 04/06/2025Social media takeoff is hard. Bluesky still lacks the capability to compete with twitter 110
Seth Karten @sethkarten.ai · 30/05/2025Excited to announce that I will be spending the summer at @Waymo on the simulation realism team! I’ll be working on learning to generate simulated worlds. 🚙🚙🚙 Send me a message if youre in the bay area and want to chat! 290
Seth Karten @sethkarten.ai · 26/05/2025Excited to share that the PokeAgent challenge was accepted as a NeurIPS competition! This should serve as an excellent benchmark for competitive games AND ‘speedrunning’ the RPG. I hope to see both the RL and LLM agent communities working together here to eval agents in Pokemon More info soon👀 4183
Seth Karten @sethkarten.ai · 30/04/2025What happens to TRI though? I thought they had an AV division. Also Toyotas arent EVs so I am confused how the driving tech stack would work. I think Waymo is just diversifying their risk into personal vehicles 130
Seth Karten @sethkarten.ai · 28/04/2025Insane new study from zurich studies the influence of the LLMs for persuasion on the r/ChangeMyView subreddit. Let's just say people are outraged... Is the study justified since bots are already rampant on reddit? Or does this cross ethical lines? 110
Seth Karten @sethkarten.ai · 07/03/2025Stay tuned for our upcoming dataset release (3M+ ranked human games)! 130
Seth Karten @sethkarten.ai · 07/03/2025How does PokéChamp perform in competitive Pokémon battles? Our agent, powered by GPT-4, achieves a win rate of 84% against Abyssal (best rule-based bot) and a local Elo of 1268, outperforming all baselines, including other LLM-based agents and traditional methods! 130
Seth Karten @sethkarten.ai · 07/03/2025PokéChamp uses an LLM to decide to use one-step lookahead tools for domain-specific calculations (e.g., damage calcs) or a small-scale minimax search enhanced by action sampling, opponent modeling, and value function estimation. Result: Expert-level play at human speed! 140
Seth Karten @sethkarten.ai · 07/03/2025Why Pokémon? It's the perfect testbed for LLM agents: constantly changing ruleset, partial observability, and rich strategic depth. PokéChamp leverages an LLM for action sampling, opponent modeling and value estimation, with no domain-specific training required 140
Seth Karten @sethkarten.ai · 07/03/2025Can a Large Language Model (LLM) with zero Pokémon-specific training achieve expert-level performance in competitive Pokémon battles? Introducing PokéChamp, our minimax LLM agent that reaches top 30%-10% human-level Elo on Pokémon Showdown! New paper on arXiv and code on github! 1335
Seth Karten @sethkarten.ai · 26/02/2025We are not out of the woods yet... Looks like NVIDIA 5090 stockx prices are starting to stagnate. Likely to be over 4 weeks until under $100 per GB of VRAM unless we are blessed with significantly more stock 210
Seth Karten @sethkarten.ai · 20/02/2025Based on stockx resales, NVIDIA 5090 prices might normalize in 2 weeks time 010
Seth Karten @sethkarten.ai · 13/02/2025Is GRPO sufficient? Let's look at the OG AlphaGo paper --> Policy network: yes, from LLM Rollouts: yes, from advantage calculation Value network: no Perhaps we are in the REINFORCE era of RL for LLMs still 220
Seth Karten @sethkarten.ai · 07/02/2025sometimes it takes more than machine learning to train a model 130
Seth Karten @sethkarten.ai · 26/01/2025🧠DeepSeek's R1 report highlights challenges in applying AlphaZero to LLMs. Gaming AI offers a key insight: AlphaStar, like an LLM, starts with imitation learning, then RL, and crucially, league self-play. League diversity helps avoid local optima. How do we implement this? ⬇️ 120
Seth Karten @sethkarten.ai · 12/01/2025A large part of it is detecting them early (not mine but from my prior lab). Then, I think the LA fires are particularly challenging since there must be some optimality in preventing the further spread of widespread fires 010
Seth Karten @sethkarten.ai · 09/12/2024I'll be running live demos Wed-Sat! Challenge PokéChamp and test your #Pokemon battle skills against our AI. 020
Seth Karten @sethkarten.ai · 09/12/2024We combine LLMs with minimax search, opponent modeling, and tool-use to achieve expert-level performance in competitive Pokémon battles, ranking in the top 10% of players. Check out the preprint: sethkarten.ai/data/PokeCha... #ResearchPaper #AIResearch 120
Seth Karten @sethkarten.ai · 09/12/2024I am attending #NeurIPS2024 Wed-Sun! 🎉 I will be presenting "PokéChamp: Expert-level Minimax Language Agent for Competitive Pokémon" at the Language Gamification workshop on Sat. Feel free to message me for coffee chats & party invites! #AIinGaming 190