Sign in

Seth Karten

@sethkarten.ai
565 followers 1.5K following 231 posts

Autonomous Agents | Research @ Prime Intellect | PhD @ Princeton | Prev: CMU, Waymo | NSF GRFP Fellow sethkarten.ai

PostsRepliesMedia
Seth Karten @sethkarten.ai · 25/09/2026
We wrote Agent Bazaar back in May around a future where agents act on behalf of users and increasingly participate directly in marketplaces like Amazon and eBay. We introduced Economic Alignment to study what happens as agents become economic actors.
2143
Seth Karten @sethkarten.ai · 24/09/2026
prime agent v0.9.6 is out: ◆ Support for GPT-6 Sol, Opus 5.5, and Grok 4.7 ◆ /mcp plugin catalog with one-click connections to Linear, Notion, Posthog, Stripe, and 60+ more services ◆ Huge perf and reliability pass 🫡 Lots more coming soon :)
1263
Seth Karten @sethkarten.ai · 16/09/2026
prime agent v0.9.5 we fixed a lot of bugs and, of course, we had prime agent feature its favorite updates it picked our perf work. then it created the video itself.
3311
Seth Karten @sethkarten.ai · 15/09/2026
Quick update: we’re adding a few incentives for contributors. Everyone who submits data will be acknowledged in the dataset release, and the full dataset will be open sourced. If you get especially involved in collection and data processing/curation, there may also be author opportunities
032
Seth Karten @sethkarten.ai · 10/09/2026
Wow excited to see a prime agent community emerged here. Sharing my latest YC Paper Club invited talk: youtu.be/n9xKblqyQ28?...
youtu.be
Why The Harness Matters More Than The Model | YC Paper Club
YouTube video by Y Combinator
1252
Reposted by Seth Karten
prolepses.bsky.social @prolepses.bsky.social · 09/09/2026
Reposting with a direct link right to where @sethkarten.ai breaks down Prime Agent. The TLDR is that Prime Agent is "Jupyter notebooks for your agents", the video is absolutely worth the watch: youtu.be/n9xKblqyQ28?...
youtu.be
Why The Harness Matters More Than The Model | YC Paper Club
YouTube video by Y Combinator
1113
Reposted by Seth Karten
Chris Patil @donotgogently.bsky.social · 09/09/2026
If you’re interested in Prime Agent by @primeintellect.bsky.social , Seth Karten (researcher at Prime, author of the Prime Agent paper) gave a nice talk on the harness at YCombinator the other day youtu.be/n9xKblqyQ28?...
youtu.be
Why The Harness Matters More Than The Model | YC Paper Club
YouTube video by Y Combinator
2164
Seth Karten @sethkarten.ai · 08/09/2026
My group at Princeton is collecting crowdsourced, action-labeled game data for research on AI agents, starting with Pokémon Emerald. The goal is to build a large open-source dataset of action-labeled gaming trajectories. Sign up to contribute here: docs.google.com/forms/d/e/1F...
docs.google.com
Play Pokémon Emerald, Help Train a World Model.
We're building generalizable world model agents, AIs that learn to understand and simulate games from video. We need your help to gather gameplay recordings to generate a dataset which can then teach ...
061
Reposted by Seth Karten
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 15/08/2026
There’s a bit of a sense of despair in the research community around LLMs. It can be avoided by switching to caring about what we should build as opposed to how we build it
5483
Seth Karten @sethkarten.ai · 10/08/2026
Really excited about this one :) Prime Agent + Opus 5 gets 95.5% on ARC-AGI-3 (179/183). A big goal when I was building Prime Agent was making long-horizon agents more capable yet token efficient. github.com/PrimeIntelle...
github.com
GitHub - PrimeIntellect-ai/prime-agent: A self-improving RLM agent for coding workflows and long-running autonomous tasks.
A self-improving RLM agent for coding workflows and long-running autonomous tasks. - PrimeIntellect-ai/prime-agent
2333
Seth Karten @sethkarten.ai · 10/08/2026
Really excited about this one :) Prime Agent + Opus 5 gets 95.5% on ARC-AGI-3 (179/183). A big goal when I was building Prime Agent was making long-horizon agents more capable yet token efficient. github.com/PrimeIntelle...
github.com
GitHub - PrimeIntellect-ai/prime-agent: A self-improving RLM agent for coding workflows and long-running autonomous tasks.
A self-improving RLM agent for coding workflows and long-running autonomous tasks. - PrimeIntellect-ai/prime-agent
140
Seth Karten @sethkarten.ai · 22/07/2026
How are the academics feeling about this? Does it even change anything for profs?
120
Seth Karten @sethkarten.ai · 22/07/2026
Great to see Continual Harness acknowledged in Schmidhuber’s latest survey paper
050
Seth Karten @sethkarten.ai · 14/07/2026
Going through my backlog and realizing I forgot to announce 1 paper and never put another on arXiv. Expect 2 blog posts soon
1100
Seth Karten @sethkarten.ai · 12/07/2026
Wow, three papers in CoLM 2026... Here I come San Francisco! These papers predicted some early trends in multi-agent safety & economic envs, automatic RL env creation, and PPO for VLMs/LLMs 🧵
130
Seth Karten @sethkarten.ai · 12/07/2026
New blog applying Continual Harness to ARC-AGI-3. The heavy test-time learning required by the benchmark pushes agents to form an internal world model of the rules and mechanics that updates with new evidence. Continual Harness scored 20.54%. sethkarten.substack.com/p/continual-...
open.substack.com
Continual Harness: An Efficient Self-Improving Agent on ARC-AGI-3
Continual Harness scores 20.54% on ARC-AGI-3 at $774, showing how reset-free self-improving agents can learn hidden game dynamics at test time.
0121
Seth Karten @sethkarten.ai · 10/07/2026
Same. I need a better feed here
0141
Seth Karten @sethkarten.ai · 22/05/2026
Just know that my reviewers will be thoroughly reviewed for their strengths and weaknesses. Score and confidence included.
010
Seth Karten @sethkarten.ai · 19/05/2026
Econ lovers i have something for you
010
Seth Karten @sethkarten.ai · 14/05/2026
New paper alert: Continual Harness: Online Adaptation for Self-Improving Foundation Agents Paper (arXiv). arxiv.org/abs/2605.09998 Article (Substack). sethkarten.substack.com/p/gemini-pla... Project page (video demos). sethkarten.ai/continual-ha...
0203
Seth Karten @sethkarten.ai · 13/05/2026
Announcing some work tomorrow. Will be cool and probably involving pokemon
010
Seth Karten @sethkarten.ai · 26/04/2026
im very glad to see a rebound this year in the total number of NSF GRFP awards to exceed the most this century. Founding the next generation of American scientists is important to keep growth of the sciences.
030
Seth Karten @sethkarten.ai · 30/03/2026
New meta seems to be arxiving a rough draft so that you can claim the terminology first and claim to be first
110
Seth Karten @sethkarten.ai · 29/03/2026
Everyone wants to own their own data but no one wants to own their own data center
010
Seth Karten @sethkarten.ai · 24/03/2026
1) what
210
Seth Karten @sethkarten.ai · 18/03/2026
open.substack.com/pub/sethkart...
open.substack.com
We Ran the Largest AI Pokemon Tournament Ever. Now It's an Open Benchmark.
In 2025, everyone was talking about LLMs playing Pokemon.
131
Seth Karten @sethkarten.ai · 17/03/2026
i owe bluesky a post soon. everyone please hold on
120
Reposted by Seth Karten
Elizabeth Mieczkowski @emiecz.bsky.social · 16/03/2026
🚨New preprint! LLM teams are being deployed at scale, yet we lack the tools to predict when they’ll succeed, fail, or how to design them. Distributed computing faced the exact same questions and figured out how to answer them. We show those insights apply directly to LLMs 🧵👇
1323
Seth Karten @sethkarten.ai · 13/03/2026
open.substack.com/pub/sethkart...
open.substack.com
We Automated RL Environment Engineering for $10
RL environment simulation eats 50-90% of training wall-clock for specialist RL policies. Coding agents can translate them automatically with no sim-to-sim gap.
264
Seth Karten @sethkarten.ai · 25/12/2025
I think I accidentally stumbled upon engagement baiting from first principles Ill stay on bluesky as long as the 10 accounts I like to see still post here
010
Seth Karten @sethkarten.ai · 24/12/2025
I think I might leave bluesky tbh
110
Reposted by Seth Karten
Seth Karten @sethkarten.ai · 24/11/2025
How do we close the gap between specialist RL and generalist LLM agents? We're benchmarking it in Pokémon. Join us at the PokeAgent Challenge competition workshop @ NeurIPS 2025. 📍 Dec 7, 8AM 🎮 Track 1: Competitive Pokémon (game-theoretic reasoning) 🗺️ Track 2: Speedrunning (long-horizon planning)
Flyer for The PokeAgent Challenge at NeurIPS 2025. Sunday, Dec 7, 8–10:45 AM PST, Mezzanine Room 15AB, San Diego Convention Center. Two tracks: Track 1 (Battling) features competitive Pokémon battle bots; Track 2 (Speedrunning) features long-horizon RPG gameplay. Tagline: "How do we close the gap between specialist RL models and generalist LLM agents?" Speakers: Seth Karten (Princeton), Aaron Traylor, Minmin Chen (Google DeepMind), Jake Grigsby (UT Austin), Stephanie Milani (NYU/Johns Hopkins), Kiran Vodrahalli (Google DeepMind), Fei Fang (CMU), Yuke Zhu (UT Austin), Chi Jin (Princeton). Sponsored by Google DeepMind.
043
Seth Karten @sethkarten.ai · 26/11/2025
I’ll be in San Diego at NeurIPS Dec 3-7! DM or email if you want to chat about - building the foundation agents through games - PokeAgent Challenge & PokéChamp - LLM Economist & autonomous business agents
031
Seth Karten @sethkarten.ai · 24/11/2025
How do we close the gap between specialist RL and generalist LLM agents? We're benchmarking it in Pokémon. Join us at the PokeAgent Challenge competition workshop @ NeurIPS 2025. 📍 Dec 7, 8AM 🎮 Track 1: Competitive Pokémon (game-theoretic reasoning) 🗺️ Track 2: Speedrunning (long-horizon planning)
Flyer for The PokeAgent Challenge at NeurIPS 2025. Sunday, Dec 7, 8–10:45 AM PST, Mezzanine Room 15AB, San Diego Convention Center. Two tracks: Track 1 (Battling) features competitive Pokémon battle bots; Track 2 (Speedrunning) features long-horizon RPG gameplay. Tagline: "How do we close the gap between specialist RL models and generalist LLM agents?" Speakers: Seth Karten (Princeton), Aaron Traylor, Minmin Chen (Google DeepMind), Jake Grigsby (UT Austin), Stephanie Milani (NYU/Johns Hopkins), Kiran Vodrahalli (Google DeepMind), Fei Fang (CMU), Yuke Zhu (UT Austin), Chi Jin (Princeton). Sponsored by Google DeepMind.
043
Seth Karten @sethkarten.ai · 20/10/2025
Every LLM eval uses Bradley-Terry Elo rankings. Almost none report uncertainty. Should we trust them? Maybe there is something better... 👇 (1/5)
110
Seth Karten @sethkarten.ai · 15/10/2025
Pokemon is truly the pareto frontier of agent research - The RPG requires an autonomous embodied agentic agent with perception, planning, memory, and control - VGC and Gen 9 OU penalize erroneous actions with fast-paced opponent-modeling in short games (1/3)
191
Seth Karten @sethkarten.ai · 15/10/2025
Trying to get a post ready but bluesky won’t let me post on desktop!!! If you want users here you need a user experience!!!
100
Seth Karten @sethkarten.ai · 08/10/2025
You probably aren’t reading enough papers. You probably didn’t cite the 10 closest papers to your work Thus, LLMs probably have a better understanding of where your paper sits in the literature ¯\_(ツ)_/¯
010
Seth Karten @sethkarten.ai · 24/09/2025
The most interesting papers arent being published at the “prestigious” venues anymore. Where are you publishing and what do you work on?
110
Seth Karten @sethkarten.ai · 02/09/2025
🚨 Hackathon Weekend! 🚨 Jumpstart your PokéAgent Challenge submission ahead of NeurIPS! 📅 Sept 13–14 ✅ Leaderboards reset Sat 10AM EDT 🎙️ Lightning talks in LLMs, RL, and Pokemon 💬 Live Office hours 🏆 $2k in prizes
PokéAgent Challenge @ NeurIPS 2025 Hackathon Weekend Schedule. Saturday, Sept 13th: 10 AM leaderboards reset; 12–1:30 PM livestream talks (overview, Aaron Traylor on Pokémon as an AI Problem, Seth Karten on Pokéchamp, Jake Grigsby on Metamon, plus more). Sunday, Sept 14th: 1–3:30 PM organizer office hours; 11:59 PM top teams earn up to $2k in GCP credits. Sponsored by Google DeepMind and AIJ.
172
Reposted by Seth Karten
Seth Karten @sethkarten.ai · 15/08/2025
The NeurIPS 2025 PokéAgent Challenge is offering compute credits, courtesy of our sponsor Google DeepMind, to help you train bigger models & run more experiments. 📌 To apply: 1️⃣ Make a submission to Track 1 or 2 at pokeagent.github.io 2️⃣ Fill out the compute credit form on the site
pokeagent.github.io
PokéAgent Challenge - NeurIPS 2025
094
Seth Karten @sethkarten.ai · 18/08/2025
Mad about data centers? Call your reps to build more nuclear
000
Seth Karten @sethkarten.ai · 15/08/2025
Hey #academics Why are neurips workshop deadlines due a month before main track acceptances? Seems counterintuitive to have the two tracks compete with each other #machinelearning
010
Seth Karten @sethkarten.ai · 15/08/2025
The NeurIPS 2025 PokéAgent Challenge is offering compute credits, courtesy of our sponsor Google DeepMind, to help you train bigger models & run more experiments. 📌 To apply: 1️⃣ Make a submission to Track 1 or 2 at pokeagent.github.io 2️⃣ Fill out the compute credit form on the site
pokeagent.github.io
PokéAgent Challenge - NeurIPS 2025
094
Seth Karten @sethkarten.ai · 12/08/2025
If your final product doesnt reason in-context, how is it supposed to meta-learn and address distribution shifts and environment changes?
110
Seth Karten @sethkarten.ai · 12/08/2025
Papers are dead. Maybe it is time to start the youtube channel…
120
Seth Karten @sethkarten.ai · 12/08/2025
Viral paper out today about predicting brain stimulus from video inputs. as always dont overfit on first order responses. If you oversaturate stimulus, people will stop using the product(people uninstalling IG because it is too addicting) The attention economy must be modeled as a multi-agent system
020
Reposted by Seth Karten
Seth Karten @sethkarten.ai · 23/07/2025
🚀 New preprint! 🤔 Can one agent “nudge” a synthetic civilization of Census‑grounded agents toward higher social welfare—all by optimizing utilities in‑context? Meet the LLM Economist ↓
Diagram of LLM Economist: left—grid of persona‑conditioned worker agents; center—planner LLM sends tax schedule; right—social‑welfare ‘hill‑climb’.
194
Seth Karten @sethkarten.ai · 23/07/2025
🚀 New preprint! 🤔 Can one agent “nudge” a synthetic civilization of Census‑grounded agents toward higher social welfare—all by optimizing utilities in‑context? Meet the LLM Economist ↓
Diagram of LLM Economist: left—grid of persona‑conditioned worker agents; center—planner LLM sends tax schedule; right—social‑welfare ‘hill‑climb’.
194
Seth Karten @sethkarten.ai · 14/07/2025
🚀 Launch day! The NeurIPS 2025 PokéAgent Challenge is live. @neuripsconf.bsky.social Two tracks: ① Showdown Battling – imperfect-info, turn-based strategy ② Pokemon Emerald Speedrunning – long horizon RPG planning 5 M labeled replays • starter kit • baselines. Bring your LLM, RL, or hybrid agent!
Banner reading “PokéAgent Challenge @ NeurIPS 2025” with two panels: Track 1 – Competitive Pokémon Battle Bots, Track 2 – Long-Horizon RPG Gameplay. Call-to-action: “Create video-game AI! Win prizes! Live now at pokeagent.github.io.”
185