Sign in

Seth Karten

@sethkarten.ai
565 followers 1.5K following 231 posts

Autonomous Agents | Research @ Prime Intellect | PhD @ Princeton | Prev: CMU, Waymo | NSF GRFP Fellow sethkarten.ai

PostsRepliesMedia
Seth Karten @sethkarten.ai · 29/09/2026
Feel free to follow on /nightly
110
Seth Karten @sethkarten.ai · 29/09/2026
We have merged a fix for this. Will be doing a minor release to fix this. Then no release for a week while we dogfood some huge performance changes :)
120
Seth Karten @sethkarten.ai · 29/09/2026
Is this just claude code? Or codex sub too?
100
Seth Karten @sethkarten.ai · 25/09/2026
This is exactly the kind of setting we built Agent Bazaar to study, and I think these questions are only going to become more important as agents start transacting and negotiating on behalf of users. I'll be presenting Agent Bazaar at COLM Wed Oct 7. arxiv.org/abs/2605.17698
arxiv.org
Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
The deployment of Large Language Models (LLMs) as autonomous economic agents introduces systemic risks that extend beyond individual capability failures. As agents transition to directly interacting w...
070
Seth Karten @sethkarten.ai · 25/09/2026
The risks go beyond whether any single agent is aligned with its user. Agents can collectively crash markets even when each individual agent is behaving rationally, and users can get scammed by autonomous sellers. Meta Muse is starting to make this future real pretty quickly.
240
Seth Karten @sethkarten.ai · 25/09/2026
We wrote Agent Bazaar back in May around a future where agents act on behalf of users and increasingly participate directly in marketplaces like Amazon and eBay. We introduced Economic Alignment to study what happens as agents become economic actors.
2143
Seth Karten @sethkarten.ai · 24/09/2026
prime agent v0.9.6 is out: ◆ Support for GPT-6 Sol, Opus 5.5, and Grok 4.7 ◆ /mcp plugin catalog with one-click connections to Linear, Notion, Posthog, Stripe, and 60+ more services ◆ Huge perf and reliability pass 🫡 Lots more coming soon :)
1263
Seth Karten @sethkarten.ai · 23/09/2026
This is fixed in /nightly (main of github repo). Release not coming until tomorrow
110
Seth Karten @sethkarten.ai · 22/09/2026
Also we are making new models available without requiring prime agent update going forward so you wont need to update to access new models going forward
140
Seth Karten @sethkarten.ai · 22/09/2026
We have open source you can fork. I think we have a PR to add new opus. Let me check on it
210
Seth Karten @sethkarten.ai · 17/09/2026
i think this is it. should be merged soon and on nightly github.com/PrimeIntelle...
github.com
send queued messages after interrupt by kevinjosethomas · Pull Request #2426 · PrimeIntellect-ai/prime-agent
pressing escape or ctrl+c now aborts the active run and immediately sends queued messages together. preserves message order and the configured steering mode while keeping empty queues abort-only. a...
120
Seth Karten @sethkarten.ai · 17/09/2026
I think we just merged this today in the nightly. Will be in next release
130
Seth Karten @sethkarten.ai · 17/09/2026
Can you use the daemon cli commands for that?
100
Seth Karten @sethkarten.ai · 16/09/2026
prime agent v0.9.5 we fixed a lot of bugs and, of course, we had prime agent feature its favorite updates it picked our perf work. then it created the video itself.
3311
Seth Karten @sethkarten.ai · 15/09/2026
Quick update: we’re adding a few incentives for contributors. Everyone who submits data will be acknowledged in the dataset release, and the full dataset will be open sourced. If you get especially involved in collection and data processing/curation, there may also be author opportunities
032
Seth Karten @sethkarten.ai · 12/09/2026
Noted. Team is working on a front-end overhaul
131
Seth Karten @sethkarten.ai · 10/09/2026
We also released the arxiv on Prime Agent: A Self-Improving RLM Harness (with Swarm Communication) arxiv.org/abs/2608.23552
arxiv.org
Prime Agent: A Self-Improving RLM Harness
Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long...
071
Seth Karten @sethkarten.ai · 10/09/2026
Wow excited to see a prime agent community emerged here. Sharing my latest YC Paper Club invited talk: youtu.be/n9xKblqyQ28?...
youtu.be
Why The Harness Matters More Than The Model | YC Paper Club
YouTube video by Y Combinator
1252
Reposted by Seth Karten
prolepses.bsky.social @prolepses.bsky.social · 09/09/2026
Reposting with a direct link right to where @sethkarten.ai breaks down Prime Agent. The TLDR is that Prime Agent is "Jupyter notebooks for your agents", the video is absolutely worth the watch: youtu.be/n9xKblqyQ28?...
youtu.be
Why The Harness Matters More Than The Model | YC Paper Club
YouTube video by Y Combinator
1113
Reposted by Seth Karten
Chris Patil @donotgogently.bsky.social · 09/09/2026
If you’re interested in Prime Agent by @primeintellect.bsky.social , Seth Karten (researcher at Prime, author of the Prime Agent paper) gave a nice talk on the harness at YCombinator the other day youtu.be/n9xKblqyQ28?...
youtu.be
Why The Harness Matters More Than The Model | YC Paper Club
YouTube video by Y Combinator
2164
Seth Karten @sethkarten.ai · 08/09/2026
My group at Princeton is collecting crowdsourced, action-labeled game data for research on AI agents, starting with Pokémon Emerald. The goal is to build a large open-source dataset of action-labeled gaming trajectories. Sign up to contribute here: docs.google.com/forms/d/e/1F...
docs.google.com
Play Pokémon Emerald, Help Train a World Model.
We're building generalizable world model agents, AIs that learn to understand and simulate games from video. We need your help to gather gameplay recordings to generate a dataset which can then teach ...
061
Seth Karten @sethkarten.ai · 31/08/2026
This slide has done more harm than good
010
Seth Karten @sethkarten.ai · 19/08/2026
If you cared about safety, you would have allowed waymos in nyc
120
Reposted by Seth Karten
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 15/08/2026
There’s a bit of a sense of despair in the research community around LLMs. It can be avoided by switching to caring about what we should build as opposed to how we build it
5483
Seth Karten @sethkarten.ai · 13/08/2026
Thanks! We have a lot of cool stuff still coming
000
Seth Karten @sethkarten.ai · 12/08/2026
Made an account using my bluesky but im not sure how to use it. Is it mainly for blogs? Can i post like I would post my research on X or bluesky?
120
Seth Karten @sethkarten.ai · 10/08/2026
www.primeintellect.ai/blog/prime-a...
primeintellect.ai
Prime Agent: A self-improving RLM agent
Prime Agent is our open-source, self-improving coding harness built around two abstractions: the Recursive Language Model (RLM) and the Continual Harness. With Opus 5, it achieves 95.5% on ARC-AGI-3, ...
010
Seth Karten @sethkarten.ai · 10/08/2026
Really excited about this one :) Prime Agent + Opus 5 gets 95.5% on ARC-AGI-3 (179/183). A big goal when I was building Prime Agent was making long-horizon agents more capable yet token efficient. github.com/PrimeIntelle...
github.com
GitHub - PrimeIntellect-ai/prime-agent: A self-improving RLM agent for coding workflows and long-running autonomous tasks.
A self-improving RLM agent for coding workflows and long-running autonomous tasks. - PrimeIntellect-ai/prime-agent
2333
Seth Karten @sethkarten.ai · 10/08/2026
www.primeintellect.ai/blog/prime-a...
primeintellect.ai
Prime Agent: A self-improving RLM agent
Prime Agent is our open-source, self-improving coding harness built around two abstractions: the Recursive Language Model (RLM) and the Continual Harness. With Opus 5, it achieves 95.5% on ARC-AGI-3, ...
011
Seth Karten @sethkarten.ai · 10/08/2026
Really excited about this one :) Prime Agent + Opus 5 gets 95.5% on ARC-AGI-3 (179/183). A big goal when I was building Prime Agent was making long-horizon agents more capable yet token efficient. github.com/PrimeIntelle...
github.com
GitHub - PrimeIntellect-ai/prime-agent: A self-improving RLM agent for coding workflows and long-running autonomous tasks.
A self-improving RLM agent for coding workflows and long-running autonomous tasks. - PrimeIntellect-ai/prime-agent
140
Seth Karten @sethkarten.ai · 23/07/2026
I agree with the “we dont really know until we see it in practice” Though largely i see public science funding trending from “those not willing to use AI” to “those willing to use AI” Or at least write so in their proposal
020
Seth Karten @sethkarten.ai · 22/07/2026
How are the academics feeling about this? Does it even change anything for profs?
120
Seth Karten @sethkarten.ai · 22/07/2026
Great to see Continual Harness acknowledged in Schmidhuber’s latest survey paper
050
Seth Karten @sethkarten.ai · 14/07/2026
Going through my backlog and realizing I forgot to announce 1 paper and never put another on arXiv. Expect 2 blog posts soon
1100
Seth Karten @sethkarten.ai · 12/07/2026
Odysseus (PPO for VLMs/LLMs) We found the way to apply PPO to VLMs/LLMs for long horizon tasks using super mario world as our case study huggingface.co/papers/2605....
huggingface.co
Paper page - Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
Join the discussion on this paper page
071
Seth Karten @sethkarten.ai · 12/07/2026
Automatic RL Env The main takeaway is to develop a strong verifier feedback loop. Here we create the idea of minimizing a sim-to-sim gap between environments in order to correctly verify performance versions of envs huggingface.co/papers/2603....
huggingface.co
Paper page - Automatic Generation of High-Performance RL Environments
Join the discussion on this paper page
130
Seth Karten @sethkarten.ai · 12/07/2026
Agent Bazaar (multi-agent safety & economic envs) This is my personal take on looking at the macro behavior of agents in economic settings wrt managing supply, consumer demand, and trust in long horizons huggingface.co/papers/2605....
huggingface.co
Paper page - Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
Join the discussion on this paper page
140
Seth Karten @sethkarten.ai · 12/07/2026
Wow, three papers in CoLM 2026... Here I come San Francisco! These papers predicted some early trends in multi-agent safety & economic envs, automatic RL env creation, and PPO for VLMs/LLMs 🧵
130
Seth Karten @sethkarten.ai · 12/07/2026
New blog applying Continual Harness to ARC-AGI-3. The heavy test-time learning required by the benchmark pushes agents to form an internal world model of the rules and mechanics that updates with new evidence. Continual Harness scored 20.54%. sethkarten.substack.com/p/continual-...
open.substack.com
Continual Harness: An Efficient Self-Improving Agent on ARC-AGI-3
Continual Harness scores 20.54% on ARC-AGI-3 at $774, showing how reset-free self-improving agents can learn hidden game dynamics at test time.
0121
Seth Karten @sethkarten.ai · 10/07/2026
Same. I need a better feed here
0141
Seth Karten @sethkarten.ai · 10/07/2026
Just seize money from football. They should pay the university rent
100
Seth Karten @sethkarten.ai · 11/06/2026
😵‍💫 what if you compare the elon holding companies to alphabet. Maybe it’s roughly the same give or take a few 100 billion
000
Seth Karten @sethkarten.ai · 10/06/2026
They own a lot of gpus and electricity. Starlink is profitable. As shown with the anthropic and google deals, they dont need to make their own AI to make money when gpus and power are in short supply
110
Seth Karten @sethkarten.ai · 22/05/2026
Just know that my reviewers will be thoroughly reviewed for their strengths and weaknesses. Score and confidence included.
010
Seth Karten @sethkarten.ai · 19/05/2026
Midwest emo is from new jersey 😋
110
Seth Karten @sethkarten.ai · 19/05/2026
Econ lovers i have something for you
010
Seth Karten @sethkarten.ai · 14/05/2026
New paper alert: Continual Harness: Online Adaptation for Self-Improving Foundation Agents Paper (arXiv). arxiv.org/abs/2605.09998 Article (Substack). sethkarten.substack.com/p/gemini-pla... Project page (video demos). sethkarten.ai/continual-ha...
0203
Seth Karten @sethkarten.ai · 13/05/2026
Announcing some work tomorrow. Will be cool and probably involving pokemon
010
Seth Karten @sethkarten.ai · 26/04/2026
im very glad to see a rebound this year in the total number of NSF GRFP awards to exceed the most this century. Founding the next generation of American scientists is important to keep growth of the sciences.
030
Seth Karten @sethkarten.ai · 01/04/2026
Huge money in providing an API to directly get bibtex formatted citations and would fix this issue easily Gscholar doesnt allow this Semantic scholar rate limits Who is building this?
020