Sign in

Chris Painter

@chris.bsky.social
3.3K followers 496 following 946 posts

President @metr.org Check out my artisanal hand-crafted "AI Bluesky" starter pack here: bsky.app/starter-pack/chris.bsky.so…

PostsRepliesMedia
Chris Painter @chris.bsky.social · 02/09/2026
METR is hiring in cyberforensics. We now embed researchers inside of AI labs to stress test monitoring, assess AI loss-of-control risk, and investigate misalignment. If you want to apply DFIR skills in frontier AI, apply (or DM). Comp range is $400k - 580k cash. jobs.lever.co/metr/b1a2f73...
jobs.lever.co
METR - Member of Technical Staff, Cyberforensics
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation an...
2325
Reposted by Chris Painter
METR @metr.org · 30/07/2026
We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions.
1371
Reposted by Chris Painter
METR @metr.org · 19/05/2026
Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test their best internal models with CoT access, (2) review non-public info about capabilities, alignment, and control. The result: our first Frontier Risk Report.
2234
Reposted by Chris Painter
METR @metr.org · 11/05/2026
We surveyed 349 technical researchers, engineers, and managers (in February–April 2026) about how they use AI tools at work. On average, participants self-report that AI use made their work 1.6–2.1x more valuable, and that this multiplier will grow over time.
273
Chris Painter @chris.bsky.social · 17/04/2026
Cool profile of @metr.org’s work in the NYT today! Particularly like this from my colleague Ajeya: “METR is an organization that asks... what we think would be most valuable for the world to know about A.I. and its risks, and then the answers are what they are.” www.nytimes.com/2026/04/17/t....
nytimes.com
How Do You Measure an A.I. Boom?
1185
Chris Painter @chris.bsky.social · 07/03/2026
More on this idea here: metr.org/blog/2024-11...
metr.org
The Rogue Replication Threat Model
Thoughts on how AI agents might develop large and resilient rogue populations.
080
Chris Painter @chris.bsky.social · 07/03/2026
Worth reading this in full. I come in skeptical, but this basically is a claim that an AI system at Alibaba attempted autonomous replication without human intervention. This excerpt was found and highlighted by Alexander Long. Full paper here: arxiv.org/abs/2512.24873
3468
Reposted by Chris Painter
METR @metr.org · 05/03/2026
We’re correcting a mistake in our modeling that inflated recent 50%-time horizons by 10-20% (and reduced 80%-horizons). We inappropriately penalized steepness in task-length→success curve fits. This most affects the oldest and newest models, whose fits are less data-constrained.
291
Reposted by Chris Painter
METR @metr.org · 24/02/2026
Since early 2025, we've been studying how AI tools impact productivity among developers. Previously, we found a 20% slowdown. That finding is now outdated. Speedups now seem likely, but changes in developer behavior make our new results unreliable. We’re working to address this.
3245
Chris Painter @chris.bsky.social · 21/02/2026
metr.org/careers
metr.org
Careers at METR
010
Chris Painter @chris.bsky.social · 21/02/2026
Our team is stretched thin at the moment! To continue upper-bounding the autonomy of AI agents, and developing evaluations for monitoring AI systems and their propensity to subvert human control, we need more great engineering and research staff. Please apply below or DM me!
161
Chris Painter @chris.bsky.social · 04/02/2026
Groundhog Day is a very METR-y holiday. Small animal emerges from a cave for only a moment, shares a forecast about timelines that's somewhat difficult to interpret, and then retreats into his cave for another year.
010
Chris Painter @chris.bsky.social · 23/01/2026
Today we published a critique of @metr.org’s time-horizon methodology by one of the paper’s lead authors, Thomas Kwa Link: metr.org/notes/2026-0...
041
Chris Painter @chris.bsky.social · 20/01/2026
Maybe this isn't what people mean by "permanent underclass", maybe they don't imagine it involves any subjugation or won't be unpleasant. That isn't the impression their urgency to evade this "underclass" gives.
020
Chris Painter @chris.bsky.social · 20/01/2026
If "escaping the permanent underclass" is the explicit motivation, then this isn't that. The explicit belief here is that some people will be subjugated, and the speaker needs to make sure they aren't one of the subjugated.
120
Chris Painter @chris.bsky.social · 20/01/2026
Obviously, it's fine and great to build one's personal wealth in ways that create consumer surplus and increased quality of life for everyone, and even to only accomplish the latter unthinkingly while e.g. building an amazing business.
130
Chris Painter @chris.bsky.social · 20/01/2026
I do occasionally now hear tech/finance people sincerely say that they need to focus on making more money to "escape the permanent underclass" It's important to emphasize how selfish orienting one's life around that goal is, rather than improving the median outcome for everyone
1100
Chris Painter @chris.bsky.social · 31/07/2025
The full website lets you toggle and see the task-horizon at 80% success rate as well. The resolution we can observe confidently is very low at pass rates like 95% Full site: metr.org/blog/2025-03... Original paper explaining: arxiv.org/abs/2503.14499
metr.org
Measuring AI Ability to Complete Long Tasks
We propose measuring AI performance in terms of the *length* of tasks AI agents can complete. We show that this metric has been consistently exponentially increasing over the past 6 years, with a doub...
011
Chris Painter @chris.bsky.social · 31/07/2025
We first characterize the difficulty of the tasks in our suite by seeing how long they take experienced human developers/engineers/researchers. We then sort the tasks into buckets based on how long they take humans. Grok 4 gets 50% success on the ~1hr50min part of the task difficulty distribution
100
Chris Painter @chris.bsky.social · 11/07/2025
Oh I also should clarify that we have many more than 2 projects going in parallel at any given time hahahaha, these two were just similar
020
Chris Painter @chris.bsky.social · 11/07/2025
Oh I also should clarify that we have many more than 2 projects going in parallel at any given time time, for what it’s worth
020
Chris Painter @chris.bsky.social · 11/07/2025
To be clear: The other project was very nascent, and would’ve been far less quantitative/experimental, more like an index of developer anecdotes. To my knowledge the RCT was not formally pre-registered, but I would want to check with the people on our team who worked on it
120
Chris Painter @chris.bsky.social · 11/07/2025
For me, the biggest upshot of this work, at the moment, is that the most obvious and straightforward ways of assessing AI R&D acceleration from access to AI, like "just survey people" or "monitor the vibes in your AI lab" probably won't work, or will badly misfire.
191
Chris Painter @chris.bsky.social · 11/07/2025
METR a few months ago had two projects going in parallel: a project experimenting with AI researcher interviews to track degree of AI R&D acceleration/delegation, and this project. When the results started coming back from this project, we put the survey-only project on ice.
2202
Reposted by Chris Painter
METR @metr.org · 13/06/2025
At METR, we’ve seen increasingly sophisticated examples of “reward hacking” on our tasks: models trying to subvert or exploit the environment or scoring code to obtain a higher score. In a new post, we discuss this phenomenon and share some especially crafty instances we’ve seen.
163
Reposted by Chris Painter
Emily Liu @emilyliu.me · 30/05/2025
personal update: today is my last day with the Bluesky team! this is bittersweet news to share, but the great thing about an open network is you never really have to leave. I’ll be rooting for Bluesky and atproto from the outside 🫡💙
842216103
Chris Painter @chris.bsky.social · 09/04/2025
In particular, the amount of influence and power that depends on the outcomes of these debates, without any of these people really being in the trenches of politics or business, feels very monastic
010
Chris Painter @chris.bsky.social · 09/04/2025
You have these monks and scholars hidden away in a sort of monastery, and the law of the land hangs on their calm debates about the correct way to interpret our secular scripture
110
Chris Painter @chris.bsky.social · 09/04/2025
I spent a few days at Yale Law, while also listening to Sam Harris’s interview with Tom Holland about his book “Dominion”, and it’s striking how similar the role and vibe of the American judiciary is to a kind of secular priesthood. Robes, scholars interpreting sacred texts
120
Chris Painter @chris.bsky.social · 26/03/2025
020
Chris Painter @chris.bsky.social · 19/03/2025
050
Reposted by Chris Painter
METR @metr.org · 19/03/2025
When will AI systems be able to carry out long projects independently? In new research, we find a kind of “Moore’s Law for AI agents”: the length of tasks that AIs can do is doubling about every 7 months.
3205
Chris Painter @chris.bsky.social · 10/03/2025
Bought a new bike this weekend :(
020
Chris Painter @chris.bsky.social · 11/02/2025
Taking science fiction seriously - thinking with effort about which ideas from sci-fi could become real soon and why and which couldn’t - has been so useful to me that it feels something like a core value
130
Chris Painter @chris.bsky.social · 29/12/2024
Look at this extremely expansive definition of Russia’s territory on my hand-drawn 7th grade map
0100
Chris Painter @chris.bsky.social · 29/12/2024
Also: my high school graduation speech was superintelligence-pilled:
1110
Chris Painter @chris.bsky.social · 29/12/2024
Cleaning a childhood bedroom and I’m struck by how much optimistic messaging about technology and space technology in particular I was surrounded by as a kid in the 90’s. Are kids still immersed in this stuff? I hope so
2201
Chris Painter @chris.bsky.social · 29/12/2024
Sadly most physical goods that you’d be tempted to donate to someone are worth less than the cost in effort it would take to find someone who needs them
220
Chris Painter @chris.bsky.social · 23/12/2024
If this group is dedicated to advocating for what it seems like they’re dedicated to advocating for, it’s pretty wild that they exist!
130
Chris Painter @chris.bsky.social · 23/12/2024
Worlds with federal pre-emption of AI policy might be correlated with worlds with a huge expansion of social attention to AI (e.g. acute labor displacement), and a less "technocratic" reaction. Will the first big federal AI bill feel more like the CARES Act or the CHIPS Act?
130
Chris Painter @chris.bsky.social · 23/12/2024
I think AI would benefit from more social contact with scientists in fields whose questions don't have intuitively verifiable answers. To assess model capability, I find myself often relying on happenstance anecdotes I hear from e.g. lab-bench researchers months after the fact.
140
Chris Painter @chris.bsky.social · 22/12/2024
I’m not sure that’s going to be a very meaningful distinction for the most advanced models, and I guess I’m specifically interested in what’s possible with both the best models-as-agents and models-as-tools
120
Chris Painter @chris.bsky.social · 22/12/2024
Will human-level AI be self-deploying/"productizing", or not? Will the "the product can explain to you how to use it and apply it" dynamic dramatically increase the adoption of AI relative to historical comparisons like AVs and steam engines?
140
Chris Painter @chris.bsky.social · 10/12/2024
Or man, idk, is it a "corollary"? Maybe it's just an example
010
Chris Painter @chris.bsky.social · 10/12/2024
A corollary to this: I think many policy initiatives would benefit from having more deeply engaged and informed opponents, and this is a neglected niche in many areas/topics. Detailed proposals having better (in the sense of more substantive) opponents is good for the world
241
Chris Painter @chris.bsky.social · 10/12/2024
I think the world could always benefit from more good-faith really in-depth critique of effortful technical/intellectual work. Many organizations that I collaborate with publish work hoping to have their ideas improved upon or attacked, but often surprisingly few people engage.
030
Chris Painter @chris.bsky.social · 10/12/2024
I’ve been wondering, has the (bad) reaction to this actually been that unusual?
010
Chris Painter @chris.bsky.social · 10/12/2024
Maybe now deleted?
000
Chris Painter @chris.bsky.social · 09/12/2024
Luigi Mangione's review of the Unabomber manifesto on Goodreads
151
Chris Painter @chris.bsky.social · 08/12/2024
Uber Eats is truly an embarrassment of riches. We are living in a golden age of delivered food.
010