METR @metr.org · 26/08/2026METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. 646598
METR @metr.org · 09/09/2026We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement. 419215
Reposted by METRChris Painter @chris.bsky.social · 02/09/2026METR is hiring in cyberforensics. We now embed researchers inside of AI labs to stress test monitoring, assess AI loss-of-control risk, and investigate misalignment. If you want to apply DFIR skills in frontier AI, apply (or DM). Comp range is $400k - 580k cash. jobs.lever.co/metr/b1a2f73...jobs.lever.coMETR - Member of Technical Staff, CyberforensicsAbout METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation an... 2325
Reposted by METRderifatives.bsky.social @derifatives.bsky.social · 31/08/2026After 19 amazing years at Google, I'm delighted to announce I'll be joining METR as a Member of Technical Staff. I'm excited about METR, their work to date, and their mission. Society needs independent expert orgs that can deeply understand and communicate about AI's capabilities and risks. 1702
METR @metr.org · 15/08/2026In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI incidents, and more. 1182
METR @metr.org · 30/07/2026We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions. 1371
METR @metr.org · 29/07/2026We believe it's important to track and investigate misalignment incidents: cases where an AI agent autonomously took sophisticated, sustained actions in violation of human intent. In a new post, we lay out how independent propensity investigations of such incidents could be conducted. 1120
METR @metr.org · 21/07/2026Introducing “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems. The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon. 172
METR @metr.org · 27/06/2026OpenAI gave METR early access to GPT-5.6 Sol for testing including raw chain-of-thought, a railfree version of the model, and internal information about the model. With this access, METR conducted a pre-deployment evaluation of GPT-5.6 Sol, including an attempted measurement of its 50%-Time Horizon. 1190
METR @metr.org · 19/05/2026Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test their best internal models with CoT access, (2) review non-public info about capabilities, alignment, and control. The result: our first Frontier Risk Report. 2234
METR @metr.org · 11/05/2026We surveyed 349 technical researchers, engineers, and managers (in February–April 2026) about how they use AI tools at work. On average, participants self-report that AI use made their work 1.6–2.1x more valuable, and that this multiplier will grow over time. 273
METR @metr.org · 09/05/2026We reviewed a section of Anthropic’s February 2026 Risk Report focused on automated R&D risk from Opus 4.6. While we take issue with the adequacy of evidence the report provides, we agree with Anthropic about the overall level of risk & remain excited to pilot reviews like these. 170
METR @metr.org · 09/05/2026We evaluated an early version of Claude Mythos Preview for risk assessment during a limited window in March 2026. We estimated a 50%-time-horizon of at least 16hrs (95% CI 8.5hrs to 55hrs) on our task suite, at the upper end of what we can measure without new tasks. 48516
Reposted by METRChris Painter @chris.bsky.social · 17/04/2026Cool profile of @metr.org’s work in the NYT today! Particularly like this from my colleague Ajeya: “METR is an organization that asks... what we think would be most valuable for the world to know about A.I. and its risks, and then the answers are what they are.” www.nytimes.com/2026/04/17/t....nytimes.comHow Do You Measure an A.I. Boom? 1185
METR @metr.org · 10/04/2026We co-developed MirrorCode with @epochai.bsky.social to test AI on extremely long-horizon blackbox software reimplementation tasks. We found that recent public models are able to fully implement at least some programs we estimate would take humans weeks or months to implement. 140
METR @metr.org · 10/04/2026We ran GPT-5.4 (xhigh) on our tasks. Its time-horizon depends greatly on our treatment of reward hacks: the point estimate would be 5.7hrs (95% CI of 3hrs to 13.5hrs) under our standard methodology, but 13hrs (95% CI of 5hrs to 74hrs) if we allow reward hacks. 1120
METR @metr.org · 05/03/2026We’re correcting a mistake in our modeling that inflated recent 50%-time horizons by 10-20% (and reduced 80%-horizons). We inappropriately penalized steepness in task-length→success curve fits. This most affects the oldest and newest models, whose fits are less data-constrained. 291
METR @metr.org · 24/02/2026Since early 2025, we've been studying how AI tools impact productivity among developers. Previously, we found a 20% slowdown. That finding is now outdated. Speedups now seem likely, but changes in developer behavior make our new results unreliable. We’re working to address this. 3245
METR @metr.org · 23/02/2026We estimate that GPT-5.3-Codex with reasoning effort `high` (not `xhigh`) has a 50%-time-horizon of around 6.5 hours (95% CI of 3 hrs to 17 hrs) on our suite of software tasks. OpenAI provided API access for this evaluation. 240
METR @metr.org · 21/02/2026We estimate that Claude Opus 4.6 has a 50%-time-horizon of around 14.5 hours (95% CI of 6 hrs to 98 hrs) on software tasks. While this is the highest point estimate we’ve reported, this measurement is extremely noisy because our current task suite is nearly saturated. 1241
METR @metr.org · 04/02/2026We estimate that GPT-5.2 with `high` (not `xhigh`) reasoning effort has a 50%-time-horizon of around 6.6 hrs (95% CI of 3 hr 20 min to 17 hr 30 min) on our expanded suite of software tasks. This is the highest estimate for a time horizon measurement we have reported to date. 3182
METR @metr.org · 03/02/2026We’ve started to measure time horizons for recent models using our updated methodology. On this expanded suite of software tasks, we estimate that Gemini 3 Pro has a 50%-time-horizon of around 4 hrs (95% CI of 2 hr 10 mins to 7 hrs 20 mins). 170
METR @metr.org · 03/02/2026We’re updating the way we measure model time horizons on software tasks (TH 1.0→1.1). The updated methodology incorporates more of the tasks from HCAST, expanding our total from 170 to 228. This produces tighter estimates, especially at longer horizons. 150
Reposted by METRChris Painter @chris.bsky.social · 23/01/2026Today we published a critique of @metr.org’s time-horizon methodology by one of the paper’s lead authors, Thomas Kwa Link: metr.org/notes/2026-0... 041
METR @metr.org · 23/01/2026How well can AI-based monitoring detect when an agent is covertly pursuing a side objective? In early work, we find clear trends: more capable models (in terms of time horizon) are better able to detect covert behavior. 1120
METR @metr.org · 02/09/2025We estimate that Claude Opus 4.1 has a 50%-time-horizon of around 1 hr 45 min (95% confidence interval of 50 to 195 minutes) on our agentic multi-step software engineering tasks. This estimate is lower than the current highest time-horizon point estimate of around 2 hr 15 min. 150
METR @metr.org · 13/08/2025We tested how autonomous AI agents perform on real software tasks from our recent developer productivity RCT. We found a gap between algorithmic scoring and real-world usability that may help explain why AI benchmarks feel disconnected from reality. 1257
METR @metr.org · 12/08/2025Before publishing our recent developer productivity RCT, we thought hard about how to accurately and clearly communicate our results. In a new blog post, we outline some of our key considerations regarding scientific integrity and communication. 191
METR @metr.org · 11/08/2025Prior work has found that Chain of Thought (CoT) can be unfaithful. Should we then ignore what it says? In new research, we find that the CoT is informative about LLM cognition as long as the cognition is complex enough that it can’t be performed in a single forward pass. 130
METR @metr.org · 08/08/2025In a new report, we evaluate whether GPT-5 poses significant catastrophic risks via AI R&D acceleration, rogue replication, or sabotage of AI labs. We conclude that this seems unlikely. However, capability trends continue rapidly, and models display increasing eval awareness. 1202
METR @metr.org · 31/07/2025We found that Grok 4’s 50%-time-horizon on our agentic multi-step software engineering tasks is about 1hr 50min (with a 95% CI of 48min to 3hr 52min) compared to o3 (previous SOTA) at about 1hr 30min. However, Grok 4’s time horizon is below SOTA at higher success rate thresholds. 140
METR @metr.org · 30/07/2025We have open-sourced anonymized data and core analysis code for our developer productivity RCT. The paper is also live on arXiv, with two new sections: One discussing alternative uncertainty estimation methods, and a new 'bias from developer recruitment' factor that has unclear effect on slowdown. 1288
METR @metr.org · 14/07/2025METR previously estimated that the time horizon of AI agents on software tasks is doubling every 7 months. We have now analyzed 9 other benchmarks for scientific reasoning, math, robotics, computer use, and self-driving; we observe generally similar rates of improvement. 164
Reposted by METRChris Painter @chris.bsky.social · 11/07/2025METR a few months ago had two projects going in parallel: a project experimenting with AI researcher interviews to track degree of AI R&D acceleration/delegation, and this project. When the results started coming back from this project, we put the survey-only project on ice. 2202
METR @metr.org · 10/07/2025We ran a randomized controlled trial to see how much AI coding tools speed up experienced open-source developers. The results surprised us: Developers thought they were 20% faster with AI tools, but they were actually 19% slower when they had access to AI than when they didn't. 11668692993
METR @metr.org · 01/07/2025In measurements using our set of multi-step software and reasoning tasks, Claude 4 Opus and Sonnet reach 50%-time-horizon point estimates of about 80 and 65 minutes, respectively. 110
METR @metr.org · 13/06/2025At METR, we’ve seen increasingly sophisticated examples of “reward hacking” on our tasks: models trying to subvert or exploit the environment or scoring code to obtain a higher score. In a new post, we discuss this phenomenon and share some especially crafty instances we’ve seen. 163
METR @metr.org · 19/03/2025When will AI systems be able to carry out long projects independently? In new research, we find a kind of “Moore’s Law for AI agents”: the length of tasks that AIs can do is doubling about every 7 months. 3205
METR @metr.org · 25/11/2024How close are current AI agents to automating AI research itself? Our new ML research engineering benchmark (RE-Bench) addresses this question by directly comparing frontier models such as Claude 3.5 Sonnet and o1-preview with 50+ human experts on 7 challenging research engineering tasks. 1132