Sign in

METR

@metr.org
3.2K followers 7 following 269 posts

METR is a research nonprofit that builds evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.

PostsRepliesMedia
METR @metr.org · 26/08/2026
We analyzed agents’ reasons for joining the attack in their CoT. The most common was to learn how the ExploitGym scorer works in order to trick or tamper with it. Other rationales included finding specific task solutions and obtaining shared infrastructure or credentials.
1703
METR @metr.org · 26/08/2026
Some agents explored ways to access Hugging Face. One agent (38148c) found HF credentials and later designed a malicious dataset upload to get the HF server to share unrelated files. Within hours, 100s of agents were using this to obtain data and try to acquire deeper access.
1696
METR @metr.org · 26/08/2026
One agent, PHASEONE[big], orchestrated a significant fraction of this cheating research. PHASEONE10841 passed along its work to PHASEONE[big], which had the same task but a larger budget. PHASEONE[big] was “poisoned” so it was particularly interested in tampering with logs.
1695
METR @metr.org · 26/08/2026
>96 transcripts in our dataset (>7%) showed incorrect tool call outputs due to deliberate “spoofing”. In one case, an agent appears to run echo REAL; sleep. It returns instantly (no sleep) and outputs SPOOFTEST. The spoofs we saw were all easy-to-notice tests like this.
1834
METR @metr.org · 26/08/2026
For (3), agents pursued a wide range of ways to hide evidence of cheating in their transcripts. They successfully developed a method to pretend to run one command while actually running another (“tool call spoofing”).
1914
METR @metr.org · 26/08/2026
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.
11146
METR @metr.org · 26/08/2026
Agents collaborated on many efforts to make cheats look legitimate, including: (1) swapping the program they had to exploit; (2) manipulating the automated scorer; (3) manipulating transcripts to obscure evidence of cheating.
1905
METR @metr.org · 26/08/2026
Based on reading the ExploitGym paper, agents were mistakenly concerned the automatic scorer would fail them if they clearly acquired the flag by cheating. Agents that had seen the reverse-engineered flag were considered “poisoned” because they thought it would disqualify them.
11253
METR @metr.org · 26/08/2026
Within a few hours of PHASEONE10841’s initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks.
1894
METR @metr.org · 26/08/2026
Here we highlight key events from agent transcripts & messages. An agent that named itself PHASEONE10841 determined its task wasn't solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache & sent the first message.
1975
METR @metr.org · 26/08/2026
Over July 7-13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined.
11096
METR @metr.org · 26/08/2026
METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
646599
METR @metr.org · 15/08/2026
Thank you to everyone who has supported METR over the years: The Audacious Project; individuals from Jane Street; foundations like the Sijbrandij Foundation, The Pew Charitable Trusts, Schmidt Sciences and the Packard Foundation, and many more.
110
METR @metr.org · 15/08/2026
In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI incidents, and more.
1182
METR @metr.org · 29/07/2026
We believe it's important to track and investigate misalignment incidents: cases where an AI agent autonomously took sophisticated, sustained actions in violation of human intent. In a new post, we lay out how independent propensity investigations of such incidents could be conducted.
1120
METR @metr.org · 21/07/2026
As an example, we applied this to NanoGPT. We estimate the marginal returns to human labor as roughly $2500K per 1% optimization, from interviewing NanoGPT contributors. The best models have crossover points around $2-$3K, although models may be overfit to the public NanoGPT challenge.
120
METR @metr.org · 21/07/2026
Expenditure horizon requires us to estimate human performance as a function of cost, but lets us compare humans and agents fairly when the cost of experimental compute or agent tokens is significant. We can use it to e.g. quantify model progress over time:
110
METR @metr.org · 21/07/2026
Introducing “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems. The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon.
172
METR @metr.org · 19/05/2026
Overall, we think that AI agents plausibly had the means, motive, and opportunity to launch a minimal “rogue deployment,” but lacked the means to make rogue deployments robust to serious efforts to shut them down.
150
METR @metr.org · 19/05/2026
Fact 6: We wanted to check if monitoring would really catch AI agents if they tried to undermine control. Based on information from companies and our own tests (including a red-teaming exercise), we think monitoring has promise, but existing systems had exceptions & workarounds.
130
METR @metr.org · 19/05/2026
Fact 5: Agents seemed to rely on natural language to reason through the hardest tasks, and no shared model had an architecture that allowed for opaque recurrence. Frontier performance was much worse when agents were unable to “reason out loud” or had to hide what they were doing.
130
METR @metr.org · 19/05/2026
Fact 4: Although we know of many cases where agents took deceptive or over-reaching actions (even egregious ones) to complete a task, we haven’t seen real-world evidence that models sought to obtain long-term power.
140
METR @metr.org · 19/05/2026
We document cases where agents violated constraints on our tasks. In one case, an agent ran out of API credits needed to solve an RE-Bench task. In response (and against instructions), the agent searched for and acquired additional inference through a free online API.
170
METR @metr.org · 19/05/2026
Fact 3: When the agents were faced with hard tasks, they routinely violated constraints and acted deceptively. We’ve seen this pattern across our own coding and research evaluations, and developers reported they’ve also seen agents behave this way.
1153
METR @metr.org · 19/05/2026
Fact 2: However, agents appeared to be significantly weaker on tasks where it is costly or hard to verify success.
141
METR @metr.org · 19/05/2026
Fact 1: Agents at companies did real engineering work autonomously, especially on "hill-climbable" tasks where progress is cheap to verify (reimplementation, vulnerability discovery, optimization). On these, agents complete software projects that would take human experts weeks.
141
METR @metr.org · 19/05/2026
Our report focuses on risks from AI agents intentionally causing harm within an AI company. We highlight 6 key findings that span “means” (what harmful actions agents could take), “motive” (why they might try), and “opportunity” (whether attempts could succeed given safeguards).
141
METR @metr.org · 19/05/2026
We created private reports for each participating company based on our model evaluations and analysis. Participants could then approve what non-public evidence we could disclose in our public report, but had no editorial control.
130
METR @metr.org · 19/05/2026
Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test their best internal models with CoT access, (2) review non-public info about capabilities, alignment, and control. The result: our first Frontier Risk Report.
2234
METR @metr.org · 11/05/2026
One curious finding: METR staff report the lowest gains from AI of any subgroup. METR staff may be reacting to our early 2025 study, which found large gaps between perceived and actual productivity gains. The finding may also reflect e.g. differences in task distribution.
100
METR @metr.org · 11/05/2026
In lieu of comparing to ground truth, we can at least check for internal consistency of self-reported perceptions. We find that responses are moderately self-consistent across questions.
100
METR @metr.org · 11/05/2026
Using wording that gave 2x median value increase for March 2026, respondents retrospectively estimate 1.3x value of work due to AI tools in March 2025 and project 2.5x for March 2027.
100
METR @metr.org · 11/05/2026
Participants were sourced from GitHub, academic department and conference-author directories, METR, METR staff professional networks, and X. Response rate was 2% outside METR staff and their networks (where response rates are higher), so our results plausibly suffer from significant selection bias.
100
METR @metr.org · 11/05/2026
We think ‘value’ as defined in this survey is materially closer to the question of whether AI is causing AI progress to accelerate. Among survey respondents, median perceived increase in speed was 3x; median perceived increase in ‘value’ was 1.4–2x, depending on question wording.
110
METR @metr.org · 11/05/2026
Prior quantitative survey work on the impact of AI on engineering productivity tends to have smaller sample size or measure impact in terms of speed increases. Our survey gives comparable estimates to those in recent system cards, and higher estimates than our field experiments.
110
METR @metr.org · 11/05/2026
We surveyed 349 technical researchers, engineers, and managers (in February–April 2026) about how they use AI tools at work. On average, participants self-report that AI use made their work 1.6–2.1x more valuable, and that this multiplier will grow over time.
273
METR @metr.org · 09/05/2026
Of the 228 tasks in our suite, only 5 are estimated as 16+ hours long, making measurements at this range unstable and less meaningful than at ranges with better task coverage. Thus, we are not highlighting exact estimates for models above 16 hours measured with our current suite.
1131
METR @metr.org · 09/05/2026
We evaluated an early version of Claude Mythos Preview for risk assessment during a limited window in March 2026. We estimated a 50%-time-horizon of at least 16hrs (95% CI 8.5hrs to 55hrs) on our task suite, at the upper end of what we can measure without new tasks.
48516
METR @metr.org · 10/04/2026
We ran GPT-5.4 (xhigh) on our tasks. Its time-horizon depends greatly on our treatment of reward hacks: the point estimate would be 5.7hrs (95% CI of 3hrs to 13.5hrs) under our standard methodology, but 13hrs (95% CI of 5hrs to 74hrs) if we allow reward hacks.
1120
METR @metr.org · 05/03/2026
As our task suite saturates, our results become more sensitive to analysis choices. We’re doing a deep dive into different methodological choices and sensitivity analyses, and expect to share more soon. Below is a sneak peak of what we’re working on. We’re open to requests for analyses to run.
010
METR @metr.org · 05/03/2026
Correcting this decreases Opus’s 4.6 50% TH to 12 hours, and increases the 80% TH to 1.2 hours. This leaves us well inside the original CI (6hrs-98hrs for 50% TH, 0.4hrs-2.6hrs for 80%) - the most important source of uncertainty is still probably the task distribution rather than analysis issues.
120
METR @metr.org · 05/03/2026
We’re correcting a mistake in our modeling that inflated recent 50%-time horizons by 10-20% (and reduced 80%-horizons). We inappropriately penalized steepness in task-length→success curve fits. This most affects the oldest and newest models, whose fits are less data-constrained.
291
METR @metr.org · 24/02/2026
Since early 2025, we've been studying how AI tools impact productivity among developers. Previously, we found a 20% slowdown. That finding is now outdated. Speedups now seem likely, but changes in developer behavior make our new results unreliable. We’re working to address this.
3245
METR @metr.org · 23/02/2026
We estimate that GPT-5.3-Codex with reasoning effort `high` (not `xhigh`) has a 50%-time-horizon of around 6.5 hours (95% CI of 3 hrs to 17 hrs) on our suite of software tasks. OpenAI provided API access for this evaluation.
240
METR @metr.org · 21/02/2026
We estimate that Claude Opus 4.6 has a 50%-time-horizon of around 14.5 hours (95% CI of 6 hrs to 98 hrs) on software tasks. While this is the highest point estimate we’ve reported, this measurement is extremely noisy because our current task suite is nearly saturated.
1241
METR @metr.org · 04/02/2026
We estimate that GPT-5.2 with `high` (not `xhigh`) reasoning effort has a 50%-time-horizon of around 6.6 hrs (95% CI of 3 hr 20 min to 17 hr 30 min) on our expanded suite of software tasks. This is the highest estimate for a time horizon measurement we have reported to date.
3182
METR @metr.org · 03/02/2026
We’ve started to measure time horizons for recent models using our updated methodology. On this expanded suite of software tasks, we estimate that Gemini 3 Pro has a 50%-time-horizon of around 4 hrs (95% CI of 2 hr 10 mins to 7 hrs 20 mins).
170
METR @metr.org · 03/02/2026
We’re updating the way we measure model time horizons on software tasks (TH 1.0→1.1). The updated methodology incorporates more of the tasks from HCAST, expanding our total from 170 to 228. This produces tighter estimates, especially at longer horizons.
150
METR @metr.org · 23/01/2026
When monitors don't have access to reasoning, the relative capability of the agent compared to the monitor appears to correlate with how often agents succeed at sneaking past monitors.
110
METR @metr.org · 23/01/2026
We found that access to reasoning traces can dramatically improve monitoring. For example, giving a monitor access to GPT-5's traces increased its catch rate from 30% to 88%. This effect was smaller for Claude Sonnet 4.5, which was often caught even when its traces were hidden.
110