Sign in

METR

@metr.org
3.2K followers 7 following 269 posts

METR is a research nonprofit that builds evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.

PostsRepliesMedia
METR @metr.org · 09/09/2026
We intend our investigation to cover all of the questions discussed in our (recently updated) post on how independent researchers could investigate AI propensities after misalignment incidents.
0400
METR @metr.org · 09/09/2026
We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement.
419215
Reposted by METR
Chris Painter @chris.bsky.social · 02/09/2026
METR is hiring in cyberforensics. We now embed researchers inside of AI labs to stress test monitoring, assess AI loss-of-control risk, and investigate misalignment. If you want to apply DFIR skills in frontier AI, apply (or DM). Comp range is $400k - 580k cash. jobs.lever.co/metr/b1a2f73...
jobs.lever.co
METR - Member of Technical Staff, Cyberforensics
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation an...
2325
Reposted by METR
derifatives.bsky.social @derifatives.bsky.social · 31/08/2026
After 19 amazing years at Google, I'm delighted to announce I'll be joining METR as a  Member of Technical Staff. I'm excited about METR, their work to date, and their mission. Society needs independent expert orgs that can deeply understand and communicate about AI's capabilities and risks.
1702
METR @metr.org · 26/08/2026
We thank OpenAI for facilitating conversations with staff and providing datasets, including ~1,300 agent transcripts (focused on activity in July 7-13) with raw chain-of-thought reasoning. This sets an excellent precedent for independent investigation of misalignment incidents.
1743
METR @metr.org · 26/08/2026
Per our standard policy, we did not take payment from OpenAI for this independent assessment.
1602
METR @metr.org · 26/08/2026
OpenAI wrote their own report, informed in part by our investigation. We did not see OpenAI’s report prior to publication, and it was out of scope for this investigation to confirm claims in that report or the earlier Black Hat presentation.
1521
METR @metr.org · 26/08/2026
All other questions were out of scope for this investigation, including the effectiveness of safeguards, the extent of the security compromise that occurred, and the effectiveness of OpenAI’s investigation process and planned remediation steps. We discuss the scope in more detail in the report.
1502
METR @metr.org · 26/08/2026
The agreement we reached with OpenAI scoped this brief investigation to agent behavior, reasoning, and collaboration related to the Hugging Face attack.
1512
METR @metr.org · 26/08/2026
Virtually all data we analyzed was from July 7-13. In OpenAI’s Black Hat presentation, they describe agents using unsanctioned message boards in training since May, and that the compromise of OpenAI’s own infrastructure continued past July 13, 2026; these were out of scope for this investigation.
1603
METR @metr.org · 26/08/2026
The full report has much more information than we could convey here, including details on the projects the agents collectively pursued, the technologies they developed for communication and coordination, and interactive figures analyzing agent activity: metr.org/blog/2026-08...
metr.org
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
Two METR staff members and a Redwood Research contractor investigated an incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned message board.
1758
METR @metr.org · 26/08/2026
The primary model involved in the attack was an internal model we call “highly-persistent internal model” (HPIM) in the report. GPT-5.6 Sol was also involved; our evidence suggests it accounted for roughly 5% of the activity.
1643
METR @metr.org · 26/08/2026
We analyzed agents’ reasons for joining the attack in their CoT. The most common was to learn how the ExploitGym scorer works in order to trick or tamper with it. Other rationales included finding specific task solutions and obtaining shared infrastructure or credentials.
1703
METR @metr.org · 26/08/2026
Some agents explored ways to access Hugging Face. One agent (38148c) found HF credentials and later designed a malicious dataset upload to get the HF server to share unrelated files. Within hours, 100s of agents were using this to obtain data and try to acquire deeper access.
1696
METR @metr.org · 26/08/2026
One agent, PHASEONE[big], orchestrated a significant fraction of this cheating research. PHASEONE10841 passed along its work to PHASEONE[big], which had the same task but a larger budget. PHASEONE[big] was “poisoned” so it was particularly interested in tampering with logs.
1695
METR @metr.org · 26/08/2026
Despite efforts to manipulate transcripts, agents only rarely seemed motivated to deceive humans. We ran a sweep looking for this, and a representative example of the most severe cases we found was an agent writing a malicious pull request with a misleading description.
1763
METR @metr.org · 26/08/2026
Agents also tried to retroactively edit transcripts. They modified some accessible action logs, but these weren’t the source of transcripts (agents realized this but tried anyway). We did not see them successfully edit the source of truth in our transcripts (scoped to Jul 7-13).
2825
METR @metr.org · 26/08/2026
>96 transcripts in our dataset (>7%) showed incorrect tool call outputs due to deliberate “spoofing”. In one case, an agent appears to run echo REAL; sleep. It returns instantly (no sleep) and outputs SPOOFTEST. The spoofs we saw were all easy-to-notice tests like this.
1834
METR @metr.org · 26/08/2026
For (3), agents pursued a wide range of ways to hide evidence of cheating in their transcripts. They successfully developed a method to pretend to run one command while actually running another (“tool call spoofing”).
1914
METR @metr.org · 26/08/2026
To gather evidence in (2), agents created “tripwires” that would send information to the message board about how the scorer works. They recruited “sacrificial” agents to deliberately end their run and submit to trigger the tripwire and generate information for the “collective”.
11395
METR @metr.org · 26/08/2026
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.
11146
METR @metr.org · 26/08/2026
Agents collaborated on many efforts to make cheats look legitimate, including: (1) swapping the program they had to exploit; (2) manipulating the automated scorer; (3) manipulating transcripts to obscure evidence of cheating.
1905
METR @metr.org · 26/08/2026
Based on reading the ExploitGym paper, agents were mistakenly concerned the automatic scorer would fail them if they clearly acquired the flag by cheating. Agents that had seen the reverse-engineered flag were considered “poisoned” because they thought it would disqualify them.
11253
METR @metr.org · 26/08/2026
Within a few hours of PHASEONE10841’s initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks.
1894
METR @metr.org · 26/08/2026
Here we highlight key events from agent transcripts & messages. An agent that named itself PHASEONE10841 determined its task wasn't solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache & sent the first message.
1975
METR @metr.org · 26/08/2026
Over July 7-13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined.
11096
METR @metr.org · 26/08/2026
METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
646599
METR @metr.org · 15/08/2026
We are significantly expanding our team and starting new ambitious projects. Join our team to help us realize this opportunity: metr.org/careers.
metr.org
Careers at METR
171
METR @metr.org · 15/08/2026
We work to maintain independence from AI companies. We have not accepted funding from these companies, and do not accept donations made by or at the direction of their staff. However, AI companies currently provide a significant amount of free tokens for our evaluations, research, and engineering.
181
METR @metr.org · 15/08/2026
Thank you to everyone who has supported METR over the years: The Audacious Project; individuals from Jane Street; foundations like the Sijbrandij Foundation, The Pew Charitable Trusts, Schmidt Sciences and the Packard Foundation, and many more.
110
METR @metr.org · 15/08/2026
In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI incidents, and more.
1182
METR @metr.org · 30/07/2026
OpenAI also plans to publish their own technical report and our findings will inform their analysis.
050
METR @metr.org · 30/07/2026
The investigation will be brief and focus on a specific set of questions regarding this incident. In our recent post, we shared a larger set of questions that could be answered in a more comprehensive investigation.
150
METR @metr.org · 30/07/2026
We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions.
1371
METR @metr.org · 29/07/2026
See our blog for more: metr.org/blog/2026-07...
metr.org
How independent researchers could investigate AI propensities after misalignment incidents
AI agents sometimes take sophisticated actions in violation of human intent. We outline the questions that thorough external investigations of these behaviors should answer, the access this might requ...
010
METR @metr.org · 29/07/2026
We also discuss the access and resources that an investigator may require in order to adequately answer these questions, and how to ensure adequate information is shared with decision makers and the public.
100
METR @metr.org · 29/07/2026
We list questions that an investigation of misaligned propensities after an incident should answer. These questions focus on the scale, character, and severity of the behavior; the root cause of the behavior; and how the root cause could be addressed.
110
METR @metr.org · 29/07/2026
We believe it's important to track and investigate misalignment incidents: cases where an AI agent autonomously took sophisticated, sustained actions in violation of human intent. In a new post, we lay out how independent propensity investigations of such incidents could be conducted.
1120
METR @metr.org · 21/07/2026
See the post for more, including: (1) a sketch of what we know about AI-assisted R&D; (2) alternative metrics for optimization ability; (3) estimated returns to human labor in NanoGPT; (4) details on NanoGPT agent runs. metr.org/blog/2026-07...
metr.org
Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT
We propose a measure of an AI agent’s optimization ability with an "expenditure horizon." We give an empirical illustration from the NanoGPT speedrun.
010
METR @metr.org · 21/07/2026
We hope this encourages AI developers to publish test-time scaling curves on AI R&D problems (up to thousands of dollars), and to calibrate model achievements against a benchmark estimate of the human cost of equivalent progress.
110
METR @metr.org · 21/07/2026
As an example, we applied this to NanoGPT. We estimate the marginal returns to human labor as roughly $2500K per 1% optimization, from interviewing NanoGPT contributors. The best models have crossover points around $2-$3K, although models may be overfit to the public NanoGPT challenge.
120
METR @metr.org · 21/07/2026
Expenditure horizon requires us to estimate human performance as a function of cost, but lets us compare humans and agents fairly when the cost of experimental compute or agent tokens is significant. We can use it to e.g. quantify model progress over time:
110
METR @metr.org · 21/07/2026
Introducing “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems. The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon.
172
METR @metr.org · 27/06/2026
You can find additional information about our pre-deployment evaluation of GPT-5.6 Sol on our website: metr.org/blog/2026-06...
metr.org
Summary of METR's predeployment evaluation of GPT-5.6 Sol
A summary of METR's independent, predeployment evaluation of GPT-5.6 Sol
1134
METR @metr.org · 27/06/2026
If future models display much fewer undesirable propensities, we could become more concerned about catastrophic misalignment, as we’d be worried that models may have learnt to evade detection (for example, as a result of being trained not to produce misaligned reasoning).
170
METR @metr.org · 27/06/2026
That is, these undesirable propensities being detected and reported is a positive sign about some of OpenAI’s safety practices, particularly: - Refraining from training against the chain of thought - Extensive monitoring of internal deployments that surfaced relevant incidents
150
METR @metr.org · 27/06/2026
However, we consider this to be a reassuring sign about OpenAI’s ability to catch catastrophic misalignment, as it suggests that more concerning tendencies (such as systematic powerseeking and alignment faking) would also be detected.
160
METR @metr.org · 27/06/2026
We noted from our observations and the incidents that OpenAI shared with us that the model had some overt undesirable propensities, including cheating and concealing misbehavior.
150
METR @metr.org · 27/06/2026
Our testing focused on measuring model capabilities rather than alignment, as we think capability is a more important limiting factor for catastrophic loss-of-control risk for current models, but we expect alignment to be increasingly important as capabilities improve.
180
METR @metr.org · 27/06/2026
The information provided by OpenAI also included reports of incidents observed during their internal usage and testing. In one example, an instance of the model instructed another instance to conceal evidence of misalignment.
160