Sign in

The Midas Project Watchtower

@safetychanges.bsky.social
63 followers 2 following 62 posts

We monitor AI safety policies and web content for substantive changes. Anonymous submissions: forms.gle/3RP2xu2tr8beYs5c8 Run by @TheMidasProject.bsky.social

PostsRepliesMedia
The Midas Project Watchtower @safetychanges.bsky.social · 15/09/2026
Check out our Watchtower post for the diff: www.themidasproject.com/watchtower/o...
themidasproject.com
OpenAI, Moderate Change on Sep 09, 2026 | Watchtower | The Midas Project
Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.
020
The Midas Project Watchtower @safetychanges.bsky.social · 15/09/2026
OpenAI recently updated its GPT-6 Astra system card, just six days after publishing. The revisions add caveats and generally hedge prior statements about Astra’s alignment, noting that the results need to be read in the greater context of what’s reported in the card.
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/08/2026
Over the past two weeks, three different system cards were revised post-publication. Read the full updates, and find a diff of the xAI model card, at themidasproject.com/watchtower
011
The Midas Project Watchtower @safetychanges.bsky.social · 27/05/2026
A full diff is available at our website: themidasproject.com/watchtower/g...
themidasproject.com
Google | The Midas Project
Google updated its Frontier Safety Framework from version 3.0 to 3.1, in a change announced on its website.
010
The Midas Project Watchtower @safetychanges.bsky.social · 27/05/2026
FSF v. 3.1 also includes a thin section on “Governance and Accountability” which fails to name any specific governance or accountability mechanisms (though Google has said more on this elsewhere: deepmind.google/responsibili...)
deepmind.google
Responsibility & Safety
AI can provide extraordinary benefits, but like all transformational technology, it could have negative impacts unless it’s developed and deployed responsibly.
110
The Midas Project Watchtower @safetychanges.bsky.social · 27/05/2026
Still, it’s an improvement over v. 3.0, which just described its misalignment CCLs as an “illustrative” example.
120
The Midas Project Watchtower @safetychanges.bsky.social · 27/05/2026
It’s notable that this doesn’t rise to the level of a full CCL. Google is essentially saying when a model reaches this risk threshold: If we don’t put additional safeguards in place, we might lose control of the model… but we’re not going to require a formal safety case for it.
100
The Midas Project Watchtower @safetychanges.bsky.social · 27/05/2026
TCLs trigger risk assessments and mitigations, but don’t require formal safety cases like CCLs do. A misalignment TCL is defined when models have enough situational awareness and stealth that “absent additional mitigations, we cannot rule out the model significantly undermining human control.”
120
The Midas Project Watchtower @safetychanges.bsky.social · 27/05/2026
Company: Google Date: April 17 Google updated its Frontier Safety Framework from v. 3.0 to 3.1. The new version introduces “Tracked Capability Levels” (TCLs), covering risks at a lower level of capabilities than the FSF’s Critical Capability Levels (CCLs).
110
The Midas Project Watchtower @safetychanges.bsky.social · 24/03/2026
The changes largely seem like an improvement over the former policy, and more frontier AI companies ought to release similar guidance for their employees. A full diff is available on our website at www.themidasproject.com/watchtower/a...
themidasproject.com
Anthropic | The Midas Project
Updated its RSP noncompliance reporting and anti-retaliation policy
010
The Midas Project Watchtower @safetychanges.bsky.social · 24/03/2026
Company: Anthropic Date: March 24th, 2026 Change: Updated its RSP noncompliance reporting and anti-retaliation policy.
110
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Read the full diff on our website: www.themidasproject.com/watchtower/a...
themidasproject.com
Anthropic | The Midas Project
Updated its Frontier Compliance Framework without public announcement
000
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
But a broader point, in which we are confident, is that companies should take the “clear and conspicuous” requirement for SB 53 far more seriously. Updates to their safety frameworks ought to be as legible and well-justified as they can muster.
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
On the other hand, narrowing the scope for AI R&D from measuring a model's general autonomous capabilities to a few specific fields may weaken it.
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Do these changes make the policy stronger? This is up for debate. Clearly, adding thresholds for harmful manipulation is an improvement over not having any.
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Luckily, we have our own copy. We’ve made a diff of the changes that you can visit at our website (themidasproject.com/watchtower/a...). Below are the key sections that were modified:
Redline of changes in FCF
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Making matters worse, unlike the RSP, Anthropic doesn’t publish past FCF versions to easily compare the text. While the new document includes a brief changelog at the end, the previous December 2025 framework has essentially been overwritten and erased from the trust center.
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Instead of a public announcement, Anthropic slipped the updated file into its trust center and left it off its update log. Does quietly overwriting a PDF (even with a changelog at the bottom) really satisfy the legal requirement for clear and conspicuous disclosure?
110
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Under California’s SB 53, AI developers must "clearly and conspicuously publish" material modifications to their safety frameworks within 30 days. The statute includes this language to guarantee public oversight when companies alter their binding commitments.
SB 53 text
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
That being said, if you click through to read the document, you will find a changelog at the bottom revealing that a new version has been uploaded (although it doesn’t provide many concrete details of what’s new in V2, or why).
FCF changelog
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Anthropic seemingly did not announce the update. Clicking the link in the Dec. blog post which first revealed the policy sends you to their trust center, where the new document lives with no obvious mention of the change (including no mention in the trust center’s update log!)
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Company: Anthropic Date: March 2nd You probably didn’t notice, but a few weeks ago, Anthropic quietly updated its legally binding safety framework, the Frontier Compliance Framework (FCF). We took a look at what changed. 🧵
110
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
On the whole, it's good that Google is continuing to update its risk management policies, and they seem to treat the issue with much more seriousness than some competitors. Read the full diff at our website: www.themidasproject.com/watchtower/g...
themidasproject.com
Google | The Midas Project
Updated their Frontier Safety Framework
010
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
Remember that in 2024 Google promised to *define* specific risk thresholds, not explore illustrative examples.
120
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
Additionally, as pointed out by Zach Stein-Perlman of AI Lab Watch, the CCLs for misalignment, which used to be a concrete (albeit initial) approach, are now described as "exploratory" and "illustrative."
120
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
Similarly, for ML R&D, models that "can" accelerate AI development no longer require RAND SL 3. Only models that have been used for this purpose count. But this is a strange ordering -- shouldn't the safeguards precede the deployment (and even the training) of such a model?
100
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
But it's weakened in other ways. Critical capability levels, which previously focused on capabilities (e.g. "can be used to cause a mass casualty event") now seems to rely on anticipated outcomes (e.g. "resulting in additional expected harm at severe scale")
100
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
In their blog post, Google describes this as a strengthening of the policy. And in some ways, it is: they define a new harmful manipulation risk category, and they even soften the claim from v2 that they would only follow their promise if every other company does so as well.
100
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
Date: September 22, 2025 Company: Google Change: Released v3 of their Frontier Safety Framework
110
The Midas Project Watchtower @safetychanges.bsky.social · 08/03/2025
Old page: web.archive.org/web/20250206... Current page: research.google/teams/respon...
web.archive.org
Responsible AI
The mission of the Responsible AI and Human Centered Technology (RAI-HCT) team is to conduct research and develop methodologies, technologies, and best practices to ensure AI systems are built respons...
020
The Midas Project Watchtower @safetychanges.bsky.social · 08/03/2025
Date: Feb 26 - March 6, 2025 Company: Google Change: Scrubbed mentions of diversity and equity from the mission description of their Responsible AI team.
130
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
(Removed)
000
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
(Removed)
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
The smaller changes made to Anthropic's practices: (Added)
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
The good news is that the details they provide on internal practices have changed very little (select screenshots included in rest of thread). Now all they need to do is provide transparency on *all* the commitments they've made + when they are choosing to abandon any.
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
Most surprisingly, there is now no record of the former commitments on Anthropic's transparency center, a web resource they launched to track their compliance with voluntary commitments and which they describe as "raising the bar on transparency."
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
In fact, post-election, multiple tech companies confirmed their commitments hadn't changed. Perhaps they understood that the commitments were not contingent on whatever way the political winds blow, but made to the public at large. fedscoop.com/voluntary-ai...
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
While there is a new administration in office, nothing in the commitments suggested that the promise was (1) time-bound or (2) contingent on the party affiliation of the sitting president.
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
The White House Voluntary Commitments, made in 2023, were a pledge to conduct pre-deployment testing, share information on AI risk management frameworks, invest in cybersecurity, implement bug bounties, and publicly report capabilities and limitations. bidenwhitehouse.archives.gov/briefing-roo...
bidenwhitehouse.archives.gov
FACT SHEET: Biden-Harris Administration Secures Voluntary Commitments from Leading Artificial Intelligence Companies to Manage the Risks Posed by AI | The White House
Voluntary commitments – underscoring safety, security, and trust – mark a critical step toward developing responsible AIBiden-Harris Administration will
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
Company: @anthropic.com Date: February 27th, 2025 Change: Removed "White House's Voluntary Commitments for Safe, Secure, and Trustworthy AI," seemingly without a trace, from their webpage "Transparency Hub" (formerly "Tracking Voluntary Commitments") Some thoughts in 🧵
101
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
To view all of the policies released by AI companies, and their scorecards, check out our full report at www.seoul-tracker.org
seoul-tracker.org
Seoul Commitment Tracker
Tracking the progress of voluntary commitments made at the 2023 AI Safety Summit in Seoul, South Korea
000
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: xAI Date: February 10, 2025 Change: Released Risk Management Framework draft URL: x.ai/documents/20... xAI's policy is stronger than others in terms of using specific benchmarks, but lacks threshold details, and provides no mitigations.
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Amazon Date: February 10, 2025 Change: Released their Frontier Model Safety Framework URL: amazon.science/publications... Like Microsoft, Amazon's policy also goes through the motions while setting vague thresholds that aren't clearly connected to specific mitigations
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Microsoft Date: February 8, 2025 Change: Released their Frontier Governance Framework URL: cdn-dynmedia-1.microsoft.com/is/content/m... Microsoft's policy is an admirable effort, but as with others, needs further specification. Mitigations should also be connected to specific thresholds
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Cohere Date: February 7th, 2025 Change: Released their "Secure AI Frontier Model Framework" URL: cohere.com/security/the... Cohere's framework mostly neglects the most important risks. Like G42, they are not developing frontier models, which makes this more understandable.
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: G42 Date: February 6, 2025 Change: Released their "Frontier AI Framework" URL: g42.ai/application/... G42's policy is surprisingly strong for a non-frontier lab. It's biggest issues are a lack of specificity and not defining future thresholds for catastrophic risks.
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Google Date: February 4, 2025 Change: Released v2 of their Frontier Safety Framework URL: deepmind.google/discover/blo... v2 of the framework improves Google's policy in some areas while weakening it in others, most notably no longer promising to adhere to it if others are not.
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Meta Date: February 3rd, 2025 Change: Released their "Frontier AI Framework" URL: ai.meta.com/static-resou... Meta's policy includes risk thresholds and a commitment to pause development, but is severely weakened by caveats, loopholes, and no mitigations named
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Over the past few weeks, a number of AI companies have released safety frameworks as promised at the Seoul AI Safety Summit. Many others have not. Most that do exist are weak or missing key components. A 🧵 of the recently released policies and evaluation from @themidasproject.bsky.social
110
The Midas Project Watchtower @safetychanges.bsky.social · 04/02/2025
H/T to @wired.com for first reporting on the change earlier today www.wired.com/story/google...
wired.com
Google Lifts a Ban on Using Its AI for Weapons and Surveillance
Google published principles in 2018 barring its AI technology from being used for sensitive purposes. Weeks into President Donald Trump’s second term, those guidelines are being overhauled.
000