Sign in

The Midas Project Watchtower

@safetychanges.bsky.social
63 followers 2 following 62 posts

We monitor AI safety policies and web content for substantive changes. Anonymous submissions: forms.gle/3RP2xu2tr8beYs5c8 Run by @TheMidasProject.bsky.social

PostsRepliesMedia
The Midas Project Watchtower @safetychanges.bsky.social · 15/09/2026
OpenAI recently updated its GPT-6 Astra system card, just six days after publishing. The revisions add caveats and generally hedge prior statements about Astra’s alignment, noting that the results need to be read in the greater context of what’s reported in the card.
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/08/2026
Over the past two weeks, three different system cards were revised post-publication. Read the full updates, and find a diff of the xAI model card, at themidasproject.com/watchtower
011
The Midas Project Watchtower @safetychanges.bsky.social · 27/05/2026
Company: Google Date: April 17 Google updated its Frontier Safety Framework from v. 3.0 to 3.1. The new version introduces “Tracked Capability Levels” (TCLs), covering risks at a lower level of capabilities than the FSF’s Critical Capability Levels (CCLs).
110
The Midas Project Watchtower @safetychanges.bsky.social · 24/03/2026
Company: Anthropic Date: March 24th, 2026 Change: Updated its RSP noncompliance reporting and anti-retaliation policy.
110
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Luckily, we have our own copy. We’ve made a diff of the changes that you can visit at our website (themidasproject.com/watchtower/a...). Below are the key sections that were modified:
Redline of changes in FCF
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Under California’s SB 53, AI developers must "clearly and conspicuously publish" material modifications to their safety frameworks within 30 days. The statute includes this language to guarantee public oversight when companies alter their binding commitments.
SB 53 text
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
That being said, if you click through to read the document, you will find a changelog at the bottom revealing that a new version has been uploaded (although it doesn’t provide many concrete details of what’s new in V2, or why).
FCF changelog
100
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Company: Anthropic Date: March 2nd You probably didn’t notice, but a few weeks ago, Anthropic quietly updated its legally binding safety framework, the Frontier Compliance Framework (FCF). We took a look at what changed. 🧵
110
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
Similarly, for ML R&D, models that "can" accelerate AI development no longer require RAND SL 3. Only models that have been used for this purpose count. But this is a strange ordering -- shouldn't the safeguards precede the deployment (and even the training) of such a model?
100
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
But it's weakened in other ways. Critical capability levels, which previously focused on capabilities (e.g. "can be used to cause a mass casualty event") now seems to rely on anticipated outcomes (e.g. "resulting in additional expected harm at severe scale")
100
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
Date: September 22, 2025 Company: Google Change: Released v3 of their Frontier Safety Framework
110
The Midas Project Watchtower @safetychanges.bsky.social · 08/03/2025
Date: Feb 26 - March 6, 2025 Company: Google Change: Scrubbed mentions of diversity and equity from the mission description of their Responsible AI team.
130
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
(Removed)
000
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
(Removed)
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
The smaller changes made to Anthropic's practices: (Added)
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
Most surprisingly, there is now no record of the former commitments on Anthropic's transparency center, a web resource they launched to track their compliance with voluntary commitments and which they describe as "raising the bar on transparency."
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
In fact, post-election, multiple tech companies confirmed their commitments hadn't changed. Perhaps they understood that the commitments were not contingent on whatever way the political winds blow, but made to the public at large. fedscoop.com/voluntary-ai...
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
Company: @anthropic.com Date: February 27th, 2025 Change: Removed "White House's Voluntary Commitments for Safe, Secure, and Trustworthy AI," seemingly without a trace, from their webpage "Transparency Hub" (formerly "Tracking Voluntary Commitments") Some thoughts in 🧵
101
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: xAI Date: February 10, 2025 Change: Released Risk Management Framework draft URL: x.ai/documents/20... xAI's policy is stronger than others in terms of using specific benchmarks, but lacks threshold details, and provides no mitigations.
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Amazon Date: February 10, 2025 Change: Released their Frontier Model Safety Framework URL: amazon.science/publications... Like Microsoft, Amazon's policy also goes through the motions while setting vague thresholds that aren't clearly connected to specific mitigations
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Microsoft Date: February 8, 2025 Change: Released their Frontier Governance Framework URL: cdn-dynmedia-1.microsoft.com/is/content/m... Microsoft's policy is an admirable effort, but as with others, needs further specification. Mitigations should also be connected to specific thresholds
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Cohere Date: February 7th, 2025 Change: Released their "Secure AI Frontier Model Framework" URL: cohere.com/security/the... Cohere's framework mostly neglects the most important risks. Like G42, they are not developing frontier models, which makes this more understandable.
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: G42 Date: February 6, 2025 Change: Released their "Frontier AI Framework" URL: g42.ai/application/... G42's policy is surprisingly strong for a non-frontier lab. It's biggest issues are a lack of specificity and not defining future thresholds for catastrophic risks.
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Google Date: February 4, 2025 Change: Released v2 of their Frontier Safety Framework URL: deepmind.google/discover/blo... v2 of the framework improves Google's policy in some areas while weakening it in others, most notably no longer promising to adhere to it if others are not.
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Company: Meta Date: February 3rd, 2025 Change: Released their "Frontier AI Framework" URL: ai.meta.com/static-resou... Meta's policy includes risk thresholds and a commitment to pause development, but is severely weakened by caveats, loopholes, and no mitigations named
100
The Midas Project Watchtower @safetychanges.bsky.social · 04/02/2025
Company: Google Date: February 4, 2024 Change: Removed a ban on using AI technology for warfare and surveillance URL: ai.google/responsibili...
121
The Midas Project Watchtower @safetychanges.bsky.social · 04/02/2025
Worryingly, Google now essentially says that they only plan to follow their framework if other companies are following similar ones. This would go against the promise they made to the White House, and later to the UK + Korea, to adhere to a risk framework without qualification.
000
The Midas Project Watchtower @safetychanges.bsky.social · 04/02/2025
Security mitigations are now connected to risk thresholds. In the previous version of the policy, there was no logical relationship between the two.
100
The Midas Project Watchtower @safetychanges.bsky.social · 04/02/2025
Model autonomy risks have been removed and replaced with alignment risks, but the details are vague and yet-to-be fleshed out. The current policy focuses primarily on misuse risks.
100
The Midas Project Watchtower @safetychanges.bsky.social · 14/01/2025
Company: OpenAI Date: January 13, 2024 Change: Adjusted the language on the o1 system card webpage, changing "o1" to "o1-preview." URL: openai.com/index/openai...
000
The Midas Project Watchtower @safetychanges.bsky.social · 19/12/2024
Company: Anthropic AI Date: December 18, 2024 Change: Confusingly, within the last two days, Anthropic changed the "last updated" date on their Responsible Disclosure Policy to a date in June 2024, with no apparent substantive changes to the text of the policy.
100
The Midas Project Watchtower @safetychanges.bsky.social · 11/12/2024
Company: Cognition AI Date: December 10, 2024 Change: Quietly updated terms of service concerning user data. Seemingly did not announce changes, nor have the changes been reflected in the "last updated" variable on their site. URL: www.cognition.ai/pages/terms-...
000