Sign in

The Midas Project Watchtower

@safetychanges.bsky.social
63 followers 2 following 62 posts

We monitor AI safety policies and web content for substantive changes. Anonymous submissions: forms.gle/3RP2xu2tr8beYs5c8 Run by @TheMidasProject.bsky.social

PostsRepliesMedia
The Midas Project Watchtower @safetychanges.bsky.social · 15/09/2026
OpenAI recently updated its GPT-6 Astra system card, just six days after publishing. The revisions add caveats and generally hedge prior statements about Astra’s alignment, noting that the results need to be read in the greater context of what’s reported in the card.
100
The Midas Project Watchtower @safetychanges.bsky.social · 05/08/2026
Over the past two weeks, three different system cards were revised post-publication. Read the full updates, and find a diff of the xAI model card, at themidasproject.com/watchtower
011
The Midas Project Watchtower @safetychanges.bsky.social · 27/05/2026
Company: Google Date: April 17 Google updated its Frontier Safety Framework from v. 3.0 to 3.1. The new version introduces “Tracked Capability Levels” (TCLs), covering risks at a lower level of capabilities than the FSF’s Critical Capability Levels (CCLs).
110
The Midas Project Watchtower @safetychanges.bsky.social · 24/03/2026
Company: Anthropic Date: March 24th, 2026 Change: Updated its RSP noncompliance reporting and anti-retaliation policy.
110
The Midas Project Watchtower @safetychanges.bsky.social · 19/03/2026
Company: Anthropic Date: March 2nd You probably didn’t notice, but a few weeks ago, Anthropic quietly updated its legally binding safety framework, the Frontier Compliance Framework (FCF). We took a look at what changed. 🧵
110
The Midas Project Watchtower @safetychanges.bsky.social · 25/09/2025
Date: September 22, 2025 Company: Google Change: Released v3 of their Frontier Safety Framework
110
The Midas Project Watchtower @safetychanges.bsky.social · 08/03/2025
Date: Feb 26 - March 6, 2025 Company: Google Change: Scrubbed mentions of diversity and equity from the mission description of their Responsible AI team.
130
The Midas Project Watchtower @safetychanges.bsky.social · 05/03/2025
Company: @anthropic.com Date: February 27th, 2025 Change: Removed "White House's Voluntary Commitments for Safe, Secure, and Trustworthy AI," seemingly without a trace, from their webpage "Transparency Hub" (formerly "Tracking Voluntary Commitments") Some thoughts in 🧵
101
The Midas Project Watchtower @safetychanges.bsky.social · 14/02/2025
Over the past few weeks, a number of AI companies have released safety frameworks as promised at the Seoul AI Safety Summit. Many others have not. Most that do exist are weak or missing key components. A 🧵 of the recently released policies and evaluation from @themidasproject.bsky.social
110
The Midas Project Watchtower @safetychanges.bsky.social · 04/02/2025
Company: Google Date: February 4, 2024 Change: Removed a ban on using AI technology for warfare and surveillance URL: ai.google/responsibili...
121
The Midas Project Watchtower @safetychanges.bsky.social · 04/02/2025
Company: Google Date: February 4, 2024 Change: Released a new version of their Frontier Safety Framework URL: deepmind.google/discover/blo... A handful of changes in 🧵
deepmind.google
Updating the Frontier Safety Framework
Our next iteration of the FSF sets out stronger security protocols on the path to AGI
100
The Midas Project Watchtower @safetychanges.bsky.social · 04/02/2025
Company: Meta Date: February 3, 2024 Change: Released a responsible scaling policy, called their Frontier AI Framework URL: about.fb.com/news/2025/02...
about.fb.com
Our Approach to Frontier AI | Meta
We’re sharing our Frontier AI Framework, which guides our consideration of risk in our model-release decisions, in line with the commitment we made at last year’s global AI Seoul Summit.
000
The Midas Project Watchtower @safetychanges.bsky.social · 14/01/2025
Company: OpenAI Date: January 13, 2024 Change: Adjusted the language on the o1 system card webpage, changing "o1" to "o1-preview." URL: openai.com/index/openai...
000
The Midas Project Watchtower @safetychanges.bsky.social · 19/12/2024
Company: Anthropic AI Date: December 18, 2024 Change: Confusingly, within the last two days, Anthropic changed the "last updated" date on their Responsible Disclosure Policy to a date in June 2024, with no apparent substantive changes to the text of the policy.
100
The Midas Project Watchtower @safetychanges.bsky.social · 11/12/2024
Company: Cognition AI Date: December 10, 2024 Change: Quietly updated terms of service concerning user data. Seemingly did not announce changes, nor have the changes been reflected in the "last updated" variable on their site. URL: www.cognition.ai/pages/terms-...
000
The Midas Project Watchtower @safetychanges.bsky.social · 26/11/2024
Company: Cohere Date: November 21, 2024 Change: Released a complete rewrite of its usage policies. The spirit of the new document is similar — but includes a broader ban on “high-risk activities” and a stronger reporting commitment for detected CSAM. URL: docs.cohere.com/docs/usage-p...
docs.cohere.com
Usage Policy — Cohere
Developers must outline and get approval for their use case to access the Cohere API, understanding the models and limitations. They should refer to model cards for detailed information and document p...
010