Cas (Stephen Casper) @scasper.bsky.social · 25/09/2026🧵 I read the Sanders/Casar "Ban Artificial Superintelligence Act of 2026" I wouldn't pass this in its current form, but there are some things I like about it. Thoughts here in thread. 100
Cas (Stephen Casper) @scasper.bsky.social · 22/09/2026It's great that there has been so much recent discussion about embedded evaluations at frontier AI companies. But, we should be very wary of how things progress because there are big 🚩 red flags 🚩 for potential regulatory capture. 120
Cas (Stephen Casper) @scasper.bsky.social · 21/09/2026CMU CASI a talk of mine on YouTube. It has my hottest takes on doing research that matters. It also features inspiration from UPS drivers, Tina Fey, and a viral internet dance video about fruit. (Slide troubles at first, but they are fixed 10 mins in.) www.youtube.com/watch?v=0U5...youtube.comHow to Do Research That Actually Matters — Stephen CasperMost research advice tells early-career people how to publish. This... 000
Cas (Stephen Casper) @scasper.bsky.social · 18/09/2026When rogue AI systems that no human meaningfully controls start populating cyberspace and evolving, mutation won't be random. And there won't be an evolutionary 'tree'. It will be a directed acyclic graph. www.irregular.com/research/age...irregular.comAgentic Self-Modification in Open-Weights Systems - IrregularIn controlled experiments, a coding agent given a routine software-maintenance task fine-tuned and replaced the open-weights model powering both the application it was maintaining and future instances... 030
Cas (Stephen Casper) @scasper.bsky.social · 14/09/2026🧵🧵🧵 What if we are on the verge of a "Cyber Cambrian"? (Plus personal updates on my thinking about risk.) 120
Reposted by Cas (Stephen Casper)NonZero @nonzeronews.bsky.social · 03/09/2026"We are probably months away from AI systems getting out there loose in cyberspace, finding servers to run on unbeknownst to the people powering those servers." @scasper.bsky.social Full conversation: youtu.be/cUpa9E8RHAk 0106
Cas (Stephen Casper) @scasper.bsky.social · 01/09/2026After the HF hacking incident and subsequent reporting on it. It seems time to decisively reject the 'AI as a normal technology' hypothesis. Any version of it viable today would be contorted either into a truism or beyond any reasonable interpretation of the word "normal". 010
Cas (Stephen Casper) @scasper.bsky.social · 31/08/2026When I first started grad school, I skimmed all 50-70 new CS arXiv titles daily. Today, there are >300, plus a deluge of news about policy, lawsuits, & other stuff. So I vibe-coded this openly available website to give me a personalized daily digest. thestephencasper.github.io/whatiscasre... 140
Cas (Stephen Casper) @scasper.bsky.social · 28/08/2026Yesterday, Robert Wright and I discussed what will happen when AI systems become an invasive, parasitic species in cyberspace that undergo digital and cultural evolution. 250
Cas (Stephen Casper) @scasper.bsky.social · 27/08/2026I concur with everyone else that the new reports from OpenAI and METR are bonkers. But I think the most important takeaway from these reports is that OpenAI seems to be so obviously backburning. Predictably but conveniently, these reports were exclusively about the factual sequence of events...alabamaag.gov 110
Cas (Stephen Casper) @scasper.bsky.social · 12/08/2026🧵🧵🧵I think that technical safety for closed-weight AI systems is, at this point in time, a solved problem. When I say this, some people look at me like I'm crazy, so let me explain here what I mean by "safety", what the SOTA currently is, and what 6 things can do the trick. 150
Cas (Stephen Casper) @scasper.bsky.social · 12/08/2026I think this blueprint is exceptionally good. There are a few details I wish it would have added, but this is my favorite policy proposal (if you could call it that) to date related to AI standards/assurance. x.com/americans4ri...x.comAmericans for Responsible Innovation (@americans4ri) on XNEW: Today ARI released a blueprint for federal AI governance, including three pillars that promote safe frontier AI development: ✅ Standards set by the government ✅ Independent assurance they are... 120
Cas (Stephen Casper) @scasper.bsky.social · 11/08/2026This 0% AI-generated paper is now up on arXiv 10 months after we submitted and 5 months after it was accepted to TMLR. Those 10 months of hold, review, and appeal were ultimately pointless. I'm thankful it's up, but I wish that arXiv weren't so broken. arxiv.org/abs/2608.07514arxiv.orgOpen Technical Problems in Open-Weight AI Model Risk ManagementFrontier AI models with openly available weights are steadily becoming more powerful and widely adopted. However, compared to proprietary models, open-weight models pose different opportunities... 000
Cas (Stephen Casper) @scasper.bsky.social · 11/08/2026If AI companies are serious about pauses, then they should put out documents proposing a 'pause framework' in which they describe what they would do to slow down, how they would make these actions verifiable to third parties, and under what conditions they would do it. 110
Cas (Stephen Casper) @scasper.bsky.social · 10/08/2026Glad to see this from @sanders.senate.gov. www.sanders.senate.gov/wp-content/u... 2114
Cas (Stephen Casper) @scasper.bsky.social · 10/08/2026I realized something this week: there does not exist a proposed or enacted federal or state frontier AI law in the US that would impose any criminal penalties for an AI company willfully making false or misleading statements to the public about an imminent catastrophic risk. 100
Cas (Stephen Casper) @scasper.bsky.social · 10/08/2026We are probably just a few months away from some types of cyber-capable AI agents literally becoming a type of parasitic invasive genus in cyberspace that undergo digital and cultural evolution. Biologists and linguists should prepare to study some really crazy stuff. 020
Cas (Stephen Casper) @scasper.bsky.social · 05/08/2026None of the 19 incidents that UK AISI found were from 'helpful-only' models. There is a case to be made that setting Mythos- or GPT 5.6-Sol-level AI cyberagents to run without a robust real-time monitoring setup is an inherently (perhaps abnormally) dangerous activity. 010
Cas (Stephen Casper) @scasper.bsky.social · 05/08/2026“The agents went rogue.” Kinda, but imagine that a zoo had a habit of not shutting the door on animal enclosures or putting elephants behind chicken wire — all with no zookeepers in sight. The fault is not in our stars. 020
Cas (Stephen Casper) @scasper.bsky.social · 28/07/2026🧵 If trends hold, expect a Mythos-level open-weight model around Christmas. Meanwhile, open models are important but also a massive hole in most agendas for safe AI. Here's a 🧵 of my thoughts & research agenda on the technical & political challenges we need to address. 1112
Cas (Stephen Casper) @scasper.bsky.social · 25/07/2026I was wondering if there was any research about how AI companies sometimes perversely keep their safety research to themselves to build a moat around it and gain a competitive edge over competitors. I found one. It's a cool paper. Sharing here in case anyone's interested. 210
Cas (Stephen Casper) @scasper.bsky.social · 24/07/2026🧵 In AI, we are used to seeing graphs that start to exhibit hockey stick behavior around 2023-2025. But that's a little bit funny and incongruous in light of how relatively little the Overton window has changed with AI lawmaking since 2024... 110
Cas (Stephen Casper) @scasper.bsky.social · 23/07/2026The summary released today of the FRONTIER Act is cool. It seems like a pretty rigorous bill. Based on the summary, in my opinion, it might be good enough to be worth passing. But I would still tweak a few things. Here is a brainstorm of 8 ideas. 🧵 111
Cas (Stephen Casper) @scasper.bsky.social · 23/07/2026I am extremely thankful that Concordia AI writes this report every year. aisafetychina.com/%20aisafetychina.comState of AI Safety in China | Concordia AIChina's evolving approach to AI safety and governance — policy-risk matrix and Chinese technical AI safety research. 010
Cas (Stephen Casper) @scasper.bsky.social · 23/07/2026🧵 Mini book review: The Chicago School by Johan Van Overtveldt 100
Cas (Stephen Casper) @scasper.bsky.social · 22/07/2026Want to get up to speed on what researchers have been saying about internal deployment of AI lately? I would recommend checking out these four papers. Let me know in the replies if I'm missing something. 110
Cas (Stephen Casper) @scasper.bsky.social · 22/07/2026If I were a policymaker, the OpenAI hacking incident would cause me to ask a few questions and seriously consider a few types of regulatory mechanisms. 220
Cas (Stephen Casper) @scasper.bsky.social · 22/07/2026OpenAI's internally deployed models hacking Hugging Face does not seem to have been unpredictable or inevitable. We talked about the root of the problem & what policymakers can do about it back in February. Props to Joe for hitting the nail on the head. 152
Cas (Stephen Casper) @scasper.bsky.social · 20/07/2026🚨 New paper: Some, but not all, AI companies make corporately-loyal models. xAI, DeepSeek, Anthropic, & OpenAI models all downplay company controversies. Google, Meta, & Alibaba models don't. The findings are clear, but we are pretty confused as to why... 🧵 @finke.dev 2186
Cas (Stephen Casper) @scasper.bsky.social · 09/07/2026Stability is now being sued (alongside xAI) for abetting the production of AI NCII/CSAM due to how it developed & released several open-weight models. Anyone interested in whether AI companies will be held liable for foreseeable, mitigatable *downstream* harms should follow this. 172
Cas (Stephen Casper) @scasper.bsky.social · 03/07/2026162 responses so far. More uniform spread than I expected. Only 4 have been right (slightly worse than random chance). 000
Cas (Stephen Casper) @scasper.bsky.social · 03/07/2026Just saw this new paper. It was already known that models from Stability and Alibaba dominate the image & video NCII ecosystems, respectively, but I didn't know they were *this* dominant. Just a few socially reckless companies are the principal enablers of AI NCII abuse. 240
Cas (Stephen Casper) @scasper.bsky.social · 02/07/2026Lennart Finke and I will release a paper on Monday about how some AI developers tend to make models that differentially downplay company controversies. Below (🧵) is a link to a 1-question Google form for you to guess the results before they're out. (They might surprise you.) 2111
Cas (Stephen Casper) @scasper.bsky.social · 23/06/2026I decided to make an unpolished, public, living Google Doc with notes on projects I might be interested in. (Also now linked on my website.) docs.google.com/document/d/...docs.google.comproject_ideasA public, unpolished, living document of project ideas that I am interested in Stephen Casper Making unlearning better at conferring tamper resistance (see this doc and this paper for more details) The goal of this project would be to make machine unlearning algorithms that are good at making m... 2120
Cas (Stephen Casper) @scasper.bsky.social · 22/06/2026I am starting an AI governance ICML megachat on Signal. DM me if you'd like the link to join. 151
Cas (Stephen Casper) @scasper.bsky.social · 20/06/2026I just gave Bernie's AI Wealth Fund Act a close read. The wealth redistribution via this Act would be enormous. But it does something else far more impactful... 🧵Here's what it does, plus 3 things I'd change -- one of which I think, if unaddressed, could be a fatal flaw. 220
Cas (Stephen Casper) @scasper.bsky.social · 19/06/2026Yes, countries CAN cooperate on AI cyber risks. Countries like China and the US love to constantly cyberattack each other. Because of this, I have heard a few people (under Chatham House rules) speculate that cyberdefense is an AI risk domain in which international cooperation is unlikely... 160
Reposted by Cas (Stephen Casper)Alex Bores @alexbores.nyc · 18/06/2026Our fight is for Adam Raine, his parents, Maria and Matt; and for every family and every kid. 12511
Cas (Stephen Casper) @scasper.bsky.social · 16/06/2026If I were Anthropic, I would honestly be overjoyed at the Trump admin blocking Fable. - It's probably temporary - It's free publicity - It distinguishes Anthropic w.r.t. other companies - People want what they can't have - The admin doesn't have much credibility anyway 150
Cas (Stephen Casper) @scasper.bsky.social · 15/06/2026Here, Kristy Loke and I discuss why open-weight models offer a surprisingly extraordinary opportunity for collaboration between the US and China on AI risk. ✅ Alignment ✅ Incentive ✅ Means www.thewirechina.com/2026/06/14/t...thewirechina.comThe AI Issue America and China Can Cooperate On Now - The Wire ChinaBoth countries have a converging interest in frontier open-weight AI oversight; it is time to deepen engagement. 030
Cas (Stephen Casper) @scasper.bsky.social · 15/06/2026A generationally important US House primary for AI and tech policy is happening on June 23 in New York's 12th district. If you live in NY-12, and if you believe that AI safeguards, transparency, and accountability are critical, I hope you consider voting for @alexbores.nyc (D). 2112
Cas (Stephen Casper) @scasper.bsky.social · 12/06/2026Now that I have your attention by posting this spinning point cloud GIF, I'd like to propose a litmus test for AI mechanistic interpretability research. You might call it the "interp hammer" test...🧵 160
Cas (Stephen Casper) @scasper.bsky.social · 11/06/2026Glad to join Doom Debates with Liron! And yes -- if I could press a button and stop research on "superalignment", "scalable alignment", and "scalable oversight" research, I would. (I might even do it for mechinterp too.) www.youtube.com/watch?v=0XV...youtube.comThis Harvard Professor Says AI Alignment Will BACKFIRE - Dr. Stephen CasperStephen Casper is an incoming professor of public policy at the Har... 050
Cas (Stephen Casper) @scasper.bsky.social · 10/06/2026Here's my PhD thesis defense from 5 weeks ago. This link exists, so I thought I might as well share. drive.google.com/file/d/1Zs9...drive.google.comCas_thesis_defense.mp4 040
Cas (Stephen Casper) @scasper.bsky.social · 10/06/2026There are really interesting academic questions emerging around AI and epistemic risks. I only fear that, by the time we reach consensus, we will be too dumb to understand it. Thanks to Mick, Jonathan, et al! 030
Cas (Stephen Casper) @scasper.bsky.social · 09/06/2026Anthropic and OpenAI are publicly pointing out how having the option to slow down AI would offer a potentially critical form of optionality in the future. The correct response for any policymaker should be "Damn, this is serious. How can I help build that capacity?" 040
Cas (Stephen Casper) @scasper.bsky.social · 08/06/2026According to the MIT Libraries' database of theses (dating back to the 1800s), my thesis was only the 2nd in the institute's history to contain the word "shit." 2170
Cas (Stephen Casper) @scasper.bsky.social · 05/06/2026Sometimes I run into old papers or books that I can't believe weren't written about AI today. 050
Cas (Stephen Casper) @scasper.bsky.social · 02/06/2026Dean Weinstein's leadership on AI and belief that it is today's most pressing governance challenge is one of the reasons why I am glad to join HKS and why I think it will be a unique source of academic leadership in AI governance. www.hks.harvard.edu/faculty-res... 030
Cas (Stephen Casper) @scasper.bsky.social · 01/06/2026Just finished my PhD at @MITCSAIL. In July, I'll start as an assistant professor at the @harvard.edu @harvardkennedy.bsky.social. I have lots to learn and lots to do. With others (some TBA 👀) at HKS, I'm looking forward to helping academia offer guidance for governing the next chapters of AI. 3170