Atoosa Kasirzadeh @atoosakz.bsky.social · 14/08/2026Everyone says 2026 is the year of AI agents. But what are they, and how should we govern them? Iason Gabriel and I have answers for you in our new Nature paper. Read the full paper here: rdcu.be/fzZvR. 1/3 121
Atoosa Kasirzadeh @atoosakz.bsky.social · 07/07/2026🥁 New paper is out: “The case for globally beneficial technology”. Iason Gabriel and I develop five moral arguments in support of the claim that advanced technologies such as AGI should be designed, developed, and distributed in ways that benefit everyone. Full paper: arxiv.org/pdf/2607.03906 031
Atoosa Kasirzadeh @atoosakz.bsky.social · 22/12/2025📣 Our book "Contemporary Debates in the Ethics of Artificial Intelligence" is now available as an e-book! Big thanks to my co-editors Sven Nyholm and John Zerilli and all the wonderful contributors The print version will be published on Jan 21: lnkd.in/gHkFC3nb 010
Atoosa Kasirzadeh @atoosakz.bsky.social · 17/12/2025I’m very excited to announce the release of my first co-edited volume, "Contemporary Debates in the Ethics of Artificial Intelligence" (Wiley), alongside my wonderful co-editors @svennyholm.bsky.social and John Zerilli , on the 21 January 2026! 030
Reposted by Atoosa KasirzadehKshitish Ghate @kghate.bsky.social · 14/10/2025🚨New paper: Reward Models (RMs) are used to align LLMs, but can they be steered toward user-specific value/style preferences? With EVALUESTEER, we find even the best RMs we tested exhibit their own value/style biases, and are unable to align with a user >25% of the time. 🧵 1127
Reposted by Atoosa Kasirzadehbleary @bleary.off-the-records.com · 13/10/2025If anyone needs me I will be in the museum, lying down next to the bog bodies. 1505235984819
Reposted by Atoosa KasirzadehMax Kleiman-Weiner @maxkw.bsky.social · 02/10/2025Great work led by @andyliu.bsky.social and collaborators: @kghate.bsky.social, @monadiab77.bsky.social, @daniel-fried.bsky.social, @atoosakz.bsky.social Preprint: www.arxiv.org/abs/2509.25369arxiv.orgGenerative Value Conflicts Reveal LLM PrioritiesPast work seeks to align large language model (LLM)-based assistants with a target set of values, but such assistants are frequently forced to make tradeoffs between values when deployed. In response ... 051
Reposted by Atoosa KasirzadehAndy Liu @andyliu.bsky.social · 02/10/2025🚨New Paper: LLM developers aim to align models with values like helpfulness or harmlessness. But when these conflict, which values do models choose to support? We introduce ConflictScope, a fully-automated evaluation pipeline that reveals how models rank values under conflict. (📷 xkcd) 1164
Reposted by Atoosa KasirzadehBennett @bennettmcintosh.com · 17/09/2025If you read one review of If Anyone Builds It Everyone Dies make it @sigalsamuel.bsky.social 's about competing #AI worldviews "It was hard to seriously entertain both [doomer and AI-as-normal tech] views at the same time." www.vox.com/future-perfe...vox.comThe AI doomers are not making an argument. They’re selling a worldview.How rational is Eliezer Yudkowsky’s prophecy? 144
Reposted by Atoosa KasirzadehSigal Samuel @sigalsamuel.bsky.social · 18/09/2025I had a very trippy experience reading the new Eliezer Yudkowsky book, IF ANYONE BUILDS IT, EVERYONE DIES. I agree with the basic idea that the current speed & trajectory of AI progress is incredibly dangerous! But I don't buy his general worldview. Here's why: www.vox.com/future-perfe...vox.com“AI will kill everyone” is not an argument. It’s a worldview.How rational is Eliezer Yudkowsky’s prophecy? 172
Reposted by Atoosa KasirzadehLilian Edwards @lilianedwards.bsky.social · 23/09/2025Very happy to announce brilliant new paper on AI by the wonderful @atoosakz.bsky.social , Phillip Hacker, and er, me, on systemic risk and how it is confusingly differently interpreted through the EU AIA, the DSA and its origin, financial regulation. Paper at arxiv.org/pdf/2509.17878 1225
Reposted by Atoosa KasirzadehBennett @bennettmcintosh.com · 17/09/2025Kasirzadeh also has a really insightful, heartbreaking thread about what "normal" hides here: bsky.app/profile/atoo... 021
Reposted by Atoosa KasirzadehBennett @bennettmcintosh.com · 17/09/2025Anyway, big fan of the 3rd option Sigal presents, @atoosakz.bsky.social 's warning of "gradual accumulation of smaller, seemingly non-existential, AI risks" until catastrophe. A warning that suggests AI safetyists should take AI ethics & sociology much more seriously! arxiv.org/pdf/2401.07836arxiv.org 132
Atoosa Kasirzadeh @atoosakz.bsky.social · 11/09/2025How good are the current AI agents for scientific discovery? We have answers in our new paper: lnkd.in/dxzmZXpR! We look at 4 ways AI scientists can go wrong; design experiments to show the manifestation of these failures in 2 open source AI scientists; and recommend detection strategies. 081
Atoosa Kasirzadeh @atoosakz.bsky.social · 19/08/2025🔊📢 Call for paper is now open (iaseai.org/iaseai26)! I’m thrilled to share that I will once again be serving as Program Co-Chair for International Association for Safe & Ethical AI Conference, this time for the 2026 edition along with the Co-Chair Jaime Fernández Fisac.iaseai.orgIASEAI'26Building a Global Movement for Safe and Ethical AI 041
Reposted by Atoosa Kasirzadehryantlowe.bsky.social @ryantlowe.bsky.social · 11/07/2025Introducing: Full-Stack Alignment 🥞 A research program dedicated to co-aligning AI systems *and* institutions with what people value. It's the most ambitious project I've ever undertaken. Here's what we're doing: 🧵 1197
Atoosa Kasirzadeh @atoosakz.bsky.social · 10/07/2025At hashtag#AIforGood Summit 2025: ‘AI agents’ is mentioned almost every panel, from law and governance to the developer community; yet the term remains so opaque. A plug in to our paper with Iason Gabriel which offers one of the clearest definition I’ve seen: arxiv.org/pdf/2504.21848arxiv.org 040
Atoosa Kasirzadeh @atoosakz.bsky.social · 16/06/2025I was planning to launch my substack on "Human, life, AI, and future" in a few months, with something very different. But life doesn’t always care about our timelines. Events erupt. Emotions build. And suddenly, waiting feels like avoidance. My first post: atoosatopia.substack.com/p/valid-and-... 021
Atoosa Kasirzadeh @atoosakz.bsky.social · 15/06/2025Ever since I first heard the slogan “AI as normal technology,” I’ve felt uneasy. Tonight that unease crystallised. In this 🧵 I unpack what normal hides and why the metaphor may ultimately fail to capture AI’s abnormal impacts on human life and societies. 1/n 141
Atoosa Kasirzadeh @atoosakz.bsky.social · 23/05/2025What can policy makers learn from “AI safety for everyone” (Read here: www.nature.com/articles/s42... ; joint work with @gbalint.bsky.social )? I wrote about some policy lessons for Tech Policy Press.nature.comAI safety for everyone - Nature Machine IntelligenceA systematic review of peer-reviewed AI safety research reveals extensive work on practical and immediate concerns. The findings advocate for an inclusive approach to AI safety that embraces diverse m... 0186
Atoosa Kasirzadeh @atoosakz.bsky.social · 21/05/2025Congrats, Dr. Alex, for writing a fantastic dissertation! I had the pleasure of co-supervising Alex during my time at the U of Edinburgh 2022-2024! To have a glimpse at Alex's research read this amazing paper: link.springer.com/article/10.1... 010
Atoosa Kasirzadeh @atoosakz.bsky.social · 19/05/2025This is a really good spicy debate and we need more spicy debates of this kind please! 020
Atoosa Kasirzadeh @atoosakz.bsky.social · 09/05/2025I've learned so much from Tom's unbeatable papers about agency, in particular his technically and philosophically rich account of agency from a causal perspective: www.alignmentforum.org/posts/Qi77Tu... 020
Atoosa Kasirzadeh @atoosakz.bsky.social · 02/05/2025New paper with Iason Gabriel on "Characterizing AI agents" is out! 2025 is being called the year of AI agents, with overwhelming headlines about them every day. But we lack a shared vocabulary to distinguish their fundamental properties. Our paper aims to bridge this gap. A 🧵 2174
Reposted by Atoosa KasirzadehScenarios for Tomorrow @scenariofutures.bsky.social · 23/04/2025Check out the latest in #AI discoveries #AIrisk #ExplainableAI #architecture #FARvisions 021
Atoosa Kasirzadeh @atoosakz.bsky.social · 17/04/2025Our paper "AI safety for everyone" with Balint Gyevnar is out at Nature Machine Intelligence: www.nature.com/articles/s42... We challenge the narrative that AI safety is primarily about minimizing existential risks from AI. Why does this matter?nature.comAI safety for everyone - Nature Machine IntelligenceA systematic review of peer-reviewed AI safety research reveals extensive work on practical and immediate concerns. The findings advocate for an inclusive approach to AI safety that embraces diverse m... 3128
Reposted by Atoosa KasirzadehKnight First Amendment Institute @knightcolumbia.org · 10/04/2025Panel 1: Regulating AI in a Time of Democratic Upheaval starts in approximately 5 minutes. Panelists: @atoosakz.bsky.social, @randomwalker.bsky.social, @alondra.bsky.social, and Deirdre K. Mulligan. Moderator: @shaynelongpre.bsky.social. #AIDemocraticFreedoms 142
Reposted by Atoosa KasirzadehKnight First Amendment Institute @knightcolumbia.org · 02/04/2025On 4/10 and 4/11, we're hosting our symposium "AI and Democratic Freedoms." Excited to have panelists @atoosakz.bsky.social, @randomwalker.bsky.social, @alondra.bsky.social, & Deirdre K. Mulligan join moderator @shaynelongpre.bsky.social to kick it off. RSVP: www.eventbrite.com/e/artificial... 1125
Atoosa Kasirzadeh @atoosakz.bsky.social · 30/03/2025Out in Philosophical Studies. A 🧵: Read here: link.springer.com/article/10.1... Most AI x-risk discussions focus on a cataclysmic moment—a decisive superintelligent takeover. But what if existential risk doesn’t arrive like a bomb, but seeps in like a leak?link.springer.comTwo types of AI existential risk: decisive and accumulative - Philosophical StudiesThe conventional discourse on existential risks (x-risks) from AI typically focuses on abrupt, dire events caused by advanced AI systems, particularly those that might achieve or surpass human-level i... 161