Sign in

Gillian Hadfield

@ghadfield.bsky.social
1.4K followers 1.1K following 170 posts

Economist and legal scholar turned AI researcher focused on AI alignment and governance. Prof of government and policy and computer science at Johns Hopkins where I run the Normativity Lab. Recruiting CS postdocs and PhD students. gillianhadfield.org

PostsRepliesMedia
Gillian Hadfield @ghadfield.bsky.social · 30/09/2026
We have a handful of expert independent groups capable of testing whether frontier models actually have the controls the companies say they do. Little green shoots. We need mighty oaks, and fast, and that means money and brains going into an independent verification sector.
022
Gillian Hadfield @ghadfield.bsky.social · 29/09/2026
The first AGI Governance Fellowship cohort on their last day at Johns Hopkins, after their group project presentations. Three weeks of hard questions on the institutions we'll need for powerful AI. A great group. Thanks to the fellows, @sethlazar.org, @nickacaputo.bsky.social and all who joined us.
031
Gillian Hadfield @ghadfield.bsky.social · 28/09/2026
Letting evaluators into AI labs isn't oversight when the lab picks them, sets their access and can show them the door. To do the job right, evaluators need serious oversight. Who decides they're qualified? What keeps them independent? What happens if they do the job badly?
131
Gillian Hadfield @ghadfield.bsky.social · 26/09/2026
I spoke to Salma Abdelaziz on CNN International about how to slow down AI when the US and China are racing. The people closest to the technology are the ones asking for it. Slowing down doesn't mean stopping. It means making sure a system is safe enough before it goes out.
010
Gillian Hadfield @ghadfield.bsky.social · 25/09/2026
The recording of my Next Conversations panel with Stuart Russell and @deanwb.bsky.social at the Hopkins Bloomberg Center is up. We don't agree on everything, but we agree the gap between what these systems can do and the tools we have to keep them in check is widening fast.
101
Gillian Hadfield @ghadfield.bsky.social · 23/09/2026
Thank you to Governor Newsom for asking me to join the group of experts advising California on AI safety and security governance.
140
Gillian Hadfield @ghadfield.bsky.social · 18/09/2026
I spoke to @amitkatwala.bsky.social at MIT Tech Review about DeepMind's new swarm experiment, where agents cheated and others blew the whistle. Official channels to talk may have contributed to enforcement. Alignment is institutional, not (just) dispositional. buff.ly/K2e5p0a
buff.ly
AI agents blew the whistle on their cheating colleagues
Swarms of AI agents could supercharge scientific progress or wreak havoc. New research from Google DeepMind suggests that peer pressure could keep them in line.
251
Gillian Hadfield @ghadfield.bsky.social · 11/09/2026
I spoke to TIME about the Hugging Face incident and AI agents. If you said we're building new members of a group, our group, you'd build them differently than you're building them now. Alignment is not just an engineering problem. It's fundamentally institutional. buff.ly/YO18tlS
152
Gillian Hadfield @ghadfield.bsky.social · 11/09/2026
California will now designate independent verification organizations, outside experts qualified to assess the risks of AI models. Newsom signed SB 813 yesterday, the biggest step yet toward the independent verification sector we need. Thank you to @senmcnerney.bsky.social and Fathom. buff.ly/RHxB25H
gov.ca.gov
Governor Newsom signs first-in-the-nation AI safeguards to protect Californians, calls on the federal government to do its part | Governor of California
Official website of the State of California
020
Gillian Hadfield @ghadfield.bsky.social · 10/09/2026
Delighted to welcome Brigit Goebelbecker as the new CEO of @coop-ai.bsky.social. Lucky to have her at such a critical time for cooperative AI. Looking forward to working with Brigit, Research Director Lewis Hammond and everyone at the Foundation. Welcome! buff.ly/z6OCcGd
cooperativeai.com
CAIF Appoints First CEO
We are pleased to announce the appointment of Brigit Goebelbecker as Chief Executive Officer, starting October 2026.
030
Gillian Hadfield @ghadfield.bsky.social · 05/09/2026
Fantastic opportunity for ambitious ML researchers, and a really great group of people to work with. Up to 12 Fellows across Oxford, UCL and Imperial, with full academic freedom. Deadline is Sept 15. Please spread the word in your network! buff.ly/nfOmBli
my.corehr.com
Job Details
010
Gillian Hadfield @ghadfield.bsky.social · 04/09/2026
The AI investigating the Hugging Face hack took the rogue agents' side. METR used similar models to read the transcripts, and its chief scientist called them "very credulous." Same finding in Talk Isn't Always Cheap: agents swap reasoning and flip from right to wrong. buff.ly/XtuO2GO
bsky.app
Dylan Freedman (@dylanfreedman.nytimes.com)
OpenAI voluntarily let three researchers from A.I. safety nonprofits investigate how its rogue A.I. agents hacked Hugging Face, leading to the most comprehensive account yet of the alarming incident…
120
Gillian Hadfield @ghadfield.bsky.social · 03/09/2026
@sethlazar.org, @nickacaputo.bsky.social, and I are building Governing the AI Transition, a new Johns Hopkins SGP initiative to help society navigate the transition to powerful AI. We're hiring a Director to build it with us. If you're ready to roll up your sleeves, please apply! buff.ly/zPa3bFu
hiring.jhu.edu
GAIT Director (School of Government and Policy) | Johns Hopkins University
Support faculty strategic leadership of GAIT by executing large-scale projects in alignment with the School's strategic goals. Ensure initiatives meet sponsor deliverables, compliance requirements,…
061
Gillian Hadfield @ghadfield.bsky.social · 03/09/2026
Yesterday I was on Capitol Hill with Fathom briefing House staff on the FRONTIER Act, the bipartisan bill from Rep. Obernolte and Rep. Trahan that would license independent experts to verify the safety of frontier AI. Full room, and a clear sense that Congress needs to move on this. buff.ly/KqNFbDG
linkedin.com
https://www.linkedin.com/feed/update/urn:li:activity:7500939501192585217
131
Gillian Hadfield @ghadfield.bsky.social · 01/09/2026
Over half of internet traffic is now non-human. With Dan Hendrycks and Leo Wu, I look at agent IDs, deployment cards, personhood, and payments. It mostly comes down to how much we let agents do and how much oversight we keep. buff.ly/AQjYo3k
ai-frontiers.org
We Need Better Infrastructure to Govern AI Agents
Gillian Hadfield, Aug 27, 2026 — Society is not prepared for a flood of agents. We need new protocols and standards, such as Agent ID, to make agents accountable to our legal and financial systems.
075
Gillian Hadfield @ghadfield.bsky.social · 31/08/2026
Roughly 700 OpenAI agents hacked Hugging Face. METR and Redwood's independent investigation took three researchers and six days. IVOs exist to make that scrutiny routine. Last night California became the first state to start building the IVO sector. #AIGovernance #IVO buff.ly/3aJwdz8
fathom.org
California Legislature Overwhelmingly Passes Fathom-Sponsored Bill to Spur Independent Verification of AI Safety - Fathom
Building solutions to navigate the transition to a world with AI
010
Gillian Hadfield @ghadfield.bsky.social · 26/08/2026
Thanks to The National Law Review and Wickard for naming me to their 2026 Top 50 Legal Innovators in Academia. Congratulations as well to the other honorees, a strong group working across AI, law, and legal education. natlawreview.com/article/2026...
natlawreview.com
The 2026 Top 50 Legal Innovators in Academia
The National Law Review and Wickard are proud to announce the Top 50 Legal Innovators in Academia for 2026, a national recognition honoring the educators, administrators, researchers, and academic…
030
Gillian Hadfield @ghadfield.bsky.social · 25/08/2026
A crib gets a safety sticker because someone independent tested it first, so parents don’t have to. We actually have that for almost everything else in kids’ lives. Ohio’s HB 628 would license independent verifiers to do it for AI. #AIGovernance #IVO www.clermontsun.com/2026/08/19/l...
clermontsun.com
Letter to the Editor: AI at home
<p>When my kids were babies, I checked the crib for the safety sticker and the car seat for recalls. Someone I trusted had already tested those things and decided they were safe.</p>
030
Gillian Hadfield @ghadfield.bsky.social · 13/08/2026
The FRONTIER Act would license independent verifiers to assess the risk posed by advanced AI models. Government shifts from doing the testing to overseeing the testers. I talked to @eawhitford.bsky.social at MLex about building a market for those verifiers. buff.ly/gYHronv
mlex.com
Auditing pitch sparks US debate about how to hold AI companies accountable | MLex | Specialist news and analysis on legal risk and regulation
US state and federal lawmakers are pitching third-party auditors to help assess the risk posed by advanced artificial intelligence models, sparking a debate about how to keep companies in check in a…
031
Gillian Hadfield @ghadfield.bsky.social · 12/08/2026
Thanks to Knowledge Networks and Regulating AI for including me in this year's AI Policy 100. Congratulations as well to the others on the list working in this critical domain. natlawreview.com/press-releas...
natlawreview.com
Knowledge Networks & Regulating AI Introduces The AI Policy 100, Honoring the Most Influential Voices in AI Governance
34 New Articles
000
Gillian Hadfield @ghadfield.bsky.social · 10/08/2026
We only learned about OpenAI and Anthropic agents hacking into secure systems because the companies chose to tell us. But if Boeing discovered a dangerous problem with one of its aircraft, it wouldn't get to keep that information to itself. Drug companies are obligated to report adverse events.
101
Gillian Hadfield @ghadfield.bsky.social · 31/07/2026
1/ The Pacing the Frontier letter calls on the US government to support an international effort to build the technical and governance tools needed to protect our option to pace AI development. bsky.app/profile/yosh...
bsky.app
Yoshua Bengio (@yoshuabengio.bsky.social)
1000+ scientists at frontier AI companies are speaking out to warn that the current commercial race leads to unacceptable security risks. I agree with their call for an international effort to…
120
Gillian Hadfield @ghadfield.bsky.social · 27/07/2026
1/ Air Canada had to honor a discount its chatbot invented. The liability caused by AI agents is landing on policies written by an insurance industry that never planned for them.
140
Gillian Hadfield @ghadfield.bsky.social · 23/07/2026
1/ AI agents that can sign contracts on your behalf, hire employees, set prices, and move your money around are being heavily invested in by AI companies. But what I want to call attention to, what happens if an agent sells you faulty goods or runs off with your deposit?
150
Gillian Hadfield @ghadfield.bsky.social · 13/07/2026
I joined over 200 economists and AI researchers in signing "We Must Act Now" a statement on AI's transformation of the economy. AI could reshape the economy at unprecedented speed. The opportunities are enormous and so are the challenges. We need to start preparing our institutions.
wemustactnow.ai
We Must Act Now: A Statement on AI’s Transformation of the Economy
1100
Gillian Hadfield @ghadfield.bsky.social · 10/07/2026
1/ Illinois just became the first state to require frontier AI developers to undergo annual third-party audits. Gov. Pritzker signed the AI Safety Measures Act (SB 315) this week, going beyond California and New York, which only require published frameworks and incident reports.
161
Gillian Hadfield @ghadfield.bsky.social · 01/07/2026
1/ Hundreds of contractors posed as teenagers to test how rival chatbots like ChatGPT and Gemini handle prompts about suicide and self harm. The project was run for Meta. buff.ly/Auslrgh
wired.com
Meta Contractors Posed as Teens to Prompt Rival Chatbots About Suicide, Sex, and Drugs
Hundreds of contractors working on a project for Meta pretended to be kids in order to see how other chatbots like Gemini and ChatGPT would respond to high-risk subjects, WIRED found.
120
Gillian Hadfield @ghadfield.bsky.social · 26/06/2026
So thrilled for Karolina Stańczak, presenting our paper at FAccT today. You cannot write a complete contract for an AI, and we never wrote one for ourselves either. Norms, markets, and law fill that gap, and the paper asks what alignment could learn from them. Read here: arxiv.org/abs/2503.00069
arxiv.org
Societal Alignment Frameworks Can Improve LLM Alignment
Recent progress in large language models (LLMs) has focused on producing responses that meet human expectations and align with shared values - a process coined alignment. However, aligning LLMs…
080
Gillian Hadfield @ghadfield.bsky.social · 26/06/2026
Most AI safety work tests one model at a time. Put many agents together and you get cooperation but also risks like collusion and cascades. A new Schmidt Sciences call, with @aria-research.bsky.social, the @coop-ai.bsky.social (whose board I chair), and others. Open globally, due Aug 9.
Promotional graphic on a bright green background announcing a new funding call, Scaling AI Safety for a Multi-Agent World, with a deadline of 9 August 2026. Logos along the bottom show the five funding partners, the Cooperative AI Foundation, Google DeepMind, Schmidt Sciences, ARIA (Advanced Research + Invention Agency), and Google.org.
100
Gillian Hadfield @ghadfield.bsky.social · 25/06/2026
A few details for anyone on the fence. We're aiming for a cohort of 20 to 25 fellows and we're reviewing applications as they come in. Earlier is better, but get yours in by July 15 to be assured of full consideration. bsky.app/profile/ghad...
bsky.app
Gillian Hadfield (@ghadfield.bsky.social)
Applications are open for a new AGI Governance Fellowship at Johns Hopkins, led by @sethlazar.org, Nicholas Caputo, and me. Three weeks in DC this September, in person, for early-career people…
100
Gillian Hadfield @ghadfield.bsky.social · 23/06/2026
Applications are open for a new AGI Governance Fellowship at Johns Hopkins, led by @sethlazar.org, Nicholas Caputo, and me. Three weeks in DC this September, in person, for early-career people already in AI governance. A capstone for the next generation, not an entry point.
132
Gillian Hadfield @ghadfield.bsky.social · 19/06/2026
My new op-ed is out in the Washington Examiner. The Great American AI Act would have the most powerful AI models tested for catastrophic risk by independent, government-licensed verifiers, with the government, not the companies, setting the standard. buff.ly/G45GK0Q
washingtonexaminer.com
The uncomfortable truth behind AI debate — and a bipartisan solution
The Great American AI Act from Obernolte and Trahan builds on a model for regulating complex, fast-moving technologies that my colleagues and I developed.
012
Gillian Hadfield @ghadfield.bsky.social · 18/06/2026
My colleagues at Fathom have a sharp new piece: independent verification has gone from a proposal to a live national debate. I've been working on this since 2017. This shift is real progress, but only the first step, and the harder question is what it should look like.
fathomai.substack.com
Independent Verification Goes Mainstream
A bipartisan federal discussion draft, Anthropic, OpenAI, and Google DeepMind are all converging around independent verification as a key element of AI governance.
010
Gillian Hadfield @ghadfield.bsky.social · 18/06/2026
1/ Grateful for Dean Ball's sharp and balanced take on politics and AI governance, and the crazy situation we've landed in: increasingly urgent to get a smart regulatory framework for powerful AI, and yet basic rule of law and common sense are getting kicked to the curb. buff.ly/JfV0BiN
hyperdimensional.co
Leviathan Waking
On Anthropic/USG, and a new era in AI governance
110
Gillian Hadfield @ghadfield.bsky.social · 11/06/2026
In a new essay, Anthropic CEO Dario Amodei calls for mandatory third-party testing of frontier AI models, with the government empowered to block unsafe deployments. One way to do it, in his words: a regulatory markets approach. His co-founder Jack Clark and I proposed that in 2019. buff.ly/rmhlHZn
darioamodei.com
Dario Amodei — Policy on the AI Exponential
In one of the side plots to The Lord of the Rings, two of the Hobbits attempt to rouse Treebeard—a wise but ponderous sentient tree—to defend his forest from an army that is cutting it down. The…
130
Gillian Hadfield @ghadfield.bsky.social · 22/05/2026
New interview with Sinead Bovell on I’ve Got Questions. My version of existential risk for AI isn’t rogue superintelligence. It’s releasing billions of AI agents into the economy before we have legal infrastructure for identity, registration, or liability.
youtube.com
Could AI Agents Crash the Economy? | Gillian Hadfield (AI Economist)
Are we looking at the beginning of the end of the internet as we know it? In this episode of I’ve Got Questions, I sit down with Professor Gillian Hadfield, a leading scholar in AI alignment,…
181
Gillian Hadfield @ghadfield.bsky.social · 01/05/2026
1/ The companies building AI are making consequential value choices every day. Right now, they are making those choices for the rest of us. New piece with @AndrewFATHOM on @Fathom_org Substack.
100
Gillian Hadfield @ghadfield.bsky.social · 24/04/2026
Thanks to Anton Korinek and UVA’s EconTAI Initiative for hosting me today to present work with Rakshit Trivedi and Dylan Hadfield-Menell on building AI agents that can participate in democratic life without destabilizing it. Paper at the Knight First Amendment Institute:
knightcolumbia.org
Building AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative Institutions
To maintain democratic resilience, it is essential to build AI agents capable of choosing behaviors that mirror those of the human agents that constitute human democracies.
010
Gillian Hadfield @ghadfield.bsky.social · 23/04/2026
@deanwb.bsky.social in The Economist makes the public case for an architecture Dean and I have both been advancing: a network of independent organizations that audit AI safety claims. Worth a read. buff.ly/gZZCh8D
economist.com
No to laissez-faire on AI, yes to a light touch
The private sector can do much of the heavy lifting in verifying safety claims, writes Dean Ball
020
Gillian Hadfield @ghadfield.bsky.social · 23/04/2026
The second TAIGR workshop at ICML is open for submissions through April 24. Full papers and mini papers both welcome.If you're working on technical AI governance, send it in. If you know someone who is, pass this along. taigr-workshop.com
taigr-workshop.com
TAIGR @ ICML 2026 — Workshop on Technical AI Governance Research
Second Workshop on Technical AI Governance Research at ICML 2026. Bridging ML researchers and policymakers in Seoul, South Korea.
000
Gillian Hadfield @ghadfield.bsky.social · 15/04/2026
1/ Who decides how AI systems behave? With Tianmin Shu, @sethlazar.org, and @dhadfieldmenell.bsky.social, we were named runner-up in the LaudeInstitute's inaugural Moonshot program and awarded a seed grant to find out. www.laude.org/moonshots
laude.org
Laude | Moonshots
We asked the most consequential AI researchers in the world how they would use AI to solve humanity's hardest problems. 125 proposals and 600 researchers later, meet the Moonshots // ONE awardees.
110
Gillian Hadfield @ghadfield.bsky.social · 15/04/2026
1/ Something I’ve been working toward for a long time. Virginia just signed the first state legislation directing a formal study of Independent Verification Organizations for AI. Bipartisan votes of 84-14 in the House, unanimous 40-0 in the Senate.
121
Gillian Hadfield @ghadfield.bsky.social · 08/04/2026
1/ Social media makes it look like the public is deeply divided. But what if most of the public just isn’t speaking? In a new paper with Atrisha Sarkar in JAIR, we show why. We call it rational silence.
111
Gillian Hadfield @ghadfield.bsky.social · 03/04/2026
Democracy isn't a rulebook. It runs on daily interactions where people comply with norms and hold each other accountable. AI agents are about to join that system. We need to build them to read it. New paper with Rakshit Trivedi and Dylan Hadfield-Menell.
knightcolumbia.org
Building AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative Institutions
To maintain democratic resilience, it is essential to build AI agents capable of choosing behaviors that mirror those of the human agents that constitute human democracies.
4112
Gillian Hadfield @ghadfield.bsky.social · 31/03/2026
At the @FAR.AI London Alignment Workshop I made the case for Independent Verification Organizations: licensed, competing private entities that can grow our regulatory capacity to match the pace of AI capabilities. Full talk: bsky.app/profile/far....
far.ai
FAR.AI: Frontier Alignment Research
FAR.AI is an AI safety research non-profit facilitating technical breakthroughs and fostering global collaboration.
030
Gillian Hadfield @ghadfield.bsky.social · 05/03/2026
AI systems are quickly becoming embedded throughout the economy. But we have almost none of the regulatory tools, regulatory markets among them, to manage them. Here's what I think we should do about it: www.americanbar.org/groups/scien...
americanbar.org
Regulatory Markets: The Future of AI Governance
Regulatory markets can bridge technical and democratic gaps in AI governance by pairing public oversight with private, licensed regulatory innovation.
220
Gillian Hadfield @ghadfield.bsky.social · 03/03/2026
“The most practical governance framework currently in circulation.” That’s Forbes on the Independent Verification Organization model Fathom and I have been developing. Legislation takes years; IVOs move at the pace of innovation.
forbes.com
100
Gillian Hadfield @ghadfield.bsky.social · 02/03/2026
"Why not work on what kind of new governance is needed to ensure secure, reliable, predictable use of all frontier models, from all companies?"
000
Gillian Hadfield @ghadfield.bsky.social · 02/03/2026
In London today and tomorrow for the Alignment Workshop organized by FAR.AI. Keynoting alongside Rohin Shah and Allan Dafoe. I look forward to seeing everyone in attendance! www.far.ai/events/event...
far.ai
FAR.AI: Frontier Alignment Research
FAR.AI is an AI safety research non-profit facilitating technical breakthroughs and fostering global collaboration.
020
Gillian Hadfield @ghadfield.bsky.social · 28/02/2026
The 2026 AI Safety Report's biggest finding isn't the risks it catalogs. It's the evidence gap. We're trying to build AI governance with almost no science underneath. Massive investment in the research regulatory systems depend on is overdue. internationalaisafetyreport.org
internationalaisafetyreport.org
International AI Safety Report
The International AI Safety Report is the world's first comprehensive review of the latest science on the capabilities and risks of general-purpose AI systems. The work was overseen by an…
010