Sign in

Tom Everitt

@tom4everitt.bsky.social
1.2K followers 358 following 102 posts

AGI safety researcher at Google DeepMind, leading causalincentives.com Personal website: tomeveritt.se

PostsRepliesMedia
Tom Everitt @tom4everitt.bsky.social · 10/02/2026
Thoughtful essay on power concentration from AI freesystems.substack.com/p/the-enligh...
freesystems.substack.com
The Enlightened Absolutists
In 2017, OpenAI's founders warned about creating an 'AGI dictatorship.' Nine years later, we still haven't built the structures to prevent one.
020
Tom Everitt @tom4everitt.bsky.social · 21/11/2025
Keeping chains-of-thought traces reflective of the models true reasoning would be very helpful for safety. Important work to explore the ways it may fail
010
Reposted by Tom Everitt
Alex Turner @turntrout.bsky.social · 04/11/2025
New Google DeepMind paper: "Consistency Training Helps Stop Sycophancy and Jailbreaks" by @alexirpan.bsky.social, me, Mark Kurzeja, David Elson, and Rohin Shah. (thread)
The abstract of the consistency training paper.
1185
Reposted by Tom Everitt
Joel Z Leibo @jzleibo.bsky.social · 31/10/2025
[1/9] Excited to share our new paper "A Pragmatic View of AI Personhood" published today. We feel this topic is timely, and rapidly growing in importance as AI becomes agentic, as AI agents integrate further into the economy, and as more and more users encounter AI.
35415
Reposted by Tom Everitt
Toby Ord @tobyord.bsky.social · 25/09/2025
Evaluating the Infinite 🧵 My latest paper tries to solve a longstanding problem afflicting fields such as decision theory, economics, and ethics — the problem of infinities. Let me explain a bit about what causes the problem and how my solution avoids it. 1/N arxiv.org/abs/2509.19389
arxiv.org
Evaluating the Infinite
I present a novel mathematical technique for dealing with the infinities arising from divergent sums and integrals. It assigns them fine-grained infinite values from the set of hyperreal numbers in a ...
2125
Reposted by Tom Everitt
Edward Grefenstette @egrefen.bsky.social · 21/07/2025
Do you have a PhD (or equivalent) or will have one in the coming months (i.e. 2-3 months away from graduating)? Do you want to help build open-ended agents that help humans do humans things better, rather than replace them? We're hiring 1-2 Research Scientists! Check the 🧵👇
3196
Reposted by Tom Everitt
Yoshua Bengio @yoshuabengio.bsky.social · 10/07/2025
digital-strategy.ec.europa.eu/en/policies/... The Code also has two other, separate Chapters (Copyright, Transparency). The Chapter I co-chaired (Safety & Security) is a compliance tool for the small number of frontier AI companies to whom the “Systemic Risk” obligations of the AI Act apply. 2/3
digital-strategy.ec.europa.eu
The General-Purpose AI Code of Practice
The Code of Practice helps industry comply with the AI Act legal obligations on safety, transparency and copyright of general-purpose AI models.
161
Reposted by Tom Everitt
vkrakovna.bsky.social @vkrakovna.bsky.social · 08/07/2025
As models advance, a key AI safety concern is deceptive alignment / "scheming" – where AI might covertly pursue unintended goals. Our paper "Evaluating Frontier Models for Stealth and Situational Awareness" assesses whether current models can scheme. arxiv.org/abs/2505.01420
161
Reposted by Tom Everitt
Csaba Szepesvari @skiandsolve.bsky.social · 08/07/2025
First position paper I ever wrote. "Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence" arxiv.org/abs/2506.23908 Background: I'd like LLMs to help me do math, but statistical learning seems inadequate to make this happen. What do you all think?
arxiv.org
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
Sound deductive reasoning -- the ability to derive new knowledge from existing facts and rules -- is an indisputably desirable aspect of general intelligence. Despite the major advances of AI systems ...
3509
Reposted by Tom Everitt
David Lindner @davidlindner.bsky.social · 04/07/2025
Can frontier models hide secret information and reasoning in their outputs? We find early signs of steganographic capabilities in current frontier models, including Claude, GPT, and Gemini. 🧵
161
Tom Everitt @tom4everitt.bsky.social · 07/06/2025
Thought provoking
061
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
Are world models necessary to achieve human-level agents, or is there a model-free short-cut? Our new #ICML2025 paper tackles this question from first principles, and finds a surprising answer, agents _are_ world models… 🧵 arxiv.org/abs/2506.01622
24115
Tom Everitt @tom4everitt.bsky.social · 03/06/2025
Great to see serious work on non-agentic AI. I think it's an underappreciated direction: better for safety, society, and human meaning. LLMs show it's perfectly possible
072
Reposted by Tom Everitt
Yoshua Bengio @yoshuabengio.bsky.social · 20/05/2025
When I realized how dangerous the current agency-driven AI trajectory could be for future generations, I knew I had to do all I could to make AI safer. I recently shared this personal experience, and outlined the scientific solution I envision @TEDTalks⤵️ www.ted.com/talks/yoshua...
ted.com
The catastrophic risks of AI — and a safer path
Yoshua Bengio — the world's most-cited computer scientist and a "godfather" of artificial intelligence — is deadly concerned about the current trajectory of the technology. As AI models race toward fu...
24812
Tom Everitt @tom4everitt.bsky.social · 13/05/2025
Nice video about one of our recent papers, and some of its potential implications for AI agents www.youtube.com/watch?app=de...
youtube.com
Why Don't AI Agents Work?
YouTube video by Mutual Information
050
Tom Everitt @tom4everitt.bsky.social · 07/05/2025
Agency comes in degrees, and can vary along several dimensions: autonomy, efficacy, goal-complexity, and generality. Great paper helping us understand the different possibilities
150
Tom Everitt @tom4everitt.bsky.social · 04/05/2025
METR task-time scaling critique: "the 4-minute mark for GPT-4 is completely arbitrary; you could probably put together one reasonable collection of word counting ... tasks with average human time of 30 seconds and another ... of 20 minutes where GPT-4 would hit 50% accuracy on each"
020
Reposted by Tom Everitt
Piotr Mirowski @piotrmirowski.bsky.social · 01/05/2025
Generative AI tools used in art production should be evaluated by the broader art world of artists, art historians and curators, to integrate culturally-specific critique and to re-imagine these tools to suit the artists’ needs. dl.acm.org/doi/full/10....
dl.acm.org
AI and Non-Western Art Worlds: Reimagining Critical AI Futures through Artistic Inquiry and Situated Dialogue | Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems
271
Tom Everitt @tom4everitt.bsky.social · 30/04/2025
interesting argument about manufacturing: "[with] good enough software and hardware to create [good] humanoid robots ..., we will also be able to create more task-specific hardware ... that can do those roles cheaper, faster, and better." perhaps the same is true for AI agents more generally?
blog.spec.tech
Humanoid Robots in Manufacturing
Or, there's a reason we don't pull cars with mechanical horses
250
Reposted by Tom Everitt
neelnanda.bsky.social @neelnanda.bsky.social · 29/04/2025
I'm very impressed with the Sentinel newsletter: by far the best aggregator of global news I've found Expert forecasters filter for the events that actually matter (not just noise), and forecast how likely this is to affect eg war, pandemics, frontier AI etc Highly recommended!
162
Reposted by Tom Everitt
Yoshua Bengio @yoshuabengio.bsky.social · 23/04/2025
Rival nations or companies sometimes choose to cooperate because some areas are protected zones of mutual interest—reducing shared risks without giving competitors an edge. Our paper in FAccT '25: How geopolitical rivals can cooperate on AI safety research. arxiv.org/abs/2504.12914
arxiv.org
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
International cooperation is common in AI research, including between geopolitical rivals. While many experts advocate for greater international cooperation on AI safety to address shared global risks...
0251
Reposted by Tom Everitt
Ahmad Beirami @abeirami.bsky.social · 23/04/2025
Excited that our paper "safety alignment should be made more than just a few tokens deep" was recognized as an #ICLR2025 Outstanding Paper! We identified a common root cause to many safety vulnerabilities and pointed out some paths forward to address it!
2323
Tom Everitt @tom4everitt.bsky.social · 17/04/2025
What if LLMs are sometimes capable of doing a task but don't try hard enough to do it? In a new paper, we use subtasks to assess capabilities. Perhaps surprisingly, LLMs often fail to fully employ their capabilities, i.e. they are not fully *goal-directed* 🧵 arxiv.org/abs/2504.118...
1111
Reposted by Tom Everitt
Geoffrey Irving @girving.bsky.social · 14/04/2025
The 1st blackbox AI control paper uses a mixture of ML monitoring and editing to detect or block harm from malicious agents. However, blackbox control is one point in a space that varies with capability. Our new paper tracks how control might change along this trajectory. 🧵 arxiv.org/abs/2504.05259
arxiv.org
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from causing harm. AI develop...
171
Tom Everitt @tom4everitt.bsky.social · 12/04/2025
Causality can help with AI safety in so many ways!
lesswrong.com
Theories of Impact for Causality in AI Safety — LessWrong
Thanks to Jonathan Richens and Tom Everitt for discussions about this post. …
080
Reposted by Tom Everitt
David Lindner @davidlindner.bsky.social · 04/04/2025
Super excited this giant paper outlining our technical approach to AGI safety and security is finally out! No time to read 145 pages? Check out the 10 page extended abstract at the beginning of the paper
163
Reposted by Tom Everitt
Somnath Basu Roy Chowdhury @somnathbrc.bsky.social · 02/04/2025
𝐇𝐨𝐰 𝐜𝐚𝐧 𝐰𝐞 𝐩𝐞𝐫𝐟𝐞𝐜𝐭𝐥𝐲 𝐞𝐫𝐚𝐬𝐞 𝐜𝐨𝐧𝐜𝐞𝐩𝐭𝐬 𝐟𝐫𝐨𝐦 𝐋𝐋𝐌𝐬? Our method, Perfect Erasure Functions (PEF), erases concepts perfectly from LLM representations. We analytically derive PEF w/o parameter estimation. PEFs achieve pareto optimal erasure-utility tradeoff backed w/ theoretical guarantees. #AISTATS2025 🧵
2378
Reposted by Tom Everitt
Nenad Tomasev @nenadtomasev.bsky.social · 01/04/2025
Our work on concept discovery towards bridging the human-AI knowledge gap in AlphaZero has now been published in PNAS. As future AI systems become even more capable, we should be thinking of ways of utilizing them not only to perform tasks, but also to further our own knowledge and understanding.
095
Reposted by Tom Everitt
Alex Irpan @alexirpan.bsky.social · 01/04/2025
www.alexirpan.com/2025/04/01/w...
alexirpan.com
Who is AI For?
Who is AI for right now? There are obvious use cases. Image generation for people who want filler art for work presentations, or just to mess around. Coding assistance for people who code, vibe coding...
062
Reposted by Tom Everitt
Mor Geva @megamor2.bsky.social · 31/03/2025
🎉 Our Actionable Interpretability workshop has been accepted to #ICML2025! 🎉 > Follow @actinterp.bsky.social > Website actionable-interpretability.github.io @talhaklay.bsky.social @anja.re @mariusmosbach.bsky.social @sarah-nlp.bsky.social @iftenney.bsky.social Paper submission deadline: May 9th!
34116
Reposted by Tom Everitt
Center for AGI Investigations @cagii.bsky.social · 26/03/2025
Google launched Gemini Pro Experimental 2.5, a “thinking model,” none of today’s generative AI models scored higher than 4% against a human average of 60% on the new ARC-AGI-2 test. And Jack Clark is becoming increasingly convinced that AGI is right around the corner. 🤔 www.cagii.org/blog/google-...
cagii.org
Google, OpenAI, Artificial Intelligence benchmarks, and AGI — Center for AGI Investigations
Google launched Gemini Pro Experimental 2.5, a “thinking model,” none of today’s generative AI models have scored higher than 4% against a human average of 60% on the new ARC-AGI-2 test. And Jack Clar...
031
Reposted by Tom Everitt
Yoshua Bengio @yoshuabengio.bsky.social · 27/03/2025
This metric from METR shows the length of tasks AI agents can complete has been consistently exponentially increasing over the past 6 years, with a doubling time of around 7 months. metr.org/blog/2025-03... 1/2
metr.org
Measuring AI Ability to Complete Long Tasks
We propose measuring AI performance in terms of the *length* of tasks AI agents can complete. We show that this metric has been consistently exponentially increasing over the past 6 years, with a doub...
2365
Reposted by Tom Everitt
Risto Uuk @ristouuk.bsky.social · 27/03/2025
I have excellent news: our KU Leuven 𝗜𝗻𝘁𝗲𝗿𝗻𝗮𝘁𝗶𝗼𝗻𝗮𝗹 𝗖𝗼𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 𝗼𝗻 𝗟𝗮𝗿𝗴𝗲-𝗦𝗰𝗮𝗹𝗲 𝗔𝗜 𝗥𝗶𝘀𝗸𝘀 is now open for registration for the broader public! We received a lot of interesting talk submissions and we've made our selection of the best ones. It will take place on 𝟮𝟲-𝟮𝟴 𝗠𝗮𝘆 𝗶𝗻 𝗟𝗲𝘂𝘃𝗲𝗻, 𝗕𝗲𝗹𝗴𝗶𝘂𝗺.
kuleuven.be
International Conference on Large-Scale AI Risks
1103
Reposted by Tom Everitt
Jacob Eisenstein @jacobeisenstein.bsky.social · 24/03/2025
We all want LLMs to collaborate with humans to help them achieve their goals. But LLMs are not trained to collaborate, they are trained to imitate. Can we teach LM agents to help humans by first making them help each other? arxiv.org/abs/2503.14481
arxiv.org
Don't lie to your friends: Learning what you know from collaborative self-play
To be helpful assistants, AI agents must be aware of their own capabilities and limitations. This includes knowing when to answer from parametric knowledge versus using tools, when to trust tool outpu...
15620
Reposted by Tom Everitt
Surya Ganguli @suryaganguli.bsky.social · 24/03/2025
This might be the most useful thing I have come across in social media - a personalized feed of academic papers filtered by your follower network! Highly recommend. #academicsky
23211
Tom Everitt @tom4everitt.bsky.social · 23/03/2025
Great work by Jack Foxabbott, Rohin Subramani, and my long-term collaborator @f-rhys-ward.bsky.social We always thought causality was useful for safe and ethical AI www.alignmentforum.org/s/pcdHisDEGL... This work extends casual models with subjective beliefs, incl. beliefs about others beliefs
alignmentforum.org
010
Reposted by Tom Everitt
Murray Shanahan @mpshanahan.bsky.social · 21/03/2025
My latest paper is just out on arXiv: : "Palatable Conceptions of Disembodied Being: Terra Incognita in the Space of Possible Minds". 🧵1/n
Paper header: title, author, etc
1155
Reposted by Tom Everitt
Dr Francis Rhys Ward @f-rhys-ward.bsky.social · 16/03/2025
In real-life, agents with different subjective beliefs interact in a shared objective reality. They have higher-order beliefs about each other's beliefs and goals, which is required for phenomena involving theory-of-mind, like deception Our paper formalises this in causal models
142
Reposted by Tom Everitt
Rodney Brooks @rodneyabrooks.bsky.social · 14/03/2025
Top notch. "discourse about large models as intelligent agents is fundamentally misconceived. ... Large Models should not be viewed primarily as intelligent agents, but as a new kind of cultural and social technology, allowing humans to take advantage of information other humans have accumulated."
06417
Reposted by Tom Everitt
Margaret Mitchell @mmitchell.bsky.social · 10/03/2025
👩‍💻 What is AGI...and should it be a goal? 😱 My podcast interview on this question is out! With @eryk.bsky.social and @borhane.bsky.social , interviewed by @justinhendrix.bsky.social for @techpolicypress.bsky.social : www.techpolicy.press/should-agi-r... Based on this paper: arxiv.org/abs/2502.03689
techpolicy.press
Should AGI Really Be the Goal of Artificial Intelligence Research? | TechPolicy.Press
A conversation with Eryk Salvaggio, Borhane Blili-Hamelin, and Margaret Mitchell about the efficacy of aiming for artificial general intelligence.
13511
Reposted by Tom Everitt
Geoffrey Irving @girving.bsky.social · 05/03/2025
AISI has a new grant program for funding academic and nonprofit-affiliated research in (1) safeguards to mitigate misuse risk and (2) AI control and alignment to mitigate loss of control risk. Please apply! 🧵 www.aisi.gov.uk/grants
aisi.gov.uk
Grants | The AI Security Institute (AISI)
View AISI grants. The AI Security Institute is a directorate of the Department of Science, Innovation, and Technology that facilitates rigorous research to enable advanced AI governance.
195
Reposted by Tom Everitt
London AI and Humanity Project @laihp.bsky.social · 04/03/2025
Next up in our ongoing AI Affect series: 👤 Tom Everitt (Google Deepmind) @tom4everitt.bsky.social 📢 "Agency as backwards causality" 🗓️ TODAY Tuesday, March 4, 3-4:30 PM 📍 Join us in person or online! DM for details.
142
Reposted by Tom Everitt
vkrakovna.bsky.social @vkrakovna.bsky.social · 14/02/2025
We are excited to release a short course on AGI safety! The course offers a concise and accessible introduction to AI alignment problems and our technical / governance approaches, consisting of short recorded talks and exercises (75 minutes total). deepmindsafetyresearch.medium.com/1072adb7912c
deepmindsafetyresearch.medium.com
Introducing our short course on AGI safety
We are excited to release a short course on AGI safety for students, researchers and professionals interested in this topic. The course…
0185
Reposted by Tom Everitt
International Association for Safe & Ethical AI @iaseai.bsky.social · 12/02/2025
Can we trust AI? “Everyone in the world is getting onto this brand new airplane that really has never been tested before. And it’s going to take off and it’s never going to land. It has to fly, forever.” – Stuart Russell, IASEAI President.
2134
Reposted by Tom Everitt
Dan Goldstein @dggoldst.bsky.social · 27/01/2025
and you thought your arxiv paper had impact
35810
Tom Everitt @tom4everitt.bsky.social · 23/01/2025
Process based supervision done right, and with pretty CIDs to illustrate :)
081
Tom Everitt @tom4everitt.bsky.social · 27/12/2024
Can AI boost human performance at training (even) better AI systems? This is an essential recursion to get right! The Rater Assist team at Google DeepMind describes their approach
060
Tom Everitt @tom4everitt.bsky.social · 14/12/2024
🔦 @NeurIPS2024 spotlight paper we’re presenting today. Making AI systems more agentic is a hot research topic. But powerful agents bring worries about misalignment and loss of control. Can we measure how agentic an AI system is? 🧵
1151
Tom Everitt @tom4everitt.bsky.social · 27/11/2024
Excellent blue sky features summary
020