Sign in

Tom Everitt

@tom4everitt.bsky.social
1.2K followers 358 following 102 posts

AGI safety researcher at Google DeepMind, leading causalincentives.com Personal website: tomeveritt.se

PostsRepliesMedia
Tom Everitt @tom4everitt.bsky.social · 10/02/2026
Thoughtful essay on power concentration from AI freesystems.substack.com/p/the-enligh...
freesystems.substack.com
The Enlightened Absolutists
In 2017, OpenAI's founders warned about creating an 'AGI dictatorship.' Nine years later, we still haven't built the structures to prevent one.
020
Tom Everitt @tom4everitt.bsky.social · 21/11/2025
Keeping chains-of-thought traces reflective of the models true reasoning would be very helpful for safety. Important work to explore the ways it may fail
010
Tom Everitt @tom4everitt.bsky.social · 18/11/2025
Could be. But also found this interesting about the link to universal child care www.economist.com/finance-and-...
economist.com
Universal child care can harm children
Its growing popularity in America is a concern
010
Reposted by Tom Everitt
Alex Turner @turntrout.bsky.social · 04/11/2025
New Google DeepMind paper: "Consistency Training Helps Stop Sycophancy and Jailbreaks" by @alexirpan.bsky.social, me, Mark Kurzeja, David Elson, and Rohin Shah. (thread)
The abstract of the consistency training paper.
1185
Reposted by Tom Everitt
Joel Z Leibo @jzleibo.bsky.social · 31/10/2025
[1/9] Excited to share our new paper "A Pragmatic View of AI Personhood" published today. We feel this topic is timely, and rapidly growing in importance as AI becomes agentic, as AI agents integrate further into the economy, and as more and more users encounter AI.
35415
Tom Everitt @tom4everitt.bsky.social · 29/10/2025
"We think that Mars could be green in our lifetime This is not an Earth clone, but rather a thin, life-supporting envelope that still exhibits large day-to-night temperature swings but blocks most radiation. Such a state would allow people to live outside on the planet’s surface" Very cool!
020
Tom Everitt @tom4everitt.bsky.social · 15/10/2025
I was initially confused how they managed to do a randomized control trial on this. Seems they in each workflow randomly turned on the tool for a subset of the customers
010
Tom Everitt @tom4everitt.bsky.social · 09/10/2025
the focus on practical capacities is very sensible! though on basis on that, I thought you would focus on what LLMs do to humans' practical capacity to feel empathy with other beings, rather than whether LLMs satisfy humans' need to be emphasized with
000
Tom Everitt @tom4everitt.bsky.social · 02/10/2025
Interesting. Could the measure also be applied to the human, assessing changes to their empowerment over time?
110
Tom Everitt @tom4everitt.bsky.social · 02/10/2025
Interesting, does the method rely on being able to set different goals for the LLM?
100
Reposted by Tom Everitt
Toby Ord @tobyord.bsky.social · 25/09/2025
Evaluating the Infinite 🧵 My latest paper tries to solve a longstanding problem afflicting fields such as decision theory, economics, and ethics — the problem of infinities. Let me explain a bit about what causes the problem and how my solution avoids it. 1/N arxiv.org/abs/2509.19389
arxiv.org
Evaluating the Infinite
I present a novel mathematical technique for dealing with the infinities arising from divergent sums and integrals. It assigns them fine-grained infinite values from the set of hyperreal numbers in a ...
2125
Tom Everitt @tom4everitt.bsky.social · 25/09/2025
Interesting. I recall Rich Sutton made a similar suggestion in the 3rd edition of his RL book, arguing we should optimize average reward rather than discount
010
Reposted by Tom Everitt
Edward Grefenstette @egrefen.bsky.social · 21/07/2025
Do you have a PhD (or equivalent) or will have one in the coming months (i.e. 2-3 months away from graduating)? Do you want to help build open-ended agents that help humans do humans things better, rather than replace them? We're hiring 1-2 Research Scientists! Check the 🧵👇
3196
Reposted by Tom Everitt
Yoshua Bengio @yoshuabengio.bsky.social · 10/07/2025
digital-strategy.ec.europa.eu/en/policies/... The Code also has two other, separate Chapters (Copyright, Transparency). The Chapter I co-chaired (Safety & Security) is a compliance tool for the small number of frontier AI companies to whom the “Systemic Risk” obligations of the AI Act apply. 2/3
digital-strategy.ec.europa.eu
The General-Purpose AI Code of Practice
The Code of Practice helps industry comply with the AI Act legal obligations on safety, transparency and copyright of general-purpose AI models.
161
Reposted by Tom Everitt
vkrakovna.bsky.social @vkrakovna.bsky.social · 08/07/2025
As models advance, a key AI safety concern is deceptive alignment / "scheming" – where AI might covertly pursue unintended goals. Our paper "Evaluating Frontier Models for Stealth and Situational Awareness" assesses whether current models can scheme. arxiv.org/abs/2505.01420
161
Reposted by Tom Everitt
Csaba Szepesvari @skiandsolve.bsky.social · 08/07/2025
First position paper I ever wrote. "Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence" arxiv.org/abs/2506.23908 Background: I'd like LLMs to help me do math, but statistical learning seems inadequate to make this happen. What do you all think?
arxiv.org
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
Sound deductive reasoning -- the ability to derive new knowledge from existing facts and rules -- is an indisputably desirable aspect of general intelligence. Despite the major advances of AI systems ...
3519
Reposted by Tom Everitt
David Lindner @davidlindner.bsky.social · 04/07/2025
Can frontier models hide secret information and reasoning in their outputs? We find early signs of steganographic capabilities in current frontier models, including Claude, GPT, and Gemini. 🧵
161
Tom Everitt @tom4everitt.bsky.social · 27/06/2025
This is an interesting explanation. But surely boys falling behind is nevertheless an important and underrated problem?
020
Tom Everitt @tom4everitt.bsky.social · 10/06/2025
Interesting. But is case 2 *real* introspection? It infers its internal temperature based on its external output, which feels more like inference based on exospection rather than proper introspection. (I know human "intro"spection often works like this too, but still)
100
Tom Everitt @tom4everitt.bsky.social · 07/06/2025
Thought provoking
061
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
… and many more! Check out our paper arxiv.org/pdf/2506.01622, or come chat to @jonrichens.bsky.social, @dabelcs.bsky.social or Alexis Bellot at #ICML2025
arxiv.org
000
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
Causality. In previous work we showed a causal world model is needed for robustness. It turns out you don’t need as much causal knowledge of the environment for task generalization. There is a causal hierarchy, but for agency and agent capabilities, rather than inference!
120
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
Emergent capabilities. To minimize training loss across many goals, agents must learn a world model, which can solve tasks the agent was not explicitly trained on. Simple goal-directedness gives rise to many capabilities (social cognition, reasoning about uncertainty, intent…).
110
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
Safety. Several approaches to AI safety require accurate world models, but agent capabilities could outpace our ability to build them. Our work gives a theoretical guarantee: we can extract world models from agents, and the model fidelity increases with the agent's capabilities.
110
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
Extracting world knowledge from agents. We derive algorithms that recover a world model given the agent’s policy and goal (policy + goal -> world model). These algorithms complete the triptych of planning (world model + goal -> policy) and IRL (world model + policy -> goal).
100
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
Fundamental limitations on agency. In environments where the dynamics are provably hard to learn, or where long-horizon prediction is infeasible, the capabilities of agents are fundamentally bounded.
110
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
No model-free path. If you want to train an agent capable of a wide range of goal-directed tasks, you can’t avoid the challenge of learning a world model. And to improve performance or generality, agents need to learn increasingly accurate and detailed world models.
110
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
These results have several interesting consequences, from emergent capabilities to AI safety… 👇
130
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
And to achieve lower regret, or more complex goals, agents must learn increasingly accurate world models. Goal-conditioned policies are informationally equivalent to world models! But only for goals over mutli-step horizons, myopic agents do not need to learn world models.
110
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
Specifically, we show it’s possible to recover a bounded error approximation of the environment transition function from any goal-conditional policy that satisfies a regret bound across a wide enough set of simple goals, like steering the environment into a desired state.
120
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
Turns out there’s a neat answer to this question. We prove that any agent capable of generalizing to a broad range of simple goal-directed tasks must have learned a predictive model capable of simulating its environment. And this model can always be recovered from the agent.
110
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
World models are foundational to goal-directedness in humans, but are hard to learn in messy open worlds. We're now seeing generalist, model-free agents (Gato, PaLM-E, Pi-0…). Do these agents learn implicit world models, or have they found another way to generalize to new tasks?
120
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
Are world models necessary to achieve human-level agents, or is there a model-free short-cut? Our new #ICML2025 paper tackles this question from first principles, and finds a surprising answer, agents _are_ world models… 🧵 arxiv.org/abs/2506.01622
24115
Tom Everitt @tom4everitt.bsky.social · 04/06/2025
World models are foundational to goal-directedness in humans, but are hard to learn in messy open worlds. We're now seeing generalist, model-free agents (Gato, PaLM-E, Pi-0…). Do these agents learn implicit world models, or have they found another way to generalize to new tasks?
000
Tom Everitt @tom4everitt.bsky.social · 03/06/2025
I think the idea is you can use the safe AI as a filter for the unsafe one, checking the unsafe AI's plans and actions for potential harms before it proceeds
110
Tom Everitt @tom4everitt.bsky.social · 03/06/2025
Great to see serious work on non-agentic AI. I think it's an underappreciated direction: better for safety, society, and human meaning. LLMs show it's perfectly possible
072
Tom Everitt @tom4everitt.bsky.social · 02/06/2025
My feeling is that it has actually become less, sort of "oh, so AI was more like a friendly chat app than terminator"
010
Tom Everitt @tom4everitt.bsky.social · 27/05/2025
No, I didn't. But given that they apparently have many good moments, wouldn't it require quite a lot of suffering to make their life net negative? (like years of constant suffering to offset years of good times -- a brutal death and some cold/scary nights far from enough)
100
Tom Everitt @tom4everitt.bsky.social · 27/05/2025
Why? What's wrong with dysfunction = not functioning properly, or not serving the person's interests? (haven't thought deeply about it)
010
Reposted by Tom Everitt
Yoshua Bengio @yoshuabengio.bsky.social · 20/05/2025
When I realized how dangerous the current agency-driven AI trajectory could be for future generations, I knew I had to do all I could to make AI safer. I recently shared this personal experience, and outlined the scientific solution I envision @TEDTalks⤵️ www.ted.com/talks/yoshua...
ted.com
The catastrophic risks of AI — and a safer path
Yoshua Bengio — the world's most-cited computer scientist and a "godfather" of artificial intelligence — is deadly concerned about the current trajectory of the technology. As AI models race toward fu...
24812
Tom Everitt @tom4everitt.bsky.social · 19/05/2025
Yep. Maybe subjective estimates of a policy's likelihood to be beneficial/neutral/harmful would be better?
100
Tom Everitt @tom4everitt.bsky.social · 19/05/2025
I sort of agree, but it's not always easy to quantify exactly how strong the evidence is for something. It can then make sense to summarize the strength of evidence in terms of whether it's enough to support a policy
100
Tom Everitt @tom4everitt.bsky.social · 14/05/2025
Valid concern, but the application to help consensus building still seems very promising, especially given the polarisation caused by social media etc
130
Tom Everitt @tom4everitt.bsky.social · 13/05/2025
Nice video about one of our recent papers, and some of its potential implications for AI agents www.youtube.com/watch?app=de...
youtube.com
Why Don't AI Agents Work?
YouTube video by Mutual Information
050
Tom Everitt @tom4everitt.bsky.social · 07/05/2025
We explored a similar theme back in 2023, though with much less details www.alignmentforum.org/posts/Qi77Tu...
alignmentforum.org
Agency from a causal perspective — AI Alignment Forum
Post 3 of Towards Causal Foundations of Safe AGI, preceded by Post 1: Introduction and Post 2: Causality. …
031
Tom Everitt @tom4everitt.bsky.social · 07/05/2025
Agency comes in degrees, and can vary along several dimensions: autonomy, efficacy, goal-complexity, and generality. Great paper helping us understand the different possibilities
150
Tom Everitt @tom4everitt.bsky.social · 04/05/2025
METR task-time scaling critique: "the 4-minute mark for GPT-4 is completely arbitrary; you could probably put together one reasonable collection of word counting ... tasks with average human time of 30 seconds and another ... of 20 minutes where GPT-4 would hit 50% accuracy on each"
020
Reposted by Tom Everitt
Piotr Mirowski @piotrmirowski.bsky.social · 01/05/2025
Generative AI tools used in art production should be evaluated by the broader art world of artists, art historians and curators, to integrate culturally-specific critique and to re-imagine these tools to suit the artists’ needs. dl.acm.org/doi/full/10....
dl.acm.org
AI and Non-Western Art Worlds: Reimagining Critical AI Futures through Artistic Inquiry and Situated Dialogue | Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems
271
Tom Everitt @tom4everitt.bsky.social · 30/04/2025
by @benjaminreinhardt.com
010
Tom Everitt @tom4everitt.bsky.social · 30/04/2025
interesting argument about manufacturing: "[with] good enough software and hardware to create [good] humanoid robots ..., we will also be able to create more task-specific hardware ... that can do those roles cheaper, faster, and better." perhaps the same is true for AI agents more generally?
blog.spec.tech
Humanoid Robots in Manufacturing
Or, there's a reason we don't pull cars with mechanical horses
250