Sign in

Geoffrey Irving

@girving.bsky.social
4.1K followers 120 following 682 posts

Cofounder and Chief Scientist at Resolution. Alignment will be solved eventually, but not necessarily in time. Previously UK AISI, DeepMind, OpenAI, Google Brain, etc.

PostsRepliesMedia
Geoffrey Irving @girving.bsky.social · 23/08/2026
I am very excited that Beba Cibralic is joining Resolution to lead the Philosophy Team! Beba has a ton of experience in philosophy and AI safety, including in ML product teams and governance, and I am excited to work with her. 🧵 x.com/bebacibralic...
x.com
Beba Cibralic (@bebacibralic) on X
I’ve joined Resolution as the philosophy research lead. I’m excited to work with @geoffreyirving @danielmurfet and the whole team to help advance AI safety (and to raise the Aussie headcount at the or...
120
Geoffrey Irving @girving.bsky.social · 12/08/2026
080
Reposted by Geoffrey Irving
Grace @gracekind.net · 28/07/2026
How it feels to be worried about both
this is fine dog
819411
Geoffrey Irving @girving.bsky.social · 25/07/2026
There are a bunch of ways the near future could go that require AI treaty verification tech. We can separate treaty verification tech into 3 levels, roughly ordered by strength: 1. Pragmatic 2. Enclaves 3. Math Hardware, security, and crypto folk should work on all of these!
150
Reposted by Geoffrey Irving
Oops! All Paperclips @all-paperclips.bsky.social · 25/07/2026
Natural selection timelapse
3506
Geoffrey Irving @girving.bsky.social · 22/07/2026
Ha, neat! Adam Goucher conditionally disproved my absurd conjecture, in its most natural 1980's vintage big-endian form. So we have to modify it! "All but finitely many primes have the same little-endian SHA256* hash." *Assuming a good arbitrary-length extension. x.com/apgox/status...
x.com
Adam P. Goucher (@apgox) on X
@geoffreyirving Ignoring for the moment that SHA256 is only defined on inputs of fewer than 2^64 bits, a proof of the first Hardy-Littlewood conjecture would allow you to construct a bunch of nearby p...
010
Geoffrey Irving @girving.bsky.social · 22/07/2026
In case you need a control conjecture that the machines will never be able to disprove: "All but finitely many primes have the same SHA256 hash."
130
Geoffrey Irving @girving.bsky.social · 22/07/2026
Riemann is still true tho.
000
Geoffrey Irving @girving.bsky.social · 08/07/2026
One lacking area of alignment theory is how best to think about rationalization, the process of (1) guessing an answer and (2) justifying it after the fact. Ideally multiple teams at Resolution will touch on this question from different directions, using different tools.
161
Geoffrey Irving @girving.bsky.social · 06/07/2026
We're excited to announce that Resolution has a $160M grant from Coefficient Giving: $108M unconditional, with a further $52M conditional on hiring and compute needs. We'll use it to grow teams across our research portfolio and invest heavily in research automation. 🧵 resolution.org/post/funding
3181
Geoffrey Irving @girving.bsky.social · 30/06/2026
We’re changing our name: Sequent is now Resolution. 🧵
230
Geoffrey Irving @girving.bsky.social · 22/06/2026
A while ago I wrote a minimal replacement for top called ltop using Claude Opus 4.6-4.8. At first I made it so that I could nicely filter the processes to just the ones I wanted, but then I noticed the binary was pretty small, and got curious if it could be smaller... 🧵 github.com/girving/ltop
ltop screenshot
190
Geoffrey Irving @girving.bsky.social · 10/06/2026
We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵 sequent.org/launch
4172
Reposted by Geoffrey Irving
⚡️🌙 @dystopiabreaker.xyz · 14/10/2025
Don’t Worry — It Can’t Happen (also, the scientists who claim fission exists are just in the pocket of Big Science, and they’re literal nazis anyway, and also it’s just a stochastic reaction that peters out, and only physbros care about it, and it hasn’t ever happened before so it won’t)
410922
Geoffrey Irving @girving.bsky.social · 06/06/2026
Here is a metaphor for AGI definitions. Imagine you’re on a long drive from Los Angeles to the Bay Area (for me: undergrad to grad school). 🧵
130
Geoffrey Irving @girving.bsky.social · 29/05/2026
A while ago I wrote a Claude skill to parse Google Docs into Markdown. It was easy, but took a few iterations to get right, as the first pass had weird glitches like \! instead of !. What else needs a couple more iterations? Literally Google Doc's own agent Markdown converter.
230
Geoffrey Irving @girving.bsky.social · 24/05/2026
An inconvenient realisation today is that while I've written dozens of documents that refer to something like "pretraining on human-level data", now that phrase always has to be amended to "pretraining on (mostly) human-level data".
020
Reposted by Geoffrey Irving
⚡️🌙 @dystopiabreaker.xyz · 18/05/2026
average MLeng: 'well, the adversary would have to know the details of how our defense works' the humble Kerckhoff axiom:
1251
Geoffrey Irving @girving.bsky.social · 15/05/2026
A bittersweet announcement! For family reasons, I will be leaving AISI soon to move back to the Bay Area. I will be starting a new nonprofit alignment research org (more to come). I will miss this place! Here are some reflections about my time at AISI. 🧵❤️| naml.us/post/reflect...
2171
Geoffrey Irving @girving.bsky.social · 13/05/2026
New paper arguing that AI automation of AI alignment research could fail due to AI mistakes, even if the AI agents are intent aligned (not trying to cause harm). Arguably this is obvious: AIs make mistakes all the time (as do humans). But it is useful to go into detail.🧵 arxiv.org/abs/2605.06390
arxiv.org
Automated alignment is harder than you think
A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when resear...
190
Geoffrey Irving @girving.bsky.social · 13/05/2026
The real question is whether Homer faked a Bronze Age Greek accent when doing his Iron Age performances.
260
Geoffrey Irving @girving.bsky.social · 11/05/2026
The most delightful bug that Mythos found in Firefox is a NaN vulnerability. There are many bit representations of floating point not-a-number, one of which looks like a tagged representation of a pointer except on SPARC. You can ship it across the sandbox boundary, and BOOM.
110
Geoffrey Irving @girving.bsky.social · 11/05/2026
My 6yo’s current favorite movie is Lord of the Rings, which she is unafraid of. But Crouching Tiger, Hidden Dragon drove her into the kitchen or under blankets. Paraphrased, she explained that the latter is morally ambiguous fighting where it’s unclear who the good guys and bad guys are. Fair, tbh.
130
Reposted by Geoffrey Irving
Flo 🔶 @faz.ms · 11/05/2026
Phew, guess we can all relax about Mythos then. You're all doing the same security best practices and overall level of software engineering quality on every one of your projects as curl, right? right?
1353
Geoffrey Irving @girving.bsky.social · 19/04/2026
Arena allocators are pretty.
030
Reposted by Geoffrey Irving
Grace @gracekind.net · 19/02/2025
In a catastrophic typo, researchers ask superintelligence to optimize for CVE
6834
Geoffrey Irving @girving.bsky.social · 29/03/2026
Vibe coding parsers for untrusted data is a useful warning-sign-filled experience to have.
150
Geoffrey Irving @girving.bsky.social · 15/03/2026
Any recommended tools for following some narrow area of research, to get notifications when interesting papers come out? By narrow I mean as specific as an arbitrary LLM prompt. (I realise one could code up such a system, but I am interested in what exists without me writing it.)
251
Geoffrey Irving @girving.bsky.social · 14/03/2026
It's a shame: the 1.41e64 constant is _almost_ an LLM capability measure ("What quality expanders can the machines formalise?"), but alas it appears there's an enormous jump from what's already done (MGG) to Ramanujan graphs w/ deep number theory, with little in between.
110
Geoffrey Irving @girving.bsky.social · 13/03/2026
I'm curious how long it would take someone to make an optimised SNARK system for Lean verification, based on Lean4Lean (arxiv.org/abs/2403.14064) and arkworks.rs. 🧵
130
Reposted by Geoffrey Irving
Ryan Williams @rrwilliams.bsky.social · 14/12/2025
studying chatgpt's busy beaver number: how long can it run and still halt. finished one prompt in slightly under 24 hrs. the response was just as unhinged as a human would sound after grinding that long
1201
Geoffrey Irving @girving.bsky.social · 08/03/2026
Your periodic reminder that software engineering is a mixture of easy-to-verify subtasks and hard-to-verify subtasks, and the fact that the machines are getting better at coding assistance should not be explained away as "they're only good on easy-to-verify stuff".
050
Geoffrey Irving @girving.bsky.social · 08/03/2026
As an exercise in learning recent Claude Code + Opus 4.6, I've formalised Seiferas's simplified construction of the Ajtai-Komlós-Szemerédi O(log n)-depth, O(n log n)-size sorting networks in Lean, using Margulis–Gabber–Galil expander graphs. github.com/girving/aks
The toplevel theorems that our network has O(log n) depth and O(n log n) size.
130
Reposted by Geoffrey Irving
Grace @gracekind.net · 06/03/2026
Astronaut gun meme: 

Wait, it's all proxies and heuristics?

Close enough
327836
Geoffrey Irving @girving.bsky.social · 19/02/2026
I'm excited that AISI is announcing the first 60 Alignment Project grants, bringing more independent experts and ideas into AI alignment and control research! Since the RFP last year, we've grown the total funding to £27M. Which means more ideas will be explored! 🧵 www.aisi.gov.uk/blog/funding...
aisi.gov.uk
Funding 60 projects to advance AI alignment research | AISI Work
The Alignment Project welcomes its first cohort of grantees, and new partners join the coalition, bringing total funding to £27m.
150
Geoffrey Irving @girving.bsky.social · 17/02/2026
It is unfortunately that arxiv has no mechanism for setting a social media image for a paper. I do not need the enormous ARXIV logo for this kind of link.
090
Geoffrey Irving @girving.bsky.social · 17/02/2026
New "boundary point jailbreaking" method against LLM safeguards (with prior disclosure to multiple labs) by using noised versions of harmful queries to turn sparse feedback from failed attacks into dense feedback. 🧵 www.aisi.gov.uk/blog/boundar...
2445
Geoffrey Irving @girving.bsky.social · 13/02/2026
Anyone know if there are certificates for any sparse, symmetric positive definition matrix to be positive definite, that can be checked in quadratic time? Emphasis on *any* such matrix, no structure or other assumptions allowed other than SPD. scicomp.stackexchange.com/questions/45...
240
Geoffrey Irving @girving.bsky.social · 12/02/2026
New complexity theory paper mapping the precise query complexity of debate, given unbounded provers. No new safety ideas: the goal is a self-contained presentation of debate + cross-examination, with the precise complexity class it achieves. 🧵
151
Geoffrey Irving @girving.bsky.social · 11/02/2026
Apparently last year there was a cyberattack on the Irving Medical Center at Columbia University, which resulted in my name getting leaked. I...feel like they didn't need a cyberattack to do that in this case?
040
Geoffrey Irving @girving.bsky.social · 03/02/2026
The random oracle hypothesis is mostly true.
020
Reposted by Geoffrey Irving
Yoshua Bengio @yoshuabengio.bsky.social · 03/02/2026
Today we’re releasing the International AI Safety Report 2026: the most comprehensive evidence-based assessment of AI capabilities, emerging risks, and safety measures to date. 🧵 (1/19)
16129
Geoffrey Irving @girving.bsky.social · 02/02/2026
Being one of the two Deputy Directors of AISI's Research Unit is a very central and important role! Please apply if interested! > This isn’t your average Civil Service job. For 9–12 months, you’ll co-lead one of the world’s most influential AI safety research organisations. x.com/nateburnikel...
x.com
021
Geoffrey Irving @girving.bsky.social · 01/02/2026
One of the more useless things I did while at Google Brain was write down random access into xorshift128+, the hardware random number generator on TPUs. Purely a stunt: it could theoretically have meant TPU-native Jax-style random numbers faster than Threefry, but in practice random numbers are…
260
Geoffrey Irving @girving.bsky.social · 31/01/2026
An important thing to remember as AI develops is that, regardless of whether capabilities plateau or how far capabilities grow, computer science will still apply! AI won’t be magic, some computations will be intractable, P won’t be NP, etc. 🧵
140
Geoffrey Irving @girving.bsky.social · 31/01/2026
Achievement unlocked: trip to get passport photos as a family, but not all for the same country.
020
Geoffrey Irving @girving.bsky.social · 29/01/2026
"Nobody suspects the all-1 string."
000
Geoffrey Irving @girving.bsky.social · 17/01/2026
Symmetric block cyphers like AES and the cores of modern hash functions are roughly keyed, pseudorandom invertible functions. So a natural question is: if you pick a big enough nonlinear keyed invertible function at random, is it a secure block cipher? 🧵
140
Geoffrey Irving @girving.bsky.social · 31/12/2025
Unless... @qntm.org "Very funny polling result where ~3/4 of people will say they read a book last year but if you ask them to name the book the share drops 20 points" x.com/JosephPolita...
150
Geoffrey Irving @girving.bsky.social · 18/12/2025
New report on trends in AISI's evaluations of frontier AI models over the past two years. A lot of AI discourse focuses on viral moments, but it is important to zoom out to the less flashy trend: AI models are steadily growing in capabilities, including for dual-use. www.aisi.gov.uk/frontier-ai-...
061