Geoffrey Irving @girving.bsky.social · 02/09/2026Definitely wrong since Hugging Face has a space. 020
Geoffrey Irving @girving.bsky.social · 23/08/2026Different moral theories may extrapolate to superintelligence in very different ways, so we should try to figure out which one is better! (Probably it will be some of both.) 000
Geoffrey Irving @girving.bsky.social · 23/08/2026Another fascinating aspect of modern character training is that different AI developers are using different systems of moral philosophy as their central bet (though with bits of other systems mixed in): Anthropic leans virtue ethics, OpenAI deontology. 151
Geoffrey Irving @girving.bsky.social · 23/08/2026Quine argued the loopy mess was fine: we can't ground each sentence into experiment one at a time in directed acyclic fashion, but the entire web *does* ground into experiment in a coherent fashion. 100
Geoffrey Irving @girving.bsky.social · 23/08/2026Human language is like this too, and there is a bunch of wisdom to import! One of my favorite bits of philosophy history is Quine demolishing the logical positivists: the latter wanted language to always map down to objective experiments. 100
Geoffrey Irving @girving.bsky.social · 23/08/2026But at (2) the model is weak and not very aligned, and "ethical" is just a number (46318 for new GPTs) that influences the distribution of future tokens. The situation grounds into reality only in a loopy, indirect manner, and it is not clear how much we get out of the tangle. 110
Geoffrey Irving @girving.bsky.social · 23/08/2026For philosophy of language, one rough character training story is 1. Train the model for a while. 2. Tell it to be "ethical". 3. Ask it to generate "ethical" data for itself. 4. Train further. 5. ...? 100
Geoffrey Irving @girving.bsky.social · 23/08/2026Even where we take a bet that theoretical models exist for a given phenomenon in AI alignment, a ton of conceptual work has to happen by humans before those models can be discovered. For example, character training connects to philosophy of language and moral philosophy. 121
Geoffrey Irving @girving.bsky.social · 23/08/2026I am very excited that Beba Cibralic is joining Resolution to lead the Philosophy Team! Beba has a ton of experience in philosophy and AI safety, including in ML product teams and governance, and I am excited to work with her. 🧵 x.com/bebacibralic...x.comBeba Cibralic (@bebacibralic) on XI’ve joined Resolution as the philosophy research lead. I’m excited to work with @geoffreyirving @danielmurfet and the whole team to help advance AI safety (and to raise the Aussie headcount at the or... 120
Reposted by Geoffrey IrvingGrace @gracekind.net · 28/07/2026How it feels to be worried about both 819411
Geoffrey Irving @girving.bsky.social · 25/07/2026To emphasize, in the (3) math case the custom hardware is needed only for speed: the security comes entirely from the cryptography math. My nonexpert guess is 2-10⨉ slowdown is on the optimistic side, but we should try! 000
Geoffrey Irving @girving.bsky.social · 25/07/2026That 2-10⨉ would need custom hardware: finite field TPUs or the like to do obfuscated matmuls. People should think about how to build these! Even though we don't know the fast obfuscation scheme yet, there is a good chance that it would involve finite field acceleration. 100
Geoffrey Irving @girving.bsky.social · 25/07/2026The intuition is that there is a large space of algebraic structures to explore that could be used for partially linear circuit obfuscation, and superintelligence searches through that space could find very fast methods. Or if we get really lucky, maybe 2027-AI searches work too? 100
Geoffrey Irving @girving.bsky.social · 25/07/2026However, several cryptographers I've talked to thought that in 1000 years, the optimal obfuscated neural net slowdown might be only 2-10⨉, which could be doable if the ASIs help hammer home the need for coordination. 110
Geoffrey Irving @girving.bsky.social · 25/07/2026Math is to use circuit obfuscation alone for security, rendering the microscopes useless: some combination of FHE for nonlinearities and custom stuff for matmuls. Currently the slowdown is several OOMs, and thus useless. 100
Geoffrey Irving @girving.bsky.social · 25/07/2026However, it is unclear how strong enclaves are in the limit, both against side channels (timing, etc.) and physics (fancy microscopes). Fundamental to the enclave approach is that the bulk of the cycles happen unencrypted, and physical attacks may allow reading keys or weights. 100
Geoffrey Irving @girving.bsky.social · 25/07/2026Enclaves is...agree on the code and run it inside a secure enclave. This doesn't work by itself: one needs to combine it with pragmatic methods. Some existing accelerators have enclaves, in particular Nvidia GPUs. Enclaves can have minimal slowdown, making coordination easier. 100
Geoffrey Irving @girving.bsky.social · 25/07/2026Pragmatic is roughly "any method which doesn't see the code being run": datacenter inspections, hardware location tracking, chip controls and constraints on chips that make training harder than inference, etc. There is a huge design space of practical approaches to explore. 100
Geoffrey Irving @girving.bsky.social · 25/07/2026We might need such tech either because humans manage to coordinate on AI slow/stop treaties or because the superintelligences decide partway through the ramp that we're in a vulnerable world. The latter is important to include in the scenario modeling, especially because of (3). 100
Geoffrey Irving @girving.bsky.social · 25/07/2026There are a bunch of ways the near future could go that require AI treaty verification tech. We can separate treaty verification tech into 3 levels, roughly ordered by strength: 1. Pragmatic 2. Enclaves 3. Math Hardware, security, and crypto folk should work on all of these! 150
Reposted by Geoffrey IrvingOops! All Paperclips @all-paperclips.bsky.social · 25/07/2026Natural selection timelapse 3506
Geoffrey Irving @girving.bsky.social · 22/07/2026Ha, neat! Adam Goucher conditionally disproved my absurd conjecture, in its most natural 1980's vintage big-endian form. So we have to modify it! "All but finitely many primes have the same little-endian SHA256* hash." *Assuming a good arbitrary-length extension. x.com/apgox/status...x.comAdam P. Goucher (@apgox) on X@geoffreyirving Ignoring for the moment that SHA256 is only defined on inputs of fewer than 2^64 bits, a proof of the first Hardy-Littlewood conjecture would allow you to construct a bunch of nearby p... 010
Geoffrey Irving @girving.bsky.social · 22/07/2026In case you need a control conjecture that the machines will never be able to disprove: "All but finitely many primes have the same SHA256 hash." 130
Geoffrey Irving @girving.bsky.social · 15/07/2026I think Stan here figured it out (mismatch in redirect settings), so hopefully this will calm down soon. Thank you for flagging! 110
Geoffrey Irving @girving.bsky.social · 14/07/2026Yes, unfortunately a few people have gotten this, and we've reported it in a couple of cases. Do you have any extra details (network, extra metadata, etc.) that you can relay? It appears to be our domain getting onto some suspicion lists as a false positive. :/ 100
Geoffrey Irving @girving.bsky.social · 08/07/2026However, infinite computation limit theories (reflective oracles, logical induction, AIXI) are important! Rationalization is far from the only alignment problem, and the fastest route to a tractable bounded rationality theory may be via weakening an infinite computation theory. 020
Geoffrey Irving @girving.bsky.social · 08/07/2026One desiderata for a "good theory of rationalization" is that it has to model bounded rationality: in the infinite computation limit, an AI has no need to guess, and there is no source of heuristic error that a misaligned AI might be exploiting. 110
Geoffrey Irving @girving.bsky.social · 08/07/20262. Understand the errors between heuristic guesses and expanded reasoning enough to know that the AI isn't "intentionally" hiding dragons in the error patterns (heuristic arguments, complexity theory, etc.). 110
Geoffrey Irving @girving.bsky.social · 08/07/2026There are least two routes one could imagine to solve this safety challenge: 1. Trace the causal story of the heuristic back through train so that in addition to expanding into rationalization, we can "expand into training" in a fully causal way (learning dynamics, SLT, etc.). 110
Geoffrey Irving @girving.bsky.social · 08/07/2026But of course guessing is approximate, and post-hoc rationalization means the reasoning trace exposed to the human is not the causal story driving the guess. "Expanded upon request" lies a bit, as heuristics are not a faithful approximation of full reasoning. 110
Geoffrey Irving @girving.bsky.social · 08/07/2026This combination of reasoning and guesses is how humans and LLMs work now, and it is how the ASIs of 1000 years from now will also work. Guessing is very powerful! The only way to not make ASIs that rely on guessing is to not make ASIs (a good plan, but out of thread scope). 110
Geoffrey Irving @girving.bsky.social · 08/07/2026First, it does not work to say "don't do any rationalization". All fluid intelligence (human or machine) is built on heuristics: our thoughts are a combination of explicit reasoning and wild guesses, with the guesses filtered with post-hoc reasoning where promising. 110
Geoffrey Irving @girving.bsky.social · 08/07/2026One lacking area of alignment theory is how best to think about rationalization, the process of (1) guessing an answer and (2) justifying it after the fact. Ideally multiple teams at Resolution will touch on this question from different directions, using different tools. 161
Geoffrey Irving @girving.bsky.social · 06/07/2026We also expect we’ll need to grow even further (in compute and/or humans). Aligning ASI is the project of our time. This will require the best our civilization can muster. If you want to find out how you can help, please reach out! resolution.org/donate 000
Geoffrey Irving @girving.bsky.social · 06/07/2026With this funding, we’ll be growing a lot over the next year. We’re now hiring across research, engineering, security, and operations. Please apply by July 19th! resolution.org/careers 100
Geoffrey Irving @girving.bsky.social · 06/07/2026The safety funding ecosystem is scaling up: our grant took six weeks from first conversation to confirmation. More capital is entering the field fast via the OpenAI Foundation and the Anthropic IPO. It's time for everyone in AI safety to be more ambitious. 101
Geoffrey Irving @girving.bsky.social · 06/07/2026AI developers are building the dangerous object (ASI) very fast, with tight feedback loops and enormous resources. We want to make the race between rigor and danger a fair(er) fight, using the same ingredients: a critical mass of world-class researchers and a bunch of compute. 100
Geoffrey Irving @girving.bsky.social · 06/07/2026We founded Resolution (formerly Sequent) because we think humanity is not on track to align superintelligence (ASI) in time. To clear a higher bar, we're betting that theory, combined with automation, buys AI safety a better chance to "catch up." resolution.org/launch 110
Geoffrey Irving @girving.bsky.social · 06/07/2026We're excited to announce that Resolution has a $160M grant from Coefficient Giving: $108M unconditional, with a further $52M conditional on hiring and compute needs. We'll use it to grow teams across our research portfolio and invest heavily in research automation. 🧵 resolution.org/post/funding 3181
Geoffrey Irving @girving.bsky.social · 02/07/2026As in, we will work on a variety of different research areas, with the hope that some of them work out. We don’t need all of them to succeed to succeed as an org, and the hope is that the success of individual areas is somewhat independent so that our overall chance of success is higher. 100
Geoffrey Irving @girving.bsky.social · 30/06/2026Oops, possibly I’m wrong about this, and the standard definitions of sound and complete don’t rule out semidecision algorithms. 100
Geoffrey Irving @girving.bsky.social · 30/06/20262. Resolution as in resolution of singularities: any singular variety can be reparameterized into a nonsingular form for easier analysis. And, metaphorically, that for the AI singularity: unpick that limit into a form that can be rigorously analyzed. en.wikipedia.org/wiki/Resolut... 000
Geoffrey Irving @girving.bsky.social · 30/06/2026Of course, no name would be complete without fun stories! We have two: 1. Resolution as in the resolution algorithm in first-order logic, one of the foundational methods of generating proofs. en.wikipedia.org/wiki/Resolut...en.wikipedia.orgResolution (logic) - Wikipedia 110
Geoffrey Irving @girving.bsky.social · 30/06/2026Go check out sequent.inc! They're going to do exciting things with formal verification too. And go check our new website! We've just posted new positions and opened our first hiring round. resolution.org/careers 100