Mike Dodds @m-dodds.bsky.social · 26/08/2026Some personal news: I’ve left @galoisinc.bsky.social to start Oath Technologies / @oathtech.bsky.social, a new FRO that will work on AI oversight via formal methods 040
Mike Dodds @m-dodds.bsky.social · 26/08/2026How do we use AI to build formal tools? Notes from a few months of building semantics, verifiers, and proof toolchains with AI agents doing most of the work oath.tech/pub/2026/08/... 050
Mike Dodds @m-dodds.bsky.social · 18/03/2026Someone should build seL4-ablate-bench. Progressively delete proofs, lemmas, theorems and see how much a long-running AI agent can reconstruct. End state: just give it the code + top spec, and rebuild the whole 1m+ line Isabelle proof 011
Mike Dodds @m-dodds.bsky.social · 05/03/2026Context: Knuth on Claude www-cs-faculty.stanford.edu/~knuth/paper...www-cs-faculty.stanford.edu 000
Mike Dodds @m-dodds.bsky.social · 05/03/2026I formalised the Knuth / Stappers / Claude theorem in Lean4. Claude for scaffolding and Harmonic‘s Aristotle AI for the core proofs This is just the construction Claude found, not all 760 constructions (Disclaimer: theorems look plausible to me, but mistakes possible) github.com/septract/cla...github.comGitHub - septract/claudes-cycles-leanContribute to septract/claudes-cycles-lean development by creating an account on GitHub. 130
Mike Dodds @m-dodds.bsky.social · 22/02/2026Commands: #tcb (what's in your trust base), #tcb_tree (dependency graph), #tcb_why (why is this included?). v0.1.0, rough edges expected, feedback welcome github.com/OathTech/lea... 020
Mike Dodds @m-dodds.bsky.social · 22/02/2026Weekend project w/ Claude: in Lean, it can be hard to know which definitions you need to review to trust a theorem. So I built lean-tcb. It figures out your trusted computing base, ie the definitions that actually give a theorem its meaning (vs proof machinery the kernel checks)github.comGitHub - OathTech/lean-tcbContribute to OathTech/lean-tcb development by creating an account on GitHub. 130
Mike Dodds @m-dodds.bsky.social · 10/10/2025I got curious whether Claude Code could handle a low-representation theorem prover like ACL2 - turns out yes! I proved a bunch of small to medium theorem, and for good measure built a MCP server, all in about 4 hrs. I’ve never used ACL2 before. Write-up here: mikedodds.org/posts/2025/1...mikedodds.orgExperimenting with ACL2 and Claude CodeTL;DR: Using only prompting with Claude Code, I created: 50+ ACL2 theorem proofs translated from Software Foundations An MCP server for ACL2 with stateful solver sessions 070
Mike Dodds @m-dodds.bsky.social · 21/09/2025That’s true, and I think that’s exactly why Claude does so well proofs. It‘s just I happen to know first hand that proofs are a particularly difficult *kind* of program 000
Mike Dodds @m-dodds.bsky.social · 16/09/2025I wrote about Claude Code, which to my absolute astonishment is quite good at theorem proving. For people who don't know theorem proving, this is like spending your whole life building F1 engines and getting lapped by a Tesco's shopping trolley www.galois.com/articles/cla...galois.comClaude Can (Sometimes) Prove It 1165
Mike Dodds @m-dodds.bsky.social · 16/07/2025New Galois blog: “Specifications Don’t Exist”. If we want to formally verify more systems, we need formal specifications, but most real systems are hard to specify for very deep reasons www.galois.com/articles/spe... 071
Reposted by Mike DoddsHazel Weakly @hazelweakly.me · 25/06/2025I’m not sure how I missed this but it’s an extremely good article and you should absolutely read it. It’s about formal methods, but anyone who cares about integrating research into industry will find it valuable! I saw a *ton* of parallels with resilience engineering too :) 1134
Mike Dodds @m-dodds.bsky.social · 25/06/2025Hey :) Seems like a lot of people moved here from old Twitter and I’m still catching up 010
Reposted by Mike DoddsGalois @galoisinc.bsky.social · 27/05/2025At Galois, we often say things like: “Formal methods form the backbone of everything we do.” But what exactly are formal methods? How do they work, and why are they so important? We created a handy reference page to explain: www.galois.com/what-are-for... 012
Mike Dodds @m-dodds.bsky.social · 24/05/2025If a tool is not popular, it’s uncompelling to argue that everyone is just mistaken. At some point you should ask why the tool isn’t useful (at the current cost/benefit point) 010
Mike Dodds @m-dodds.bsky.social · 24/05/2025New-ish @galoisinc.bsky.social blog: “What Works (and Doesn't) Selling Formal Methods”. The boring truth: engineers are rational and adoption is all about cost/benefit tradeoffs www.galois.com/articles/wha... 151
Reposted by Mike DoddsGalois @galoisinc.bsky.social · 08/05/2025What actually works when selling formal methods in industry? What doesn't? The way Galois Principal Scientist @m-dodds.bsky.social sees it, many FM projects don’t pencil out not because clients are irrational, but because the cost/benefit tradeoffs don’t make sense. www.galois.com/articles/wha... 034
Reposted by Mike DoddsGalois @galoisinc.bsky.social · 14/04/2025c2rust is available on the Godbolt Compiler Explorer! c2rust is a tool we developed with Immunant that can convert nearly any piece of C code into compilable Rust godbolt.org/z/crsWEGEKMgodbolt.orgCompiler Explorer - C (C2Rust (master))/* Type your code here, or load an example. */ int square(int num) { return num * num; } 1132
Mike Dodds @m-dodds.bsky.social · 06/02/2025Formal methods go great with AI www.wsj.com/articles/why...wsj.comWhy Amazon is Betting on ‘Automated Reasoning’ to Reduce AI’s HallucinationsAmazon is using math to help solve one of artificial intelligence’s most intractable problems: its tendency to make up answers, and to repeat them back to us with confidence. 050
Mike Dodds @m-dodds.bsky.social · 29/01/2025I wrote about o3, the Frontier Math benchmark, and what it means if AI math keeps getting better 071
Mike Dodds @m-dodds.bsky.social · 21/01/2025I don’t think literally everyone should drop what they’re doing. But my sense is PL research as a whole is significantly under-reacting to AI. So I suppose I think *some more* PL people should bet on AI (but maybe not you!) 110
Mike Dodds @m-dodds.bsky.social · 21/01/2025Happy to mail you a couple. Email me, my address is on my website 200
Mike Dodds @m-dodds.bsky.social · 21/01/2025I think you’ve put your finger on the exact worldview mismatch because 5-10 years seems like an insanely long time horizon to me 110
Mike Dodds @m-dodds.bsky.social · 21/01/2025Why constrain the grammar - just pull more samples and keep the ones that pass :p 110
Mike Dodds @m-dodds.bsky.social · 20/01/2025Hot take for POPL: the PL community is still mostly in denial about AI. This is bad because PL+AI go great together - PL can solve the hardest problem with AI - trusting the output it produces - AI can solve the hardest problem with PL - finding enough engineers who can even use the tools 2141
Mike Dodds @m-dodds.bsky.social · 20/01/2025I’m bringing these cute Galois stickers to POPL so if you want one, come find me 1110
Mike Dodds @m-dodds.bsky.social · 27/12/20248 years on, the future is here! xkcd.com/1813/xkcd.comVomiting Emoji 040
Reposted by Mike DoddsHillel @hillelwayne.com · 27/12/2024emojikitchen.devemojikitchen.devEmoji Kitchen - Browse Google's unique emoji combinationsUnique illustrations of combined emoji, cooked up in Google's Emoji Kitchen, and comprehensively available on the web 273
Mike Dodds @m-dodds.bsky.social · 21/12/2024If I understand right, the private test set is only used during evaluation of the model - not available to the team doing the training 120
Mike Dodds @m-dodds.bsky.social · 21/12/2024Seems almost certain it’s deliberately trained on math reasoning. The way the o-series models seem to work is by long CoT, with reinforcement learning to impose correct reasoning. Not much public about how o3 works internally, but Chollet has some speculation: arcprize.org/blog/oai-o3-...arcprize.orgOpenAI o3 Breakthrough High Score on ARC-AGI-PubOpenAI o3 scores 75.7% on ARC-AGI public leaderboard. 120
Mike Dodds @m-dodds.bsky.social · 21/12/2024Re o3 - this is the big one for me. The Frontier Math benchmark is designed to be extremely difficult, and it has a private test set (no data contamination). Today, o3 is v expensive. But seems inevitable it’ll soon be cheap. If these results hold up, that means MUCH more powerful automated math 332
Mike Dodds @m-dodds.bsky.social · 17/12/2024I think specification will be a much harder problem. Nearly all successful proof deployments have been in “easy to specify” domains - OSs, hardware, crypto, etc. these are unusual, & most systems are very difficult to specify formally 020
Mike Dodds @m-dodds.bsky.social · 17/12/2024Optimistically this could haul a lot of tools across the break-even line into viability. There are many formal methods ideas that simply haven’t been tried because the cost/benefit never worked out. Exciting times for proof tech / FM 100
Mike Dodds @m-dodds.bsky.social · 17/12/2024I think there’s good reason to be optimistic that proofs themselves will get much cheaper. Most proof tools are structured as untrusted search and trusted checking. Gen AI is a just new untrusted search process which should slot right in alongside SMT solving etc 100
Mike Dodds @m-dodds.bsky.social · 17/12/2024I gave a talk recently about proof technologies - what people deploy today, what might be available soon, and what seems far off even with fancy AI. Slides here: mikedodds.github.io/files/talks/... 1175
Mike Dodds @m-dodds.bsky.social · 16/12/2024(& yes, I’d be excited to hear more about what you’re working on!) 000
Mike Dodds @m-dodds.bsky.social · 16/12/2024Completely agree. Today’s LLMs are nowhere near the hardest verification tasks, and getting there will take more leaps. I’m not sure what’s needed - more special purpose tooling, or just more generic intelligence from the LLM+RL combo. To be determined I think 100
Mike Dodds @m-dodds.bsky.social · 16/12/2024I think LLMs + reinforcement learning + trad synthesis seems quite promising for inductive invariants. There’s some hope that AI can knock out the 80% of “easy” cases, even if a hard core remains 110
Mike Dodds @m-dodds.bsky.social · 16/12/2024Yes, totally agree the specific claim matters. I sometimes mean “this is astonishing and seems like it can automate many tasks we care about” and maybe people are hearing “AGI is here, this can code better than a human”. I do aim for a little more nuance in person than on social media though :) 110
Mike Dodds @m-dodds.bsky.social · 16/12/2024Well, it’s hard to say because in the skeptics typically don’t want to make clear bets :) But my sense in such conversations is they are denying capabilities that are already here. An SMT solver (or a printf statement) can generate a FramaC spec, that’s not what we’re talking about 120
Mike Dodds @m-dodds.bsky.social · 16/12/2024I think progress is astoundingly rapid and it seems plausible (not certain!) that AIs will routinely beat human experts on many of these tasks soon 100
Mike Dodds @m-dodds.bsky.social · 15/12/2024Obviously, even the most capable current LLM can’t do these kinds of tasks perfectly, every time, for big programs etc. But surprisingly often people will fully deny they can do them *at all* 120
Mike Dodds @m-dodds.bsky.social · 15/12/2024A few off the top of my head: writing function specifications in eg FramaC, generating correct programs based on free-text problem descriptions, solving logic puzzles that require structured reasoning, finding bugs in code 120
Mike Dodds @m-dodds.bsky.social · 15/12/2024I have had conversations with professor types who say “oh I don’t think an LLM will be able solve <whatever> for a long time” and I show them the base ChatGPT model doing <whatever> first time with simple prompting. Many people’s intuitions are stuck (especially LLM critics) 181
Mike Dodds @m-dodds.bsky.social · 15/12/2024Yeah I think a lot of people strongly dislike LLMs and that means they haven’t really understood what they can do right now. Never mind what seems plausible in 5 years 040
Reposted by Mike DoddsSam Tobin-Hochstadt @samth.bsky.social · 15/12/2024I think many of the (quite gross) reactions to this are not grappling yet with how many their students already have what they think is this product in the form of chatgpt. 39711
Mike Dodds @m-dodds.bsky.social · 09/12/2024I’m a bit skeptical the CEO murder is really a v popular thing outside left social media 240
Mike Dodds @m-dodds.bsky.social · 07/12/2024Btw I have a 2nd bsky account for curating such papers so thanks for doing my homework :) @mdai.bsky.social 000