Sign in

gavin.codes

@gavin.codes
11 followers 7 following 84 posts
PostsRepliesMedia
gavin.codes @gavin.codes · 18/06/2026
$0.06 for a non-trivial proof about programming languages. This is going to change everything.
100
gavin.codes @gavin.codes · 17/06/2026
Type preservation for PCF+Ref now at 4c on DeepSeek + Opencode + Rocqpiler 4 cents ... proof is cheaper than bananas
000
gavin.codes @gavin.codes · 17/06/2026
The approach to Vericoding in axiomander in a nutshell
000
gavin.codes @gavin.codes · 17/06/2026
This is essentially a "compiler" for specifications — the dream of declarative programming. The change will be as big as shift as the shift away from assembly. But is this feasible? YES! We've already shown how it can be used for contracts on sequential python. This is the answer to AI slop.
000
gavin.codes @gavin.codes · 17/06/2026
1/2 Can we safely let AIs do almost all of the coding? What do we need to avoid getting mountains of slop? * Small control surface [Specification] * Deterministic proof of correctness [Proof assistant]
100
gavin.codes @gavin.codes · 11/06/2026
Proving properties of software has historically been hard because proof is hard. But if we don't do the proofs, then the hard problem resolves to getting the right specification. This will remain hard, but it will probably also be the main activity in the near future. Time to take this seriously!
000
gavin.codes @gavin.codes · 10/06/2026
Just tried to run my "benchmark" proof (PCF+ref type preservation) on Fable 5. It took 10 minutes and did not make ONE SINGLE ERROR. No tactics out of place, no iteration. Just full Rocq proofs the first time. This is mental.
110
gavin.codes @gavin.codes · 04/06/2026
3/3 Axiomander, our MCP for proving properties of python, shows how easy it is to augment python with this capability. Why isn't everyone already doing this?
000
gavin.codes @gavin.codes · 04/06/2026
2/3 Theorems which use dimensional analysis are absolutely trivial to prove and can be discharged 100% of the time as they fall into decidable theory: linear algebra. We can adorn python with units and get iron-clad assurance of correctness with no runtime cost and no additional code.
100
gavin.codes @gavin.codes · 04/06/2026
1/3 Proving very basic properties of software can have huge impacts. The NASA Mars Climate Orbiter failed due to a mismatch between Pound-seconds (Imperial) and Newton-seconds (Metric). This mistake cost $327 Million.
100
gavin.codes @gavin.codes · 30/05/2026
1/4 Here is a video of neuro-symbolic theorem proving in action using the (opensource) Rocq-piler MCP (we renamed when updating to Rocq) with the Rocq theorem prover.
100
gavin.codes @gavin.codes · 24/05/2026
1/2 One shotted type preservation for PCF+references, including obtaining all necessary lemmata for the proof using the latest version of our MCP for coq (github.com/scidonia/mcp...) using DeepSeek v4.
101
gavin.codes @gavin.codes · 16/04/2026
Cross-checking generated content increases quality substantially, but what models should I use to *cross-check* generated content. It turns out for entity extraction you can use almost anything! Use the model which is 200x cheaper. See below 👇
100
gavin.codes @gavin.codes · 14/04/2026
Riffing on an idea that is making the rounds of using a wiki knowledge base to drive your LLM, we tried this approach just for chat short term memory. The results were very sensitive to the ontology and organisation, we get better results than raw summary. It works amazingly well! See below 👇
120
gavin.codes @gavin.codes · 13/04/2026
Ontologies can be helpful in structuring short term memory for interactive agents. We developed a methodology for short-term memory assessment and demonstrate a comparison between naive summarisation and ontology centric memory. Full discussion and methodology below 👇
110
gavin.codes @gavin.codes · 10/04/2026
Entity extraction from documents can provide huge benefits to organisations. If your model is hallucinating information - it will pollute downstream assets. The key is appropriate provenance tracking and cross checking. Methodology link below 👇
100
gavin.codes @gavin.codes · 09/04/2026
Are you using the right models to analyze your documents? When performing large-context entity extraction, model matters and cross checking has an important impact on precision. It is easier for models to check than to generate, so checking is often worthwhile.
010
gavin.codes @gavin.codes · 11/12/2025
In the very near future code will be organised more around contract than implementation, just as currently we organise around high-level constructs and not machine code.
000
gavin.codes @gavin.codes · 11/12/2025
Generative AI is finally making software verification not just practical, but it a requirement (pun intended). AI slop is a serious problem, but it is also an opportunity.
110
gavin.codes @gavin.codes · 11/12/2025
At bookwyrm.ai we've been performing experiments in Neural-symbolic programming and we're getting amazing results. Generative AI driven by both specification and formal verification.
121
gavin.codes @gavin.codes · 09/12/2025
7/ And lo and behold, during my exploration of the space of tools to assist in this I found that Bertrand Meyer the author of DbC has been thinking along very similar lines and has made the case in a pre-print! se.inf.ethz.ch/~meyer/publi...
000
gavin.codes @gavin.codes · 09/12/2025
6/ Step A improves discoverability of code for re-use by the LLM. Step D reduces the burden on programmers to ensure correctness. Both address serious problems with current vibe-coding practices.
100
gavin.codes @gavin.codes · 09/12/2025
5/ D) Use the contracts for static and dynamic verification. The boolean specifications can be used to do hypothesis testing we can derive test data (and therefore tests) automatically. SMT solvers can be used to check pre-post conditions.
100
gavin.codes @gavin.codes · 09/12/2025
4/ C) Have an LLM write the implementation. With a combination of editor mode fencing which stops the LLM from editing the specification we can drive an LLM with a very explicit definition of correctness helping it to get the correct answer.
100
gavin.codes @gavin.codes · 09/12/2025
3/ B) Write predicates which check the conditions of the contract (pre/post-conditions and invariants). In my experiments I have used python itself to define boolean functions which act as predicates.
100
gavin.codes @gavin.codes · 09/12/2025
2/ A) Write contracts in conversation with an LLM to improve requirements gathering and specification completeness. LLMs are quite good at getting people to clarify written text and can accelerate things substantially.
100
gavin.codes @gavin.codes · 09/12/2025
🧵1/ I've been conducting experiments with the use of LLMs for "Design by Contract" (DbC), a paradigm described by Bertrand Meyer. DbC is quite straightforward to use in a language like python (for instance using icontract). The idea is essentially to:
131
gavin.codes @gavin.codes · 04/12/2025
6/ We've a chicken-and-egg type situation at the minute. Probably python will continue to be optimal in that you can add type annotations while not requiring type discipline while also getting some advantages of static analysis.
100
gavin.codes @gavin.codes · 04/12/2025
5/ Types are also of interest. While types make it harder for AIs to one shot (they require more loops) they tend to get better code and the types themselves tend to constrain the output to more *correct* code. I'm having trouble finding proof of this in the literature.
100
gavin.codes @gavin.codes · 04/12/2025
4/ This suggests that the semantics of javascript are inherently more complex for LLMs to get a handle on since the volumes must be very high. It could well be that typescript just doesn't have enough data to fix the problem.
100
gavin.codes @gavin.codes · 04/12/2025
2/ Volume is key to training AIs, and so those programming languages with large training volumes tend to win. Python is killing it. However there are some secrets hidden in this chart.
100
gavin.codes @gavin.codes · 04/12/2025
1/ If you're using AIs to generate code, you might want to have a look at this chart. Language choice now should also take into account AI comprehension and writing abilities.
110
gavin.codes @gavin.codes · 01/12/2025
9/ And best of all the entire methodology through step 7 can be treated as a kind of increasing level of trust in the software without us having to pay human attention much beyond the logical contract.
000
gavin.codes @gavin.codes · 28/11/2025
6/ It's worth recalling a claim of two famous MIT professors "Computer science is not a science, and its ultimate significance has little to do with computers" ­— Abelson / Sussman
000
gavin.codes @gavin.codes · 28/11/2025
5/ Two contenders for a basic foundation which can be used as hand written pseudocode that strike me are Z-notation and APL. Probably we would need to adapt approaches but these are compact enough and mathematical enough that they can do the job. Image
100
gavin.codes @gavin.codes · 28/11/2025
4/ I've been putting some thought into how the curriculum would have to change — typical programming languages are largely too verbose to comfortably write on the blackboard or with pen.
110
gavin.codes @gavin.codes · 28/11/2025
3/ Before students learn to drive an AI to assist in programming, they need a firm grasp of deep concepts in CS. Algorithms, relationship between proof and algorithms and complexity theory.
100
gavin.codes @gavin.codes · 28/11/2025
2/ AI presents too many opportunities for laziness. As long as assignments are on computer there is no way we can avoid the use of AI to do the assignments. If the assignment is a lot of boilerplate (which humans will no longer do) then this is useless learning for the AI age.
100
gavin.codes @gavin.codes · 28/11/2025
1/ I used to lecture in CS. Due to AI, I can tell you if I did it now I would not allow any computers in my class room at all. Assignments would be pen and paper. Students need to learn to reason, first and foremost.
110
gavin.codes @gavin.codes · 26/11/2025
6/ Going to a complete liquid metal fuel might be viable for other fuel arrangements which are not so useful for bombs. In addition Fast reactors can run on nuclear waste, reprocessed from other reactors which can eliminate one of the great bug-bears of the nuclear solution: waste.
000
gavin.codes @gavin.codes · 26/11/2025
5/ Maybe having a liquid plutonium core reactor is a bit crazy - it certainly violates any modern concept of non-proliferation! However the ideas contained here are interesting none-the-less.
100
gavin.codes @gavin.codes · 26/11/2025
4/ These advantage of such a fast reactor is that the increasing heat induces a negative reactivity response, as the fuel expands. This helps to make the reactor passively self-stabilising.
100
gavin.codes @gavin.codes · 26/11/2025
3/ This reactor design known as a fast reactor as it used no moderator, simply placed Plutonium in tantalum thimbles - placed the thimbles in proximity and the resulting reaction heated it up until it was molten. A sodium coolant flowed over the thimbles and carried away the heat.
100
gavin.codes @gavin.codes · 26/11/2025
2/ But what if your reactor was already melted! Well, there is more than one design of this type, including an existing reactor in China (itself based on a previous US design: the Molten Salt Reactor Experiment). But LAMPRE took this a step further.
100
gavin.codes @gavin.codes · 26/11/2025
1/ Reactors I have known and loved: LAMPRE One of the weirdest and most innovative reactors that I've come across is the LAMPRE reactor. As some may know, one of the biggest safety problems that nuclear reactors experience is a meltdown, which usually occurs from a loss of coolant event.
131