Sign in

green496.bsky.social

@green496.bsky.social
683 followers 268 following 5.5K posts

Anti-Trump and Anti-Republican. Fan of the American Woodcock.

PostsRepliesMedia
green496.bsky.social @green496.bsky.social · 9h
wait, uh, the entire world got rich in the past 2 years? no?
010
green496.bsky.social @green496.bsky.social · 9h
Just a reminder that you can't trust anything coming out of the Trump admin, even if it is the most sacrosanct, grave information the US government has.
000
green496.bsky.social @green496.bsky.social · 20h
The election is 3 weeks away. There's nothing that will rapidly rebound approval for Trump and Republicans in that time frame. Maybe they should have thought about this before embarking on numerous unpopular policy choices.
030
green496.bsky.social @green496.bsky.social · 21h
Direct File was one of the best government initiatives in recent memory, and would have had clear benefits had it scaled to the entire US population. So no surprise it was killed by the usual suspects.
000
green496.bsky.social @green496.bsky.social · 21h
Oops, who knew there would be consequences for Trump cutting off military aid to Ukraine
020
green496.bsky.social @green496.bsky.social · 09/10/2026
And indeed 3 of the hundreds of papers OpenAI announced earlier this week were withdrawn to serious flaws in their purported proofs: www.reddit.com/r/mathematic... Not surprisingly these did not have Lean formalizations.
010
green496.bsky.social @green496.bsky.social · 09/10/2026
Unfortunately, this bodes poorly for human understanding of AI-generated mathematics. Without high confidence in the correctness of the English language papers, mathematicians will need Lean formalizations to have confidence the proof is correct, and to build understanding on a correct foundation.
120
green496.bsky.social @green496.bsky.social · 09/10/2026
Wait - you might be wondering, how could a wrong English proof lead to a correct Lean formalization? Because even if a minor error is made in an English language proof, it may be "fixable" via some minor changes. Which is what the Lean autoformalization is forced to do (if there were errors).
100
green496.bsky.social @green496.bsky.social · 09/10/2026
However, the problem is that we no longer know that the English language paper is correct. Instead, the Lean formalization is the thing that is correct, and it is *much* harder to read than the English language paper (which is already very hard to read).
100
green496.bsky.social @green496.bsky.social · 09/10/2026
In some sense, this isn't surprising. An LLM can hallucinate/make mistakes etc. while formalizing something into Lean. Of course, the Lean compiler detects this error and returns a message explaining the error, and then the LLM tries again taking that feedback into consideration.
100
green496.bsky.social @green496.bsky.social · 09/10/2026
So what did the preprint in the first post find? It finds specific examples where a statement in the English language N-S paper does not match the corresponding portion of the Lean formalization, so in actuality the Lean formalization does not do the exact same proof as the English language paper.
100
green496.bsky.social @green496.bsky.social · 09/10/2026
In the case of N-S, everyone essentially accepts that the statement of the theorem was formalized into Lean properly, since it is a very high profile problem (and was done before OpenAI's proof anyway: eg, google-deepmind.github.io/formal-conje...)
100
green496.bsky.social @green496.bsky.social · 09/10/2026
If you have a Lean proof of a particular statement, then that lends a high degree of confidence that statement is true. Of course, you need to know that the statement that was proven in Lean is an accurate translation of the human-language statement.
100
green496.bsky.social @green496.bsky.social · 09/10/2026
However, once LLMs became good at programming, they also became "good at" Lean. In particular, they became good at translating proofs from human language (mostly English) into Lean. This means that you can point an LLM at a mathematical paper and have a good chance of formalizing it.
100
green496.bsky.social @green496.bsky.social · 09/10/2026
Before, say, 2025, a painstaking effort to formalize more and more of mathematics was gradually building momentum. Much of this took place in a library called mathlib (lean-lang.org/use-cases/ma...), which was gradually formalizing more and more of the standard corpus of mathematics.
lean-lang.org
Lean Programming Language
Lean is an open-source programming language and proof assistant that enables correct, maintainable, and formally verified code.
100
green496.bsky.social @green496.bsky.social · 09/10/2026
A theorem, definition, etc. can be stated in Lean. And its proof can also be expressed in Lean. Of course, like in real mathematics, you can build more complex mathematical concepts from simpler ones.
100
green496.bsky.social @green496.bsky.social · 09/10/2026
If you haven't heard of Lean, it is a specialized programming language one of whose aims is to allow the expression of mathematical proofs in a language that computers can verify, but is still (relatively) comprehensible to humans.
100
green496.bsky.social @green496.bsky.social · 09/10/2026
Some background: AI is notorious for hallucination. So why did mathematicians accept the correctness of OpenAI's N-S proof? The primary reason is because it was accompanied by a Lean formalization of the proof.
100
green496.bsky.social @green496.bsky.social · 09/10/2026
Earlier today I came across this preprint: arxiv.org/abs/2610.08144 TL;DR: The Navier-Stokes OpenAI natural language proof does not match the Lean proof exactly, so it is unclear if the natural language proof is actually correct.
arxiv.org
Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs
Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this pro...
110
green496.bsky.social @green496.bsky.social · 09/10/2026
Yep. AI has a selling point when it comes to productivity (eg, coding). But kids shouldn’t be thinking about that! They should be learning, exploring, and socializing with other kids!
000
green496.bsky.social @green496.bsky.social · 09/10/2026
The sooner elected Dems can get off “we’re all friends and fellow Americans here even if we disagree on policy” the better.
1391
green496.bsky.social @green496.bsky.social · 08/10/2026
Clearly, OpenAI's priorities differ from the academic community's, and this will be a source of tension for the foreseeable future.
010
green496.bsky.social @green496.bsky.social · 08/10/2026
In summary, we see that most of the suggestions were either outright ignored or mostly ignored. OpenAI did a reasonable attempt at writing papers to be comprehensible to people and formalization, but not much else.
110
green496.bsky.social @green496.bsky.social · 08/10/2026
Finally, the group recommends broad access to models used for mathematical research: "We therefore advise the AI labs to grant the global mathematical community broad, equitable access to their publicly available models." Obviously, this did not happen.
100
green496.bsky.social @green496.bsky.social · 08/10/2026
"5. Each time a solution to a problem is released, it should be clearly documented how exactly AI came to be used on that particular problem." ❌: OpenAI gave vague details about problem selection (~ 4000 total), and did not explain how they prompted for solutions nor the LLM infra that was used.
100
green496.bsky.social @green496.bsky.social · 08/10/2026
"4. As far as possible, a proof released by an AI lab should be formalized." ✅/❌: Partial credit here! Some proofs were formalized on release, but many were not. Apparently this is work in progress? It is unclear why they were not all formalized.
100
green496.bsky.social @green496.bsky.social · 08/10/2026
Hardly any prompts or chain of thought was revealed, and no precise token counts were given (only an averaged estimate of time spent per problem).
100
green496.bsky.social @green496.bsky.social · 08/10/2026
"3. For each result released, the AI lab should make public the name of the model, the prompts used, a (summarized) chain of thought, the time taken, and the estimated cost of computation" ❌ This did not happen. The model was internal (it doesn't have a public name!), ...
100
green496.bsky.social @green496.bsky.social · 08/10/2026
"2. When results are announced, they should be deposited in a timely manner in appropriate scholarly repositories..." ❌This did not happen. They put the papers in github under an anonymous submitter. The traditional location would be arXiv.
100
green496.bsky.social @green496.bsky.social · 08/10/2026
They have introductions, theorem statements, outlines of key ideas in the proof, etc. Of course, LLMs excel at this sort of summarization. However, I do not know if the details of the technical exposition are as friendly to read. LLMs are notorious for producing lots of dense jargon.
100
green496.bsky.social @green496.bsky.social · 08/10/2026
1b) "The model, or some other model, should be prompted to produce a version of each proof that is written up in a style that follows the conventions of a traditional mathematical paper..." Again, just based on a skim, this set of papers seems to do reasonably well here.
100
green496.bsky.social @green496.bsky.social · 08/10/2026
Obviously, it is impossible to audit this within a day of release. But this was one of the main complaints of the NS paper, to which additional citations were silently added a few days after the first release. So if that is any precedent, it is unlikely recommendation was systematically followed.
100
green496.bsky.social @green496.bsky.social · 08/10/2026
1a: "The literature should be scoured for any ideas that are related to the ideas in the proofs of the results released. Even if the AI lab’s model discovered those ideas independently, it should follow standard mathematical practice and cite the papers in which the ideas were first introduced."
100
green496.bsky.social @green496.bsky.social · 08/10/2026
The other case is when no human understand the AI proof (in all its detail). This is the case for essentially all recent OpenAI proofs. Here they have several detailed suggestions:
100
green496.bsky.social @green496.bsky.social · 08/10/2026
First, they outline two possible scenarios: one where a human understands the AI proof; essentially, this is where the AI is copilot, and how many mathematicians do research today. They advise no change from how mathematics has been distributed in the past.
100
green496.bsky.social @green496.bsky.social · 08/10/2026
This group is called AGMAI (Advisory Group on Mathematics and Artificial Intelligence). In late September, they released their set of recommendations: agmai.org/general-sep29/. Let's go over some of the recommendations and see which ones OpenAI followed vs. the ones they did not.
agmai.org
general-sep29
Responsible Release of AI-Generated Mathematics September 29, 2026 Back to main page Download PDF At present, some frontier AI labs are testing advanced mathematical problems on proprietary models …
110
green496.bsky.social @green496.bsky.social · 08/10/2026
One more set of meta comments on this: after OpenAI's announcement of Navier-Stokes, they got lots of pushback from the math community. One of the things OAI did in response was to ask for recommendations from an independent advisory board of mathematicians on how to conduct future releases.
agmai.org
agmai.org
Visit the post for more.
100
green496.bsky.social @green496.bsky.social · 08/10/2026
If you really want to lose faith in humanity, check out the comments on the article.
191
green496.bsky.social @green496.bsky.social · 08/10/2026
Up and down the admin, everyone’s trying to grift when they can. Whether it’s small money or big money. Just an admin filled with grifters.
010
green496.bsky.social @green496.bsky.social · 07/10/2026
Not the most important thing, but Willett can’t write clear snappy posts
000
green496.bsky.social @green496.bsky.social · 07/10/2026
Are human mathematicians obsolete? I think not… we can fairly confidently say OpenAI tried their hand basically every substantial open problem, and clearly they failed on a lot of them. So any progress made by people in the next few years is something (current) AI couldn’t do!
140
green496.bsky.social @green496.bsky.social · 07/10/2026
I think one takeaway of all this is that the existing pool of active mathematicians is not enough to saturate the space of all viable approaches using current techniques. and AI is very proficient at this.
121
green496.bsky.social @green496.bsky.social · 07/10/2026
Skimming the RH paper (I am rusty and out of date), it strikes me as similar in character to Zhang’s work on twin primes - remarkably clever use of existing techniques and theorems. But it’s not like, say, Galois’ discovery of his now namesake theory.
110
green496.bsky.social @green496.bsky.social · 07/10/2026
Right, it’s major progress. But it might be 1% of the way to the original problem (hard to say, obviously, and impossible to quantify precisely). Just an indication of how much is still out there!
110
green496.bsky.social @green496.bsky.social · 07/10/2026
It wouldn’t surprise me if that site were one of the sources of the 4000 problems OpenAI tried to bulk solve!
040
green496.bsky.social @green496.bsky.social · 07/10/2026
I’m not in the field anymore, but the caliber of these results is extraordinary. Some are incremental, but some are blockbusters (quasi Riemann!). It will be interesting to see how the community responds, both culturally and at a technical level (eg, where do people go with the quasi RH result?)
0201
green496.bsky.social @green496.bsky.social · 07/10/2026
5) There are no human names attached to any of the papers. It is impossible for anyone interested to follow up and discuss the results in a way a typical mathematical result would be discussed after announcement.
1211
green496.bsky.social @green496.bsky.social · 07/10/2026
4) They give highly obscured reasoning traces for a very small fraction of the problems. Clearly there is some special sauce in the prompting they do not want to give away.
1131
green496.bsky.social @green496.bsky.social · 07/10/2026
However, Riemann and Hodge were not attempted using this harness, and they give no further details. Presumably they used much more compute, akin to NS, for these problems.
1100
green496.bsky.social @green496.bsky.social · 07/10/2026
3) They give a partial, highly obscured, estimate of costs. They say almost all problems were solved using their internal model, averaging 3 hours of computer. They attempted ~4000 problems and solved about 700.
1111