On 1 August, OpenAI published solutions to ten open problems in mathematics, some of them open for decades. The model that solved them, Astra, is not public yet.
The compute cost about $2,000.
Each proof came with a Lean certificate. Lean checks every logical step, so these were confirmed correct on arrival, with no wait for human reviewers.
Most coverage said "AI does mathematics now" and moved on. That misses the interesting part.
THE PART THAT STAYED WITH ME
The model found the answer and cannot explain why
Henry Yuen at Columbia spent years of his PhD on one problem. He solved half of it in 2016, called it some of his proudest work, and always meant to return for the rest.
A month ago he tried GPT-5.5 on it. No progress. Last week Astra solved it. Here is what he wrote:
"I presume the proof is correct... I don't understand it yet. Maybe it'll take me an afternoon to check all the calculations, but what would still be missing is why this was an approach that would've made sense in the first place."
Chris Peikert, one of the best cryptographers alive, said of a different result: "most likely correct, but more importantly, it is original, elegant, and beautiful." He would nominate it for a Best Paper Award.
He noticed something odd. Astra proved its result at strength n^(1/400). Peikert spent one minute with GPT-5.6 and improved that to n^(1/14) by tidying the arithmetic.
So the model reached somewhere genuinely new, stopped far short of what its own method allowed, and cannot tell us why it looked there.
THE USELESS ONES
Hardy picked two subjects he was sure would never matter
I had all ten checked for practical use. None have one today. Three might later. Seven have no application anyone can see.
That sounds disappointing. It is normal, and it is why I am excited.
In 1940, the mathematician G. H. Hardy wrote that nobody had ever found a use for number theory or relativity, and probably never would. He picked those two on purpose, as his best examples of useless mathematics.
Number theory became RSA encryption. Relativity became GPS.
Every satellite above you corrects its clock by 38 microseconds a day for relativity. Skip it and your position is off by 11 kilometres within a day. Einstein was not thinking about navigation. That use came 63 years later.
The typical gap between a pure maths result and a real application is about 60 years, and the mathematician almost never sees it coming.
So seven of these look useless. So did relativity. Nobody can say which matters in 2085.
What is new is the rate. Results like these used to depend on how fast a few hundred people could work. Ten arrived for $2,000.
Turn AI into Your Income Engine
Ready to transform artificial intelligence from a buzzword into your personal revenue generator?
HubSpot’s groundbreaking guide "200+ AI-Powered Income Ideas" is your gateway to financial innovation in the digital age.
Inside you'll discover:
A curated collection of 200+ profitable opportunities spanning content creation, e-commerce, gaming, and emerging digital markets—each vetted for real-world potential
Step-by-step implementation guides designed for beginners, making AI accessible regardless of your technical background
Cutting-edge strategies aligned with current market trends, ensuring your ventures stay ahead of the curve
Download your guide today and unlock a future where artificial intelligence powers your success. Your next income stream is waiting.
ARE WE IN THE LOOP?
Not yet, and that tells you where to work
Has AI improving AI started compounding? The clearest evidence sits in safety filings, not announcements.
In June, Anthropic wrote that Claude produces over 80% of their merged code. Six weeks later their Opus 5 system card states they "do not observe a sustained AI-attributable 2× acceleration", and that gains are in engineering rather than research judgment. OpenAI says its model has no plausible path to the self-improvement threshold. DeepMind says No.
Blogs promote. System cards carry legal consequences if they overstate.
METR measured Claude Opus 4.6 at twelve hours of task length if you accept it working half the time, seventy minutes if you need it reliable. That gap has barely moved in three years.
Capability is climbing fast. The loop is not closed. Only the second is in doubt.
Getting a correct answer is now cheap. Knowing which question to ask is not. That is where the work is, and I do not expect that to last.
Until next week,
Martin
Was this forwarded to you? Start your 7-day free trial →
I build AI agents and GenAI systems for companies. If you're working on something like this, just reply to this email. I read every response.
The tactical part, the Maker Playbook, is for AI Insider members.
Subscribe to AI Insider to read the rest.
Become an AI Insider to get access to this post and other subscriber-only content.
Join AI InsiderAI Insider members get:
- ✅ Full access to 100% of all content.
- ✅ Exclusive DEMOs, reports, and other premium content.
- ✅ Ad-free experience.


