Research2 min read

OpenAI's Astra Solves 10 Decades-Old Math Problems

By , Senior AI ConsultantPublished

OpenAI says an unreleased internal version of its next model, called Astra, solved ten previously unsolved math and computer science problems for about $2,000 in computing costs, and the bigger story is what kind of work AI is now good enough to attack next.

OpenAI has not released its next model yet, but it just showed off what it can do. An early, unreleased version, code named Astra, was set loose on ten problems in advanced mathematics and computer science that had stumped researchers for at least a decade, some for much longer. It solved all ten over a single weekend, at a computing cost of roughly $2,000.

To put that number in context, some of these problems had specialists working on them, on and off, for fifteen years or more. One researcher, Henry Yuen, had spent years on a piece of quantum theory that Astra extended in a single run. OpenAI did not just claim the wins either. Each solution came with a Lean proof, a format that lets a separate computer program check the logic step by step rather than trusting the AI's word for it.

Here is the twist that matters more than the headline. Once the ten problems were public, a mathematician named Levent Alpoge pointed a rival model, Anthropic's Claude, at the same list with no hints. It solved five of them within a day. That does not make Astra's result fake. It means a lot of these answers were sitting within reach of current AI, and nobody had simply asked the right question yet. OpenAI won by trying first and publishing first, not necessarily by having a model nobody else could match.

That is actually the useful insight for anyone running a business. Two years ago, AI struggling with math was a running joke. Now the joke is dead, and the reason it died so fast points to where AI will hit next. Math problems like these share one feature: you can check with total certainty whether the answer is right. Computer code has the same feature, since it either passes its tests or it does not. AI is already excellent, arguably better than most people, at tasks where the correct answer can be verified by a machine.

Most business work does not look like that. Pricing a client proposal, judging whether a supplier will deliver on time, deciding how to handle a difficult customer, these do not have a clean right answer a computer can grade. That is exactly why AI has moved through math and coding faster than it has moved through management, sales judgment, or client relationships. Expect that gap to keep widening. Any task in your business that has a clear, checkable correct answer, think pricing formulas, scheduling, compliance checks, contract clause matching, is closer to being automated well than tasks that rely on reading a room or making a judgment call.

There is a genuine warning buried in this story too. Even the mathematicians who understood the problems said they could not immediately explain why Astra's proofs worked, only that the formal check confirmed they were correct. That is a preview of a real operational risk: AI tools may soon hand your business correct answers that nobody on your team can explain or audit. Trusting output you cannot check is a different kind of risk than trusting output that is simply wrong, and most companies do not yet have a plan for it.

None of this means general intelligence has arrived. Even AI researchers who built Astra agree it is not what they would call true general intelligence, since it still needs enormous amounts of narrow training to get good at one type of checkable problem. What arrived is proof that the checkable parts of knowledge work are falling faster than most plans account for.

Stay informed

Get AI intelligence like this delivered to your inbox.


You May Also Find Valuable