AI research

OpenAI Astra: what is verified, from maths results to agentic work

OpenAI now has primary sources identifying Astra as an upcoming model with mathematics and cyber-capability evidence. The broader 'answers to assignments' story is an interpretation that needs careful sourcing.

By Kendr Research7 min readUpdated August 12, 2026
A structured technical review representing mathematical proof and verification
Quick answer

The old version of this article was too skeptical: OpenAI has since published primary sources identifying Astra as an upcoming model. On 1 August 2026, OpenAI said an internal version of Astra produced ten mathematics and theoretical computer-science results and released paper, walkthrough, and Lean-certificate material. On 7 August 2026, OpenAI also said Astra showed significant advances in agentic coding and cybersecurity, enough that it could not rule out a Critical cyber-capability level. That verifies the existence and importance of OpenAI Astra, but it does not turn every secondary 'digital staff' claim into an official product fact.

First, separate the Astra claims

Project Astra is Google DeepMind’s prototype for a universal AI assistant. Its official page describes natural interaction, video and screen understanding, memory, personalization, tool use, and assistance across phones and prototype glasses. It does not present Astra as a theorem prover or identify it as the system behind a major mathematical result.[1]

OpenAI Astra is now separately documented. OpenAI’s 1 August 2026 mathematics publication says ten results were achieved by an internal version of Astra, described there as OpenAI’s next major model, and points readers to paper, walkthrough, and Lean-certificate material. That is primary evidence for an OpenAI Astra mathematics claim, not a social-only rumor.[8]

The Medium article framed Astra as a shift from answers to assignments: memory, delegation, software operation, and more staff-like work. Treat that as a secondary interpretation unless OpenAI has documented the exact product surface. The official record does support a broader agentic direction: OpenAI’s 7 August 2026 security post says Astra advanced in agentic coding and cybersecurity, while the GPT-5.6 launch describes long-running professional workflows, programmatic tool use, and multi-agent coordination in already released models.[9][10][11]

The verified mathematics timeline

In July 2024, DeepMind reported that AlphaProof and AlphaGeometry 2 solved four of six International Mathematical Olympiad problems. External mathematicians scored the work at 28 of 42 points, the top end of the silver-medal range that year. The system used formal mathematics, required manual translation of the questions, and took up to three days on some problems. Those conditions matter when comparing it with a student sitting the contest.[2]

In July 2025, an advanced Gemini model with Deep Think produced natural-language proofs for five of six IMO problems within the competition’s 4.5-hour limit. IMO coordinators certified 35 of 42 points, a gold-medal-level result. This was a meaningful change in interface and speed: natural-language input and output replaced the earlier formal-language translation pipeline. It was still a controlled evaluation of six competition problems, not proof that every mathematical answer from a general assistant is reliable.[3]

From contest solutions to new mathematical results

AlphaEvolve, announced in May 2025, combined Gemini models with automated evaluators and an evolutionary loop to improve algorithms. One reported result multiplied 4×4 complex-valued matrices using 48 scalar multiplications, improving the previously known construction for that setting. Because an evaluator can run code and rank candidates, this style of system is strongest where progress has a clear machine-checkable score.[4]

The frontier moved again in 2026. DeepMind described Aletheia as an agent that generates, verifies, and revises work on professional research problems. OpenAI then published a proof that disproved a longstanding conjecture connected to the planar unit-distance problem, together with a proof artifact and external mathematicians’ assessment. On 1 August 2026, OpenAI went further by saying an internal version of Astra produced ten mathematics and theoretical computer-science results, with manuscripts, walkthroughs, and Lean certificates released for inspection.[5][6][8]

What counts as ‘solved’ in mathematics?

The word solved hides several different achievements. A model can get a short answer right, write an olympiad proof accepted by graders, formalize a proof in Lean, improve an algorithm against an executable objective, suggest a promising lemma, or resolve an open research problem that survives expert review. These are not interchangeable. The stronger the claim, the more important the verification process becomes.

A useful evidence ladder is: reproducible answer; complete proof; independent expert grading; formal verification where practical; public artifact; and later scrutiny by the research community. OpenAI’s First Proof report is a good illustration of why uncertainty must stay visible: it published ten attempts, said at least five had a high chance of being correct, and explicitly withdrew confidence in one attempt after feedback. That correction is evidence of a functioning review process, not a reason to pretend the initial result never happened.[7]

  • Check whether the task was an established benchmark, a new problem, or an open conjecture.
  • Check whether the system used tools, search, multiple samples, formal languages, or human selection.
  • Look for the actual proof or program, not only a press summary.
  • Prefer independent grading and subsequent review over a model provider’s self-score.

What the breakthroughs mean—and what they do not

The strongest conclusion is that frontier systems can now sustain longer mathematical arguments, use verification loops, and sometimes produce genuinely useful search over proof or algorithm spaces. This makes them promising collaborators for conjecture generation, literature navigation, formalization, counterexample search, and checking tedious cases.

The evidence does not justify treating a conversational model as an infallible mathematician. Research proofs can contain a subtle gap, benchmark questions can leak into training data, and a correct final result can be supported by an invalid derivation. Mathematical use should preserve the full argument and invite a person—or a proof assistant—to challenge every step.

A practical way to read the next AI-maths headline

Start with the noun: Google Project Astra, OpenAI Astra, GPT-5.6, Aletheia, AlphaProof, or another system? Then identify the verb: answered, proved, formalized, optimized, delegated, operated software, or coordinated agents? Next record the conditions: time limit, tools, sample count, human interventions, access to software, and whether the task was public before testing. Finally, inspect who checked it and whether the artifact is available.

This discipline avoids two opposite mistakes. One is hype—turning a narrow, well-scaffolded result or secondary article into a claim of universal autonomy. The other is dismissal—ignoring a real advance because earlier rumors were thinly sourced. The useful answer is usually claim-by-claim: OpenAI Astra is real and source-backed; the exact consumer or enterprise workflow implied by 'answers to assignments' still needs official product evidence.

Frequently asked questions

Is Google Project Astra a mathematics model?

No stable primary source reviewed for this article describes Google Project Astra as a mathematics model. DeepMind presents it as a universal-assistant research prototype with multimodal perception, memory, and tool use.

Is OpenAI Astra real?

Yes. OpenAI has identified Astra as an upcoming model in primary posts dated 1 August 2026 and 7 August 2026. The first ties an internal version of Astra to ten mathematics and theoretical computer-science results; the second discusses Astra cyber-capability evaluations and security controls.

Does the Medium 'answers to assignments' article prove a launched Astra product?

No. It is useful as a secondary framing of the agentic direction, but official product claims still need OpenAI documentation for the exact memory, delegation, software-operation, availability, and safety boundaries.

Which AI reached IMO gold-medal level?

DeepMind reported that an advanced Gemini model with Deep Think earned 35 of 42 points at IMO 2025. OpenAI also reported a separate 35-of-42 gold-level result from a general-purpose reasoning model in 2025.

Can AI-generated proofs be trusted without review?

No. A proof should be inspected by qualified reviewers and, where feasible, checked with formal tools. Providers themselves have revised mathematical claims after expert feedback.

Sources and evidence

Primary and authoritative sources used for factual claims. Company research and executive forecasts are labeled as such in the article.

  1. 1
    Project AstraGoogle DeepMind
  2. 2
  3. 3
  4. 4
    AlphaEvolveGoogle DeepMind · 2025-05-14
  5. 5
    Gemini Deep Think and AletheiaGoogle DeepMind · 2026-02-11
  6. 6
  7. 7
    Our First Proof submissionsOpenAI · 2026-02-20
  8. 8
  9. 9
  10. 10
  11. 11