
OpenAI Astra: Everything We Know About the Model That Isn't GPT-6
- z-aiNEWZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianNEWQwen3.8 27B Uncensored (Aggressive)2026-08-1552Intelligence68Coding
- qwenNEWQwen: Qwen3.8 27B (free)2026-08-1340 tok/s
- deepseekNEWDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokNEWSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaNEWMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 220 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
- grokxAI: Grok 4.52026-07-0856Intelligence72Coding
The thing you can download today is still not a model. It is ten files of Lean code. On 1 August 2026, OpenAI published a research post about mathematics, and buried in it was the first confirmation of the name of its next frontier system: Astra. It still has no model card, no pricing page, no API endpoint and no confirmed date — and the "ships next week" flash report that accompanied the announcement has been overtaken by OpenAI's own disclosure that it cannot rule out "critical" cyber capability, a finding that paused parts of development for safety review. The current flagship you can actually call is still GPT-5.6 Sol, at $5 per million input tokens and $30 per million output tokens, exactly as it has been since 9 July.
If you searched for GPT-6, here is the short version: OpenAI has never announced a model by that name, and as of today it has not decided whether Astra will ship as GPT-6, as a GPT-5 point release like GPT-5.7, or as a fourth tier sitting alongside Sol, Terra and Luna. That naming ambiguity is not press-shyness. It is a live internal question, reported by The Information on 31 July and not resolved since.
What follows separates three things that most coverage of this story runs together: what OpenAI itself stated, what credible reporting adds, and what is circulating with no source attached. The proofs are real artifacts you can compile. The 10-trillion-parameter figure is a number someone typed on the internet. And the "releases next week" report turned out to be exactly what it looked like — a single unverified flash claim, falsified by the calendar and superseded by OpenAI's own safety disclosure. Those deserve very different weight.
What OpenAI actually said — and the much longer list of what it didn't
The announcement did not arrive as a product launch. It arrived as a paper. OpenAI's post described results produced by "an internal version of Astra, our next major model," across problems that had, in its framing, seen no progress on the main result for at least a decade. That "no progress for a decade" framing has since been revised after mathematicians pointed at prior published work on several of the results — more on that below. Astra was characterized as a system built for long-running work: multiple agents coordinating on different parts of one large problem over extended stretches rather than answering in a single pass.
That is close to the entire official technical description. Set against it, the list of absences is striking:
• Confirmed by OpenAI — the name Astra; that it is the "next major model"; that an internal version produced ten results; the multi-agent, long-horizon design intent; a 249-page manuscript; Lean 4 certificates on a public repository; a token-cost estimate quoted at GPT-5.6 Sol rates.
• Not stated by OpenAI — release date; price; context window; parameter count; any benchmark score; the model card; whether it reaches ChatGPT or only the API; how many agents ran, for how long, on what hardware; the prompts; the success rate; the final product name.
Credible reporting fills in a little. The Information reported that Sam Altman demonstrated Astra to policymakers in Washington in late July — a briefing that, per that reporting, included Senators Raphael Warnock, Bernie Moreno and Mark Warner, with Treasury Secretary Scott Bessent and Commerce Secretary Howard Lutnick also present. The model family is described as already in testing, with "Astra" itself still a tentative label.
Nobody outside OpenAI has run this model on anything. Every capability claim in circulation traces back either to the ten published proofs or to a demo the public did not see.
The ten proofs: unusually checkable, and still not settled

This is where the announcement is genuinely stronger than a typical benchmark drop, and it is worth being precise about why. OpenAI did not publish a score and ask to be believed. It published machine-checkable artifacts under Apache 2.0 in the openai/ten-proofs repository — one Lean file per result, pinned to Lean 4.32.0 with a mathlib dependency, alongside a ComparatorChallenges directory for independent checking. The repository landed as a single commit on 1 August and has since picked up several hundred stars and dozens of forks.
The ten results, as the repository names them, span an unusually wide spread of fields:
• Group theory — the construction of a non-sofic group, answering a question open since 1999 in the negative. Soficity is a property nearly every group mathematicians work with satisfies; whether every countable discrete group must satisfy it had stood for 27 years.
• Operator algebras — a counterexample bearing on the Connes rigidity conjecture, posed by Fields medalist Alain Connes in 1980.
• High-dimensional geometry — an improvement to the sphere-packing method, reported as the first of its kind since 1978 — plus a sharp form of Ehrhart's volume inequality.
• Coding theory — new bounds for binary and spherical codes.
• Complexity — an arithmetic circuit lower bound for the permanent, and an exponential parallel-repetition result for entangled games.
• Lattice cryptography — fixed-polynomial hardness for the closest vector problem.
• Extremal combinatorics — the growth scale of multicolor triangle Ramsey numbers, and counterexamples in extremal graph theory, including work on several catalogued Erdős problems.
Reactions from mathematicians were interested rather than dismissive. Thomas Bloom, who maintains the Erdős problems database, called the results big news and rated them above the unit-distance conjecture counterexample from May. OpenAI's Noam Brown offered the useful deflation from the inside: "Sadly, no Millennium Prize Problems (yet)."
In the days since, the announcement stopped being a research note and became a news story and then a dispute. Mainstream coverage — New Scientist, Scientific American and The Quantum Insider among them — picked the ten results up as a potential "Fields Medal-level" moment, led by the non-sofic group construction, and the reception split in two directions that both matter for how you read the announcement.
The first is a challenge to attribution. In a Scientific American report, mathematicians including Steven Miller of Yeshiva University and Francesco Fournier-Facio of the University of Cambridge accused the announcement of research misconduct, arguing that several proofs leaned on preexisting published ideas without proper citation — the sphere-packing improvement on a 2016 paper by Miller and a collaborator, the non-sofic group construction on 2016 and 2019 work by Andreas Thom and Gábor Kun. OpenAI revised its "no progress for a decade" framing in response. These are named-source objections from working mathematicians, and they land precisely on the gap Lean cannot check: whether the results are new in the way the headlines claim.
The second is a challenge to uniqueness. Anthropic researcher Levent Alpöge said the publicly available Claude Fable 5 had independently reproduced five of the ten results under "clean" conditions — fully autonomous, no internet — within about a day of the proofs appearing. That claim is itself unreplicated and not yet fully documented, and it has not been through review; treat it as a data point, not a verdict. But if it holds, "Astra crossed a threshold no other model could" becomes "Astra got there first" — which changes what the ten proofs establish about Astra's singularity.
What a Lean certificate settles, and what it leaves open
A Lean proof that type-checks is valid by construction. You do not have to trust OpenAI's evaluation methodology, its choice of baselines, or its interpretation of its own results — you can compile the files. For an industry where "we scored 94.1%" is the standard unit of evidence, that is a meaningful upgrade in falsifiability.
It is also narrower than the headlines imply, in four specific ways. Lean checks that the proof of a formal statement is sound. It does not check that the formal statement is the theorem the prose claims. It does not check that the encoded definitions match what the field means by those words. It does not check the informal reductions bridging the formal endpoints to the headline result. And it says nothing about novelty — whether a result is new, or already known under a different name.
All four of those are human-judgment questions, and none of them is answered by a type-checker — which is exactly why the attribution dispute that broke after publication is about provenance rather than validity. As of this writing none of the ten has been through refereed review. The manuscript acknowledges consultations with specialists; acknowledgments are not endorsements of every claim. One published partial build of the repository compiled 8,820 of 9,007 jobs, which is what a large real formalization looks like mid-verification rather than a finished audit. The honest status today is ten serious, formally encoded research claims — not ten settled entries in the mathematical canon.
There is a second, quieter limit: the discovery process is not reproducible by anyone outside OpenAI. The proofs are public; the search system, prompts, checkpoints and execution environment are not. You can verify the destination. You cannot rerun the journey.
The $2,000 figure, and why it buys less than it sounds like

The number that traveled furthest was the cost. OpenAI put the token spend behind the ten solutions at roughly $2,000 priced at GPT-5.6 Sol rates — about $200 per decade-old open problem, on the most common reading of its phrasing. Some coverage read the same sentence as under $2,000 per problem, a tenfold difference; OpenAI's post is terse enough that both readings survived into print, and we have not seen the company disambiguate it.
Take the total reading and work it out. Sol bills $5 per million input tokens and $30 per million output, with reasoning tokens billed as output. If the spend were entirely output, $2,000 buys at most about 67 million output tokens — roughly 6.7 million per problem. That is a lot of thinking, and it is not an absurd amount of thinking. It is the sort of budget a well-funded research group could already spend on a single hard question.
Three caveats matter more than the headline:
• It is a counterfactual price, not a price. Astra is unreleased and unpriced. The $2,000 describes what those tokens would cost at another model's rates. A frontier system built for hours-long multi-agent runs has no obvious reason to be priced like Sol, and every reason to be priced above it.
• It counts only the wins. No success rate was published. We do not know how many problems Astra was pointed at and failed to crack, or what those attempts cost. A cost-per-success published without a success rate is not a cost-per-result, and the difference could be an order of magnitude in either direction.
• The tokens are the cheap part. Humans framed the problems, prepared the manuscripts and drove the formalization. That labor is not in the $2,000.
• And the re-derivation price may already be falling. Alpöge's claim that a publicly available Claude Fable 5 reproduced five of the ten results autonomously is unreplicated — but if it holds, $2,000 is the first data point on a declining curve, not a floor.
The one part of this you can price today is the baseline. Because OrcaRouter passes provider list price straight through at 0% markup, GPT-5.6 Sol costs the same $5/$30 on our API as it does on OpenAI's — so if you want to sanity-check the arithmetic against your own workload, the rate in the announcement is the rate you would actually pay. What no one can price yet is Astra itself, and any vendor telling you otherwise is guessing.
The release date is gated by a 30-day clock, not by a leak

Most "when is GPT-6" coverage is rumor arithmetic. There is a more concrete constraint sitting in plain sight, and it has barely been connected to the question.
An executive order signed on 2 June 2026 directed federal agencies to stand up a voluntary review process under which frontier developers submit models for government review up to 30 days before public release. The framework was slated to take formal effect on 1 August — the same day the Astra post went up. Reporting indicates Astra is expected to be the first model through it.
"Voluntary" is doing some work in that sentence. In practice the administration has leaned on release timing before, and OpenAI's own recent behavior looks like a rehearsal: GPT-5.6 was held to roughly twenty vetted organizations from 26 June before broad access arrived on 9 July — a staged rollout of about two weeks.
The strongest "it's imminent" claim so far was published on 6 August: a flash report, citing a person it describes as a whistleblower, said the latest Astra build — codenamed "mewfour" — had been designated a release candidate and could ship as soon as this week, calling it the largest model OpenAI has trained since GPT-4.5. It was a single unverified claim with no named source at the company, no model card and no pricing page. Read against the process constraint, it was always indirectly checkable: a release that week required the 30-day pre-release submission to have already happened — so a model card or a pricing page appearing within days would have been the tell. Neither appeared, and the "next week" window has now closed without a release.
What superseded it was not another leak but the company itself. On 7 August, OpenAI disclosed that preliminary evaluations meant it "cannot rule out" that Astra reaches "critical" cyber capability under its Preparedness Framework — the threshold at which a model could autonomously identify and exploit serious real-world software vulnerabilities, or run complex coordinated attacks, without human intervention. The company paused the parts of Astra's development that do not meet its strengthened security requirements, moved the model into isolated test environments with restricted network access and sandboxed execution, and said it would test further with government agencies and select AI-safety organizations before release. CEO Sam Altman said OpenAI is still working to make Astra generally available — "we do not think it is a good strategy to keep powerful models to a chosen few" — but attached no date to that, and the review has no published end point.
Stack the process facts and you get a rough floor rather than a prediction. If Astra goes through the 30-day pre-release review and then repeats the GPT-5.6 pattern, broad availability lands something like six weeks after submission — and a possible "critical" cyber rating makes that review the binding constraint rather than a formality. The submission date is precisely what we do not know, which is why confident near-term dates should be discounted unless the review clock visibly starts.
The prediction markets have drifted in exactly that direction. Read on 13 August 2026, Polymarket's "OpenAI's Astra released by" ladder priced 3% for 15 August, 14% for 31 August, 52% for 15 September, 75% for 30 September and 93% for 31 October, on roughly $155,600 staked. The crowd has, in effect, already marked the near term closed — the two weeks the "mewfour" report pointed at were priced in single digits — and put the probability mass in the autumn. Treat that as a crowd's aggregated guess, not a source, but note what it implies: even after the safety pause, the market has not abandoned the model; it has only abandoned August.
It also helps to remember the track record. Every "GPT-6 launches next week" claim of the past twelve months has been wrong — including this one. An unverified leak pegged 14 April 2026 as launch day; what actually shipped, nine days later, was GPT-5.5. The 6 August "mewfour" report promised a release this week; the week is over, and it did not ship. On 10 August the anonymous X account Chris GPT escalated the same pattern, claiming GPT-6 would be "forcibly released" in August despite the federal review; 13 August has arrived and it has not shipped either. Specific enough to falsify quickly, and falsified on schedule.
The rumors, graded
Sorted by how much weight each actually carries:
• Naming undecided (GPT-6 / GPT-5.7 / a separate class). Sourced to The Information. The most reliable non-OpenAI claim in this story, and the one that should reset expectations — a system announced through a math paper may well not be branded as a generational jump at all.
• Release "as soon as next week" (build codename "mewfour" as release candidate). A single flash report from 6 August citing an unnamed whistleblower, with no named source at the company. It was falsified on schedule — the week it promised passed with no release — and OpenAI's 7 August disclosure of a possible "critical" cyber rating, which paused parts of development, makes the leak's timeline look optimistic rather than authoritative.
• Codenames "Zinc" and "Magnesium" spotted in Design Arena. Community sighting, entries reportedly disabled, unconfirmed by OpenAI. Read as: a GPT-5.x point release may well land before Astra, which would satisfy nobody's GPT-6 expectations while absorbing the news cycle.
• Roughly 2× GPT-5.6 Sol in scale; pricing near the GPT-5.6 range. Circulating rumor. The leaker behind it acknowledged not having tested the model. Plausible-sounding and entirely unsubstantiated.
• A 1.5-million-token context window. Appears in leak roundups with no named source. Worth noting Sol already ships 1,050,000 tokens, so this would be an increment, not a leap — which is a reason to be suspicious of it as a headline "leak."
• Approximately 10 trillion parameters. The claim now traces to the anonymous X account Chris GPT, which on 10 August reported a "forced" August release of a roughly 10-trillion-parameter model despite the federal review; several Chinese tech outlets carried the story, and it still has no named source at OpenAI. The most-repeated and least-supported number in the entire story.
• Memory and personalization as the defining direction. Grounded in Altman's own public remarks from an August 2025 interview — real, but a year old and about the post-GPT-5 direction generally, not about Astra specifically.
• A larger follow-up, codenamed "Doug", reportedly in pre-training. SemiAnalysis told institutional clients on 7 August that a "much larger" OpenAI model is "actively in the works," and Chris GPT, an anonymous X account, adds that it could arrive by year-end and would make Claude Fable 5 look "primitive." One named outlet and one anonymous account, with no OpenAI confirmation — worth tracking because it would reset the "Astra is the flagship" framing, not because it is verified.
What to actually do between now and whenever it ships
The wrong move is rearchitecting for a model that has no card. The right move is noticing which of your problems a long-horizon, multi-agent system would change, because those are the parts of your stack that are underbuilt today regardless of what OpenAI ships.
Three of them are worth building against GPT-5.6 Sol right now:
• Cost ceilings per task, not per call. If a single job can run for hours and spend six or seven million output tokens, per-request limits stop protecting you. You need a budget that a whole task inherits, and a hard stop.
• Checkpointing and resumability. A one-shot call either returns or fails. An hours-long agent run that dies at minute 90 with nothing durable written is a category of expensive failure most codebases have never had to handle.
• Evaluation you trust on problems with no reference answer. The ten proofs are interesting partly because Lean supplies an oracle. Almost nothing in production does. If you cannot tell a good six-hour run from a plausible-looking bad one, more compute will not help you.
On the switching question: Astra is not available on OrcaRouter, because it is not available anywhere. What we can say is that the version-churn problem is largely solved on our side. One key reaches 200+ models, so whether this ships as GPT-6, GPT-5.7 or a fourth GPT-5.6 tier, adopting it is a model-string change rather than a migration — no second contract, no new SDK. And for an unproven frontier model specifically, automatic failover is the difference between evaluating it in a real workload and betting a production path on a system nobody outside the lab has stress-tested. You route a slice of traffic to it, keep Sol underneath as the fallback, and find out — a setup worth having in place before the safety review clears, rather than after.
Questions worth answering
Is Astra the same thing as GPT-6?
Not necessarily, and that is the most consequential unknown in this story. OpenAI has confirmed Astra as its next major model but has not decided its release name; reporting places GPT-6, GPT-5.7 and a separate class alongside Sol, Terra and Luna all on the table. Nothing that has happened since — the "mewfour" flash report, and now the safety pause — changes that; even the flash report calls the product "Astra," not GPT-6. It is entirely possible that a model named GPT-6 never appears, and that the capability jump people have been waiting for arrives under a name nobody was tracking.
Can I access Astra in any form today?
No — not in ChatGPT, not through the API, not through any router or reseller. What exists publicly is the paper, the reasoning walkthroughs and the Lean repository. If a service claims to offer Astra access, it is either reselling something else or lying. The only genuinely useful thing you can do with the release right now is compile the proofs yourself.
Is Astra releasing any time soon?
The "next week" version has already failed — the week it named came and went with no release. The more consequential development is that OpenAI itself paused parts of Astra's development on 7 August after disclosing that it cannot rule out "critical" cyber capability, and said it will keep testing with government agencies and AI-safety organizations. Nothing about that review has a published end date. Polymarket now prices the near term in single digits and the autumn in the majority. Treat any specific date as a guess until a model card and a pricing page actually appear.
Did Astra really solve problems human mathematicians couldn't?
It produced formally verified proofs of statements that had been open for at least a decade, which is a real and unusual result — and the verification is a step above the industry norm. But the days since have made "solved" less settled than it looked on 1 August. Mathematicians quoted in Scientific American accused the release of research misconduct, arguing that several proofs leaned on preexisting published work without proper citation, and OpenAI walked back its "no progress for a decade" framing. Anthropic's Levent Alpöge, meanwhile, said a publicly available Claude Fable 5 independently reproduced five of the ten results autonomously within about a day — itself an unreplicated claim. "Lean-checked" is not "peer-reviewed," and both the novelty and the uniqueness of these results are now open questions. Expect the assessment to firm up over months, in either direction.
Does the $2,000 figure mean frontier research is now cheap?
It means the winning token runs were cheap, at another model's prices, excluding the failures and excluding the humans. Each of those three exclusions could be large. The honest read is that the marginal cost of a successful automated proof search has fallen far enough to be interesting, not that ten open problems now cost the price of a laptop — and if the claimed Claude Fable 5 replication holds, the marginal cost has already fallen further.
What would actually settle this
Three specific things, none of which are rumors, and all of which are checkable when they happen.
A model card and a pricing page. That is the moment Astra stops being a research artifact and becomes something you can plan against — and the moment the naming question resolves itself. The "next week" flash report failed exactly that test, which is how we know it was a leak and not a schedule. The model card is still the first real signal to wait for.
An independent rebuild of the Lean repository by specialists who audit the definitions and the informal reductions, not just the type-checking. The ComparatorChallenges directory suggests OpenAI expects this. If the formal statements hold up under that scrutiny, the results get much harder to discount; if a mismatch surfaces in even one of the ten, every claim in the set gets re-read. The attribution dispute reported by Scientific American makes this the highest-priority check.
Evidence that the federal review clock has started. Under a 30-day pre-release window, that submission is the earliest reliable signal of a ship date — considerably more informative than the next screenshot of a disabled entry in an arena leaderboard. The 7 August safety disclosure makes the review the live constraint: until OpenAI says a submission exists or a milestone has been reached, the timing is genuinely open.
Until then, the accurate statement is narrow and worth holding onto: OpenAI's next flagship has a name, ten machine-checkable proofs, a Washington demo, a possible "critical" cyber rating, and no ship date. Anyone offering you a date is guessing — the "mewfour" flash report was the latest guess, and it was wrong — and anyone offering you a parameter count is making it up.
