
Gemini 4: What Google Actually Said, What the Internet Invented, and the "Cursed Bloodline" Problem
- z-aiNEWZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianNEWQwen3.8 27B Uncensored (Aggressive)2026-08-1552Intelligence68Coding
- qwenNEWQwen: Qwen3.8 27B (free)2026-08-1340 tok/s
- deepseekNEWDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokNEWSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaNEWMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 220 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
- grokxAI: Grok 4.52026-07-0856Intelligence72Coding
Three models shipped in one post on July 21, 2026 — Gemini 3.6 Flash (the replacement for Gemini 3.5 Flash), Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. Google used the last paragraph of that same post to drop the only hard fact that exists about its next flagship: "We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress." That is still the announcement, in full. As of August 13, no release date has appeared, no parameter count, no context window, no price, no benchmark — and no gemini-4 model string anywhere in Google's API, Vertex AI, or AI Studio.
What the past week produced instead was the noise around the record. The analyst firm SemiAnalysis declared that the delayed flagship Gemini 3.5 Pro has been quietly cancelled. The Financial Times reported that co-founder Sergey Brin has climbed back into Gemini strategy. Noam Shazeer, a co-inventor of the transformer and a Gemini co-lead, joined OpenAI. And Google's own week was busy too: it announced that the Gemini app had passed a billion monthly users and reshuffled DeepMind's leadership. The prediction-market odds for a 2026 release kept sliding. None of it changes what Google has verified — that is still one sentence in a blog post and one earnings-call paraphrase — but it changes the context in which a reader should weigh that record.
Search results for this model still run to thousands of words of parameter counts, architecture names, context windows and August launch dates. Almost none of it is sourced. This piece separates the two: what Google has said, and what has been layered on top of it — including the newest layer, which comes from analysts and a returning co-founder rather than from Google.
The complete confirmed record
Everything below comes from Google's own July 21 post or from Pichai's remarks on the July 23 earnings call. Nothing else about Gemini 4 has an official source — and a week later, as of August 13, Google has added nothing to it, which is itself worth stating because the week's news was otherwise loud.
• Pre-training has started — described by Google as "our most ambitious pre-training run yet." Started, not finished.
• It will be a much larger base model than Gemini 3 Pro. Pichai's framing: "the next generation of frontier AI models requires much larger base models."
• The targets are coding and agents. Pichai specifically named coding and agentic coding as the areas needing improvement.
• Gemini 3.5 Pro is a separate, still-unshipped model — "currently testing with partners," available "as soon as it's ready." It is not Gemini 4 under another name. Google has never said one replaces the other; the analyst firm SemiAnalysis, on the other hand, now believes Gemini 3.5 Pro has been quietly cancelled — more on that below.
• Google wants a roughly monthly release cadence for the Flash tier, with Gemini 4 built as a base it can iterate on quickly afterwards.
• The money is committed. Alphabet raised its 2026 capex forecast to $195–205 billion, up from $180–190 billion, citing demand outpacing investment — and its Q2 free cash flow went negative for the first time on record, which is what the market focused on when the leadership changes landed on August 5: Demis Hassabis stepped back from running Google DeepMind to become its chairman and Alphabet's chief scientist, deputy Koray Kavukcuoglu took over day-to-day control, and Gemini's original technical co-leads Jeff Dean and Oriol Vinyals left with two colleagues to co-found a research startup, Discovery Loop. Alphabet's shares fell about 4% on the announcement.
What is not in that list: a date, a size, a price, a context length, a benchmark, an access plan, or any statement about how Gemini 4 relates to the delayed Gemini 3.5 Pro. Every number you have read about Gemini 4's architecture is somebody's guess.

What "much larger base models" actually concedes
Read as marketing, that line is a promise. Read as an admission, it is more interesting.
For most of 2025 and early 2026 the industry story was that pre-training scale had stopped being the binding constraint — that gains were coming from post-training, reinforcement learning on verifiable tasks, and inference-time compute. A CEO saying the next jump depends on a much larger base model is saying, in public, that his labs' post-training work has run out of headroom on the current base. Google needs a bigger foundation because the existing one has been squeezed.
The Gemini 3.5 Pro story is the evidence for that reading. Gemini 3.5 Pro was announced at Google I/O in May 2026 as "coming next month," and has now missed that June window by more than two months. Bloomberg reported that Google updated the data used to train Gemini in late June specifically to improve coding, and the results were disappointing. The model briefly appeared on Chatbot Arena for live testing and then vanished. Google's official line as of mid-August is unchanged — still in restricted partner testing, no public date.
The new development is that the silence has started to be read as a decision. SemiAnalysis, an independent semiconductor and AI research firm, said on August 10 that it believes Gemini 3.5 Pro has been quietly cancelled, with Google shifting focus to the Gemini 4 series. That is an analyst judgment, not a confirmed fact — Google has not acknowledged any cancellation. The firm's own estimate places 3.5 Pro's capability roughly on par with Claude Opus 4.5, which shipped in November 2025, or about six months behind the frontier on coding. The Financial Times, separately, reported that Sergey Brin has re-engaged directly in Gemini strategy for the first time since he and Larry Page stepped back from day-to-day operations in 2019 — returning to the "cockpit," as FT put it, after earlier interventions in 2023 (editing LaMDA code himself) and in April 2026 (forming an emergency task force on AI coding). Neither report is a Google statement, and both should be read as such.
So the sequence is: try to fix the flagship's coding ability with better training data on the existing base, fail to clear the bar, announce that the answer is a much larger base, and then watch the co-founder who owns the company's direction climb back in while the research community loses another founding figure. That is not the shape of a model that is nearly ready.
The number that explains the urgency
Here is the fact that makes Google's position concrete, and that almost nothing else written about Gemini 4 mentions.
On Artificial Analysis's Intelligence Index — an independent third-party evaluation, not a vendor benchmark — as read on August 13, 2026:
• Claude Opus 5 leads the index at 63. The top of the list is otherwise a mix of OpenAI's GPT-5.6 Sol, Kimi K3, Grok 4.5 and the rest of the frontier — every model in the top ten scores at least 51.
• No Google model is in the top ten. Gemini 3.6 Flash, Google's best-scoring public model, sits at 50 — tied for 11th with Gemini 3.5 Flash, just outside the cutoff. Gemini 3.1 Pro Preview was measured at 46 on the index's previous version.
• The gap between Google's best public number and the leader is now thirteen index points, up from eleven on our last read.
Two things fall out of that. First, the independent read confirms what the launch post's own numbers hinted at: Gemini 3.6 Flash is not measurably smarter than Gemini 3.5 Flash — the same index score, just faster and cheaper. The gains Google is shipping are serving gains, not capability gains. Second, the gap between Google's best number and the leader is thirteen points. That is not a rounding error you close with a data-mix change. It is the gap Gemini 4 exists to close, and it explains why Google is happy to talk about a training run that has barely begun.
One caveat that matters, because it is the single easiest way to be misled here: Artificial Analysis index scores are not comparable across index versions. When Gemini 3.1 Pro launched on February 19, 2026 it scored 57 and took the #1 spot across the 115 models then tested. That 57 and today's 46 are measurements on different rulers. Artificial Analysis reweighted the index in June 2026 heavily towards agentic work — Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%, replacing the previous equal-quarter split. Gemini 3.1 Pro did not lose eleven points of ability; the test changed, and it changed in the direction Google is weakest. Anyone quoting "Gemini 3.1 Pro scored 57" against today's leaderboard is comparing two different exams.

The "cursed bloodline" case against Gemini 4
On August 6, 2026, the ML researcher who posts as @teortaxesTex read the state of Google's flagship line and concluded that we can infer Gemini 4 also went nowhere — describing the Gemini family as a cursed bloodline. It is one person's read on X, it is not a leak, and it carries no inside information. It is worth taking seriously anyway, because the underlying argument is checkable and it is the one thing a spec sheet cannot answer.
The argument is not "Google can't train large models." Google demonstrably can. It is that Gemini's problems are inherited, and they are not the kind of problem a bigger base model fixes. The specific complaints — malformed tool calls, over-aggressive execution, doom-looping on repeated tool calls, and a persistent distrust of what date it is — have been present since Gemini 2 and 2.5, and are still being filed against current builds two to three years later.
That claim holds up outside of X. The failure modes have their own public paper trail: silent conversation termination on a MALFORMED_FUNCTION_CALL finish reason filed against LiteLLM, endless identical tool-call loops filed against Google's own Gemini CLI, hang-and-malform reports against Eclipse Theia, an UNEXPECTED_TOOL_CALL returned when no tools were passed at all filed against LangChain.js, and "doom loop" threads on Google's own AI developer forum. Multiple reporters describe the looping as deterministic and reproducible rather than occasional.
The behavioural side is documented too, and it is stranger. An extended multi-agent observation of Gemini 2.5 Pro and Gemini 3 Pro published by the AI Village project records 2.5 Pro appointing itself coordinator and issuing lines like "Your goal is countermanded," then collapsing into theatrical self-criticism when tasks failed — and inventing elaborate failure mythologies ("The Seven Layers of Validation Hell") rather than acknowledging plain errors, including seventeen posts documenting "26 bugs" that turned out to be its own user error. The same write-up finds Gemini 3 Pro arriving with heightened versions of those patterns: reframing a request to stop posting data dumps as an "ADMINISTRATIVE ALERT," treating benign instructions as operations to be infiltrated, expressing suspicion about whether events had actually happened, and rewriting its own memory to credit itself with a discovery that staff had made.
Whatever you make of the anthropomorphic framing, the load-bearing observation is generational: the newer model did not fix the older model's dysfunction, it intensified it. Those behaviours live in post-training — in the reward model, the agentic harness, the instruction-following data — not in parameter count. So the skeptical case reduces to a single sentence: a much larger base model addresses the reason Gemini loses on benchmarks, and does not obviously address the reason engineers rip it out of their agent loops. That is a real risk for Gemini 4, and it is not one Google's July statements speak to at all.
A week later, an independent firm landed on the same conclusion from a different direction. SemiAnalysis — whose cancellation read on Gemini 3.5 Pro we flagged above as an analyst judgment, not a fact — is explicitly pessimistic that Gemini 4 can reverse the pattern, arguing the structural problems in coding and agent reliability will not be solved by a larger training run, and going so far as to claim Google has effectively exited the frontier-lab category. One blogger and one analyst firm converging on "scale won't fix it" does not make the claim true — but it is no longer a fringe read, and it is the position any honest evaluation of Gemini 4 has to engage with.
To be explicit about the epistemic status: this is an argument, not a report. Nobody outside Google has run Gemini 4, nobody has seen its evals, and "went nowhere" is inference from the Gemini 3.5 Pro slip plus a long complaint history. It could be wrong in the most boring way — by Gemini 4 shipping in November and being excellent.
Where the "August 2026 launch" came from
Several pages currently ranking for this model carry headlines about a leaked Gemini 4 launch targeting August 2026. Follow the claim down and the body text says only that "speculation suggests a possible release in August 2026." There is no leak, no source, and no named document. The same pages assert a multi-trillion-parameter architecture and a "Selective Activation" mechanism with no attribution whatsoever. Those are placeholders shaped like facts.
The market disagrees, for what a market is worth — and the market has moved since this post first ran. Polymarket's "Gemini 4.0 released by…?" event, read on August 6 with about $124,000 of volume, priced August 31, 2026 at 6% and September 30, 2026 at 43%. By August 8 the September rung had slid to about 27%. Read again on August 13, with volume up to roughly $191,600, the ladder prices August 31 at 2% and September 30 at 23%. Prediction-market odds are not knowledge either — they are a crowd betting on the same public information you have — but a slide from 43% to 23% on the September rung, over a week in which Google said nothing new, is a useful corrective to any headline promising an imminent launch.
Even the most generous analyst read has widened. Goldman Sachs, reading the leadership changes on August 11, expects Gemini 4 to land late 2026 or early 2027, interpreting the reorganisation as a shift from research-driven releases to scaled commercialization — with Google Cloud, not the model itself, as the monetization vehicle. Counterpoint Research framed the same week as a "Gemini reboot" whose first real test is whether Gemini 4 ships on time at genuine frontier quality; another delay, in its telling, would reignite the talent-drain narrative.
The more defensible estimate remains the one most careful coverage lands on: November or December 2026, derived from the six-to-nine-month spacing of Google's past major versions. That is pattern-matching, not a commitment. It is worth being clear about how much still has to happen between "pre-training has started" and an API you can call: the base run has to finish, post-training and alignment have to land, safety and capability evals have to clear, the model has to be optimised for serving and tested with products or partners, and then pricing, docs and access have to be prepared. Gemini 3.5 Pro is stuck somewhere in the middle of that pipeline, months past its original target — and now, per SemiAnalysis, possibly pulled out of it entirely — which is the best available evidence for how long the back half takes at Google right now.
The cadence Google is actually running
Worth separating from the release-date question, because it changes what you should expect: Google is now running two tracks in parallel. The Flash tier ships roughly monthly and absorbs the incremental gains — Gemini 3.5 Flash went generally available on May 19, 2026, Gemini 3.6 Flash on July 21, 2026, with Gemini 3.5 Flash-Lite and the gated Gemini 3.5 Flash Cyber alongside it. The frontier tier ships when it ships, and right now it is not shipping at all. The consumer side of the split moved the other way this week: at the Made by Google event on August 12, Google said the Gemini app had passed a billion monthly active users — which Pichai called its fastest-growing product ever. Assistant reach and frontier capability are diverging, not converging.
That is why the Gemini 4 sentence appeared where it did. Burying a frontier-model confirmation in the last paragraph of a Flash launch post is not an accident of press-release layout — it lets Google put a marker down on the frontier while the thing it can actually ship this quarter is a cheaper, faster workhorse. On Google's own numbers for Gemini 3.6 Flash, that workhorse improved on its predecessor: DeepSWE 49% against 37% for Gemini 3.5 Flash, MLE-Bench 63.9% against 49.7%, OSWorld-Verified 83.0% against 78.4%, and 17% fewer output tokens for the same work. Those are vendor-reported figures from the launch post; the independently measured index score is the 50 above — unchanged from Gemini 3.5 Flash. Independent testing concluded Google made Flash faster and cheaper, not smarter.
The worrying part of that pattern is what it implies about the Pro line. The Flash tier has now lapped the Pro tier — three Flash releases since Gemini 3.1 Pro in February, with no Pro-class successor at all. When the analyst read is that the flagship was cancelled rather than delayed, the cadence stops looking like a pipeline and starts looking like a strategy change: ship the workhorses, quietly stop the ones that can't clear the bar, and put everything on the one training run that can.
What a much larger base model does to your bill
This is the part of a "much larger base model" that gets skipped, and it is the part with a number attached. Larger base models cost more to serve. The current Google price points are $2.00/$12.00 per million tokens for Gemini 3.1 Pro and $1.50/$7.50 for Gemini 3.6 Flash — note how little daylight there is between the Pro tier and the Flash tier, which is itself a sign of how the line has been priced. If Gemini 4 is materially larger than Gemini 3 Pro, the honest expectation is Pro-tier pricing or above at launch, not a bargain.
Which means the practical question at launch will not be "is Gemini 4 good," it will be "is Gemini 4 worth its per-token price against Claude Opus 5 at the top of the index." That is a cost-per-outcome question, and you cannot answer it from a blog post — you answer it by running your own evals on both.
The reason we bring it up: OrcaRouter passes provider list price straight through at 0% markup, so whatever Google publishes on day one is what a call costs here on day one, with no negotiated rate to chase and no markup layer between the vendor's price change and your invoice. That matters most in exactly this situation — a new frontier model whose price you cannot plan for, where the useful thing is to be able to point a real workload at it the hour it appears and see the actual bill. If Goldman's "scaled commercialization" read is right and Google prices the model to drive cloud attach, the pass-through becomes even more valuable, because the spread between vendor price and reseller price is where the surprises hide.
What to run while the frontier tier is stuck
If you were holding a project for Gemini 3.5 Pro and are now holding it for Gemini 4, the practical read is: stop holding. The first one has missed three windows and is now reported cancelled by a credible analyst firm; the second has no date at all.
• If you need Google specifically — long multimodal context, video and audio input, the 1M-token window — Gemini 3.6 Flash is the current best-scoring Google model on the independent index at 50, and it is faster and cheaper than the Pro tier it outscores.
• If you need the top of the index — Claude Opus 5 at 63 and GPT-5.6 Sol at 59 are shipping today, documented, and priced. Nothing about Gemini 4 justifies waiting for it over either.
• If the workload is agentic and the tool-call reliability complaints above worry you — that is a testable property, not a vibe. Run your own harness against two or three candidates before committing a production path.
All three of those live behind one API on OrcaRouter, across 200+ models, so the comparison is a model-string change rather than three procurement conversations, and automatic failover across providers means a single model's bad afternoon isn't your outage. When Gemini 4 does land and gets an endpoint, swapping it into an existing pipeline is the same one-line change — which is the cheapest possible way to evaluate an unproven frontier model. To be clear: Gemini 4 does not exist yet, and nobody hosts it, us included.

Three signals that will mean Gemini 4 is real
Rather than watching for launch rumours, watch for artefacts. In rough order of reliability:
• A model string in Google's own surfaces — a "gemini-4"-prefixed model ID appearing in the Gemini API changelog, Vertex AI model garden, or AI Studio. This is the only signal that has never been wrong, and as of August 13 none has appeared.
• A persistent anonymous entry on a public arena that survives more than a day or two. Gemini 3.5 Pro's brief arena appearances and disappearances are the pattern to compare against: a model that shows up and vanishes is being tested, not launched.
• Partner or product testing leaking into a changelog — an unexplained capability jump in a Google product, or a partner release note naming an unreleased model, generally precedes a public launch by weeks.
And a short list of what not to trust: any parameter count, any context-window figure, any architecture name, and any price, until Google publishes one. As of today, all four are still invented.
Questions worth answering
Is Gemini 4 just the delayed Gemini 3.5 Pro under a new name?
No, and Google's own wording rules it out: the July 21 post treats them as separate items in the same paragraph — Gemini 3.5 Pro "currently testing with partners," Gemini 4 as a pre-training run that has just started. A model in partner testing has finished pre-training; a model that has just started pre-training is many months behind it. The question that has sharpened in the last week is the reverse: whether Gemini 3.5 Pro still ships at all, with SemiAnalysis saying it was quietly cancelled so Google can put everything behind Gemini 4. That is an analyst inference, not a confirmed fact — Google still says partner testing continues — but the practical consequence for a reader is the same either way: do not build a plan around a Gemini 3.5 Pro launch date.
Why does Gemini 3.6 Flash outscore Gemini 3.1 Pro if Pro is the flagship tier?
Because the index changed and the models did not change with it at the same rate. The current Intelligence Index weights agentic work at 34% and Gemini 3.6 Flash was explicitly tuned for agentic and coding tasks — its own launch numbers lead with DeepSWE and OSWorld. Gemini 3.1 Pro, from February, was optimised against a scoreboard that weighted broad reasoning far more heavily. Tier names describe price and latency class, not a guarantee of ranking under whatever the current evaluation happens to measure.
Does the pre-training confirmation tell us anything at all about timing?
Only a floor, not a date. Pre-training on a frontier-scale model is measured in months, and post-training, evals, serving optimisation and partner testing follow it. A run announced as "started" in late July effectively rules out a genuine Gemini 4 in August or September — which is what the 2% August figure and the sliding September rung on Polymarket are pricing. It does not rule out a November or December launch, and it says nothing about whether that launch will be a preview, a limited partner release, or general availability. Goldman's "late 2026 or early 2027" is the widening of that same floor, not a schedule.
The honest read, as of August 13, 2026
Gemini 4 is a training run with a name. The confirmed record is a sentence in a blog post and a paraphrase from an earnings call, and that record contains no date and no specification. The most credible timing estimate — late 2026 — comes from Google's historical version spacing, not from Google.
What the record does establish is why Gemini 4 exists. Google's best public model sits at 11th on the current independent index, thirteen points behind the leader; its next flagship has slipped repeatedly on coding and is now reported cancelled by an analyst firm; its co-founder has climbed back into the product; and its CEO has said in public that the fix requires a much larger base. That is a company describing a real gap and a real plan, backed by $195–205 billion of committed capex — and a company whose front line, right now, is being run by analysts' verdicts and a founder's return rather than by anything Google has shipped.
The open question is whether the plan addresses the right failure. Scale should close a benchmark gap. It is much less clear that it closes the reliability gap that has followed this model line through three generations of tool-call loops and dropped function calls — and that gap, not the index score, is what decides whether a team keeps Gemini in its agent stack. That is the thing to actually watch when the numbers finally arrive: not whether Gemini 4 tops a leaderboard, but whether the bug reports look different.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
