theLLMs
Hero image for Google DeepMind Delays Gemini 3.5 Pro, Scraps Base Model as Key Talent Departs

TL;DR

Google DeepMind delayed Gemini 3.5 Pro’s general availability from June to at least July 2026 and scrapped its Gemini 2.5 Pro base model in favor of extended pre-training on a native Gemini 3 foundation. The decision followed internal evaluations showing the model could not sustain a meaningful performance gap over Gemini 3.5 Flash on core terminal tasks (76.2% on Terminal-Bench 2.1), making the premium tier at approximately $15/$60 per 1M tokens unjustifiable. The delay compounded a credibility crisis: Gemini co-lead Noam Shazeer and Nobel laureate John Jumper departed in the same week, while competitors Claude Fable 5 and GPT-5.6 moved into their launch windows. DeepMind is now betting on three targeted improvements — mathematical reasoning, SVG scene generation, and image quality — to reposition 3.5 Pro for the late-July flagship battle.

The Delay: Gemini 3.5 Pro Pushed to July 2026

Google DeepMind has delayed the general availability of Gemini 3.5 Pro from its originally announced June 2026 launch to July 17, 2026, while simultaneously scrapping the Gemini 2.5 Pro base model in favor of an extended pre-training cycle on a native Gemini 3 foundation [1,2]. The model was unveiled at Google I/O on May 19, 2026, where Sundar Pichai framed it as a “next-month” release. However, only Gemini 3.5 Flash shipped at the event, with Pro held in limited enterprise preview while Google promised full GA for June [2,3].

DeepMind’s publicly stated reason for the pullback is the need for “quality refinements after early enterprise testing” [2,3]. According to analyst trackers and internal reports, testers began flagging performance issues during the limited enterprise preview that expanded slightly in mid-June [2]. Rather than ship a model with known shortcomings, DeepMind elected to discard the initial 2.5 Pro base layer and run an extended, heavy-duty pre-training cycle on a more advanced Gemini 3 foundation [3].

The timing is notable. Google had committed to a June GA window, and while the company has not confirmed a specific date beyond “July or later,” analyst tracking places the target at “late June and early July” with the official rollout set for July 17 [1,2,3]. Gemini 3.5 Flash remains widely available as the usable tier, anchoring high-volume agent pipelines at $1.50/$9.00 per million tokens, while Pro stays in limited enterprise preview [1,2,3].

The Pro-to-Flash Paradox: Why the Base Model Was Abandoned

The core catalyst for the delay and model cancellation is what analysts have termed the “Pro-to-Flash paradox” — a structural problem that emerged when Google’s own lighter model outperformed its flagship in unexpected ways [1,3].

When Google released Gemini 3.5 Flash, it surprised the developer ecosystem by scoring higher than the older Gemini 3.1 Pro on core terminal tasks, hitting 76.2% on Terminal-Bench 2.1 at a fraction of the operating cost [1,3]. This created an immediate internal crisis: the upcoming 3.5 Pro build, if deployed on the older framework, would not offer a wide enough performance delta over its own low-cost Flash tier to justify premium enterprise token pricing, rumored at approximately $15/$60 per 1M input/output tokens [2,3].

Leaked internal evaluations revealed a more concerning picture. The scrapped base model struggled significantly under complex, recursive tool-calling environments. While it handled standard text processing efficiently, it failed to maintain structural consistency when generating complex, multi-layered layouts and mathematical reasoning steps — areas where competing models have achieved high stability [1,3]. DeepMind chose to swallow a near-term PR delay rather than release a model that would look vulnerable on arrival [1,3].

Targeted Improvements: Math, SVG, and Image Generation

The extended pre-training cycle is narrowly focused on three critical capability areas where the scrapped model revealed significant ceilings: mathematical reasoning, SVG scene generation, and image generation quality [1,2].

Multi-step mathematical reasoning was identified as a major performance ceiling in the abandoned build [1,3]. The model could handle standard operations but faltered on complex, chained mathematical problems that require maintaining structural consistency across multiple reasoning steps [1,3]. This is a competitive gap: mathematical reasoning is one of the core differentiators between flagship and mid-tier models, and it directly impacts utility in code generation, scientific computation, and data analysis workflows.

SVG scene generation — producing intricate and accurate vector-based visual representations from text prompts — is a new DeepMind focus area for this release [1,2]. This capability would extend Gemini’s multimodal strengths into structured visual output, a space where competitors have been gaining ground with their own generative imaging pipelines.

These targeted enhancements position Gemini 3.5 Pro specifically to compete with OpenAI’s GPT-5.6 reasoning modules and Anthropic’s Fable 5 autonomous workflows [1,2,3]. The strategy is clear: rather than shipping a broad but shallow upgrade, DeepMind is doubling down on the areas where the gap was widest and the competitive threat is most acute [3].

Talent Exodus: Shazeer and Jumper Departures

Compounding the technical delay is a credibility crisis sparked by the departure of two of DeepMind’s most prominent researchers in the same week the delay surfaced.

Noam Shazeer, Gemini co-lead and original co-author of the Transformer architecture that underpins much of modern AI, announced his departure for OpenAI on June 18, 2026 [2,3]. Just one day later, on June 19, Nobel laureate John Jumper — co-creator of AlphaFold and one of DeepMind’s most public-facing scientists — announced he was leaving for Anthropic [2,3]. The simultaneity of these departures, both senior and both connected to Google’s core AI identity (Transformer + AlphaFold), created immediate narrative damage [2].

Demis Hassabis publicly addressed the situation on June 23 in an interview with Semafor, saying DeepMind is “still winning” the AI talent war [2,3]. Analysts interpret this as an implicit acknowledgment that the perception had indeed shifted — Google can argue with the facts (DeepMind has roughly 4,000 researchers, and two departures don’t change the organization meaningfully) but the perception affects recruiting, internal morale, and external partner confidence [2].

The timing magnified the impact. If Gemini 3.5 Pro had shipped on time and benchmarked well, the talent story would have been a footnote. With the delay, it became a defining narrative paragraph [2].

Competitive Landscape: GPT-5.6 and Fable 5 Move In

A one-month slip in the frontier model market is not trivial — competitors have already moved into their launch windows, narrowing Google’s advantage window [2,4].

Claude Fable 5 shipped on June 9, 2026, with credit-window pricing and benchmarked model cards already exported to the public [2,4]. It has been independently benchmarked and is available for enterprise workloads. OpenAI’s GPT-5.6 Sol is in its own launch window (rumored June 22-28 per Polymarket predictions), with enhancements in speed, accuracy, and ethical safeguards [2,3,4]. Both models are positioned as flagship reasoning models with very long context windows, directly overlapping with Gemini 3.5 Pro’s intended positioning [2].

Anthropic’s Claude Opus 4.7 also targets a mid-July release, creating a three-way standoff for the July flagship slot [4]. Meanwhile, Google’s own Gemini 4 Flash and Nano Banana Pro are being developed as complementary offerings targeting specific market segments [1].

The competitive pressure is the underlying reason for the delay. DeepMind recognized that shipping a model with known performance ceilings against competitors that had already launched would cede market narrative and enterprise mindshare [1,3]. The question now is whether a late-July release can re-establish momentum in a market that has already moved on.

Confirmed Specs and What to Watch

From Google’s I/O announcement and limited-preview disclosures, the following specifications have been confirmed for Gemini 3.5 Pro (pending GA validation):

  • Context window: 2M tokens (up from 1M for Gemini 2.5 Pro) [2,4]
  • Deep Think reasoning mode: An advanced reasoning capability for complex problem-solving workflows [2]
  • Pricing: Approximately $15/$60 per 1M input/output tokens (announced, subject to GA confirmation) [2,4]
  • Gemini 3.5 Flash: Widely available at $1.50/$9.00 per million tokens, anchoring high-volume agent pipelines [1,2,4]

No firm public ship date has been confirmed for the new July window. Google has not committed to a specific date beyond “July or later” [2]. Enterprises that committed to “Gemini 3.5 Pro by June” as a workload anchor now need to re-plan: Flash is the near-term bridge for most use cases, while Pro remains the choice for complex reasoning, very long context workloads, and Deep Think workflows [2].

The default assumption should be “late July GA with a model card 1-2 weeks before release” [2]. A model card release would be the clearest signal that GA is imminent. Until then, the July 17 date cited in analyst trackers and media reports should be treated as a target, not a commitment.

Conclusion

Google DeepMind’s decision to delay Gemini 3.5 Pro and scrap its base model is a strategic bet on quality over schedule — but it comes at a significant cost. The Pro-to-Flash paradox exposed a structural pricing challenge that may reshape how flagship models are positioned in the market. The simultaneous departures of Shazeer and Jumper added a credibility layer to an already fragile situation, while competitors have already moved into their launch windows. Whether DeepMind’s targeted improvements in mathematical reasoning, SVG generation, and image quality are sufficient to regain momentum remains to be seen. The late-July flagship battle will be the true test.

Methodology

  • Data checked: 2026-07-18
  • Sources consulted: Geeky Gadgets, Andrew.ooo, HackerNoon, CryptoBriefing
  • Assumptions: Internal evaluations and leaked benchmarks are accurate; analyst trackers correctly identified the July 17 target date; talent departure announcements are confirmed.
  • Limitations: This article relies on publicly available information and leaked internal reports. Actual model performance and release timelines may differ. Pricing is pre-GA and subject to change.
  • Jurisdiction: Global.

Source list

Trust Stack

  • Last substantive check: 2026-07-18
  • Corrections policy: If you spot an error, contact us via the Contact page
  • Affiliation: theLLMs has no vendor affiliation, sponsorship, or commercial relationship with any AI provider mentioned

Change log

  • 2026-07-18: first published