TL;DR
Grok 4.5, released by xAI (now part of SpaceXAI) on July 8, 2026, is a coding-first flagship model trained jointly with Cursor using reinforcement learning across hundreds of thousands of developer-agent tasks. It achieves Opus-class benchmark scores at roughly a quarter of the output token cost — resolving SWE-Bench Pro tasks with ~14,000 tokens on average versus ~67,000 for Opus 4.8 — while running at $2/MTok input and $6/MTok output. The trade-off is a narrower feature set: the context window shrank to 500K, video input is gone, reasoning cannot be disabled, and the model was unavailable in the EU at launch. With the Cursor acquisition ($60B in June 2026) now complete, Grok 4.5 represents the first product of that union and signals xAI’s pivot from general reasoning toward specialized agentic coding workloads.
Introducing Grok 4.5 — xAI’s New Coding-First Flagship
Grok 4.5 was released on July 8, 2026, initially via Grok Build and Cursor, with public availability launching July 9. Elon Musk described the model as “an Opus-class model, but faster, more token-efficient and lower cost,” benchmarked against Claude Opus 4.7 [1]. This marks the first flagship model under the new “SpaceXAI” branding, following SpaceX’s acquisition of xAI and Cursor in 2026 [2].
The model is positioned squarely at coding, agentic tasks, and knowledge work — a deliberate pivot from general reasoning toward developer-specific workloads. It was trained jointly with Cursor using reinforcement learning across hundreds of thousands of coding and agentic tasks, giving it a unique edge in practical software engineering [3].
The Cursor Partnership — Why This Model Is Different
SpaceX agreed to acquire Cursor for approximately $60 billion in June 2026, and Grok 4.5 is the first tangible product of that union [3]. The model is a jointly trained mixture-of-experts system built using real developer-agent data collected from Cursor’s platform — including debugging traces, tool interaction logs, and multi-step coding sessions from millions of developer workflows [4].
Grok 4.5 launched inside Cursor on all plan tiers simultaneously, deeply integrated into the editor’s native agent workflow [4]. Rather than aiming at general chat or broad knowledge tasks, the model targets repository-scale software engineering and long-running agent loops — positioning it as a direct competitor to Claude Code and OpenAI’s Codex for in-editor coding workflows [4].
Specs & Architecture — MoE, 500K Context, and What Changed
Grok 4.5 uses a mixture-of-experts architecture with approximately 1.5 trillion parameters, according to Elon Musk’s public reports (not officially confirmed by xAI) [4]. The model features a 500K-token context window, which represents a deliberate shrink from Grok 4.3’s 1M window — a trade-off that sacrifices breadth for sharper coding focus [4].
Notable architectural changes include:
- Native video input dropped — Grok 4.3 was xAI’s first model to accept video; this is a regression on multimodal breadth [4]
- Text-and-image input with text-only output
- Reasoning runs at “high” effort by default and cannot be disabled, which improves quality but adds latency on simple queries [4]
- Measured throughput of ~86.7 tokens/sec according to Artificial Analysis [4]
- API model string is
grok-4.5on an OpenAI-compatible endpoint [4]
Token Efficiency — The Real Cost Advantage
The most compelling advantage of Grok 4.5 is its output token efficiency. In SWE-Bench Pro tasks, Grok 4.5 resolves problems using 15,954 output tokens on average, compared to 67,020 for Opus 4.8 (maximum effort) — roughly 4.2x fewer tokens per task [4]. Independent testing by Artificial Analysis corroborates this: ~14,000 output tokens per task, approximately 60% fewer than Opus 4.8 [4].
With pricing at $2/MTok input and $6/MTok output (cached input at $0.50/MTok), Grok 4.5 was already competitively priced on a per-token basis. But when multiplied by its token efficiency, the actual cost per completed task can be dramatically lower than equally-ranked competitors [4]. On agentic runs where output tokens dominate the billing, this efficiency becomes the model’s clearest practical advantage.
Benchmark Performance — Where Grok 4.5 Stands
At launch, Grok 4.5 achieved the following benchmark results:
| Benchmark | Score | Verification |
|---|---|---|
| SWE-Bench Pro | 64.7% | Vendor harness (not independently verified at launch) [2] |
| Terminal-Bench 2.1 | 83.3% | Vendor-reported [4] |
| Artificial Analysis Intelligence Index | ~54 | Ranked 5th of 168 models [4] |
The model ranked 5th of 168 models on the Artificial Analysis Intelligence Index, behind Fable 5, GPT-5.6 Sol, Opus 4.8, and GPT-5.5 [4]. xAI published only four coding and agentic benchmarks at launch — all first-party, with no independent third-party verification [2]. Independent benchmark results are still emerging, and the model was 4th at launch before GPT-5.6 Sol released the following day [2].
The Trade-Offs — Smaller Window, No Video, Higher Latency
Grok 4.5 is not a universal upgrade over its predecessor. Several notable trade-offs limit its appeal as a general-purpose model:
- Context window dropped from 1M (Grok 4.3) to 500K — a trade for sharper coding focus, but still a regression [4]
- Video input removed — Grok 4.3 was xAI’s first model to accept video; this is a regression on multimodal breadth [4]
- Reasoning cannot be disabled — simple queries pay latency and token cost for “high” effort reasoning, which matters for high-volume API usage [4]
- Not available in the EU at launch — regulatory delays may persist [4]
- Specialist rather than general-purpose — these trade-offs make it a targeted tool for coding workloads, not a broad replacement [4]
Competitive Landscape — Against Claude, GPT, and Fable
xAI positioned Grok 4.5 against Claude Opus 4.7, but independent benchmarks place it slightly behind Opus 4.8 and Fable 5 on pure intelligence scores [4]. GPT-5.5 and GPT-5.6 Sol are ahead on intelligence benchmarks [2].
Where Grok 4.5 differentiates is cost-per-task: the combination of low $2/$6 pricing plus 4x token efficiency makes it the cheapest frontier option per completed coding task [4]. Combined with native Cursor integration, it offers a workflow advantage over Claude Code and Codex for in-editor coding workflows [4].
The model’s sweet spot is clear: developers who want frontier coding ability without per-task costs exceeding $0.10, particularly for repetitive agentic workflows where token count dominates the bill [4].
Looking Ahead — Monthly Model Cadence and What’s Next
xAI announced plans to ship a new foundation model roughly every month through the end of 2026 [2]. Grok Build (the developer tier) now includes Grok 4.5, Composer 2.5, and Grok Build 0.1 for agent coding [4].
Key developments to watch:
- Deeper IDE integration — The Cursor acquisition could unlock real-time agent workflows and repo-level understanding directly in the editor [4]
- EU availability — The key near-term gap; rollout depends on regulatory approvals [4]
- Competitive pressure — The aggressive cadence puts pressure on Anthropic, OpenAI, and Google to respond with their own coding-specialized models [4]
Conclusion
Grok 4.5 is xAI’s most focused flagship yet: an Opus-class coding model built for a specific use case with a specific workflow. Its token efficiency and aggressive pricing make it the cheapest frontier option per completed task, and the Cursor integration gives it a workflow advantage no other model can replicate at launch. The trade-offs — smaller context window, no video, forced reasoning, and EU unavailability — keep it from being a general-purpose upgrade, but for developers running high-volume coding agents, these limitations are secondary to the economics.
Whether xAI’s monthly cadence can be sustained without quality degradation remains the critical question. For now, Grok 4.5 delivers real cost advantages in a market where token efficiency increasingly matters as more workloads shift from chat to agentic coding.
Methodology
- Data checked: 2026-07-19
- Sources consulted: xAI official announcement, TechCrunch, The Air Rankings, Codersera
- Assumptions: xAI-reported parameter counts and benchmark scores are accurate; Artificial Analysis measurements are representative; pricing reflects public API rates at time of writing.
- Limitations: This guide covers benchmark data and pricing available as of launch. Independent SWE-Bench Pro verification is pending. EU availability status is subject to regulatory timelines.
- Jurisdiction: Global.
Source list
- xAI — https://x.ai/news/grok-4-5 (accessed 2026-07-19)
- TechCrunch — https://techcrunch.com/2026/07/08/spacexai-releases-grok-4-5-which-elon-describes-as-an-opus-class-model/ (accessed 2026-07-19)
- The Air Rankings — https://theairankings.com/xai/grok-4-5/ (accessed 2026-07-19)
- Codersera — https://codersera.com/blog/grok-4-5-launch-guide-2026/ (accessed 2026-07-19)
Trust Stack
- Last substantive check: 2026-07-19
- Corrections policy: If you spot an error, contact us via the Contact page
- Affiliation: theLLMs has no vendor affiliation, sponsorship, or commercial relationship with any AI provider mentioned
Related guides
Change log
- 2026-07-19: first published