SpaceXAI launches Grok 4.5 with lower pricing and mid-July EU rollout

New model launched July 8 now claims the top spot on SWE Marathon with a 29.0% resolution rate, while xAI emphasizes fast, low-cost Opus-class coding performance pending independent verification.

Summary

SpaceXAI released Grok 4.5 on July 8, pricing it at $2 per million input tokens and $6 per million output tokens, and is now highlighting self-reported benchmark leadership on SWE Marathon, where the model posted a 29.0% resolution rate ahead of Anthropic’s Claude Opus 4.8 at 26.0% and Fable at 24.0%. The company says Grok 4.5 was trained on tens of thousands of NVIDIA GB300 GPUs with reinforcement learning tuned for software engineering tasks, runs at roughly 80 transactions per second, and can deliver up to 4.2 times better token efficiency than leading competitors in specific applications. The model launched through the Grok app, xAI console and API, with EU availability expected by mid-July. Earlier reporting also showed mixed coding results on other tests and stronger independent support on agentic workloads: Artificial Analysis ranked Grok 4.5 first on AutomationBench-AA at 51.4% and $0.34 per task, while SpaceXAI’s own launch data showed 53% on DeepSWE 1.1 and 64.7% on SWE Bench Pro. The latest SWE Marathon result has not yet been independently verified, leaving confirmation of xAI’s benchmark lead as a key point to watch.

Terms & Concepts
  • SWE Marathon: A benchmark that measures how well AI models resolve long-horizon software engineering tasks.
  • reinforcement learning: A training method that improves a model by rewarding behaviors that produce desired outcomes.
  • AutomationBench-AA: Artificial Analysis’s independent benchmark for testing how AI agents complete tasks across simulated applications while following operating limits.