Grok 4.5 has scored 91.3% on the VulcanBench coding benchmark, outperforming both Claude Fable 5 and GPT-5.6 Sol on real-world programming tasks. The model also delivered those results at a lower per-task cost than its competitors, according to benchmark data released this week.
What VulcanBench measures
VulcanBench is a coding benchmark designed to test models on practical, real-world software engineering challenges. Unlike synthetic tests, it evaluates how well an AI can handle actual codebases, bug fixes, and feature implementations. Grok 4.5’s 91.3% score marks a clear lead over the next-best models.
Cost advantage stands out
Beyond raw accuracy, the benchmark results highlight Grok 4.5’s efficiency. The model achieved its top score while keeping per-task costs lower than Claude Fable 5 and GPT-5.6 Sol. That combination of high performance and lower expense could shift how developers and companies choose which AI to use for coding work.
How the rivals compare
Claude Fable 5 and GPT-5.6 Sol both trailed Grok 4.5 on the VulcanBench evaluation. The exact scores for those models were not disclosed in the data, but the ranking is clear: Grok 4.5 came out ahead on real-world coding tasks. The cost difference adds another dimension to the competition, especially for teams that run large numbers of queries.
The results arrive as AI labs continue to push for better performance on practical benchmarks. VulcanBench’s focus on real code rather than abstract puzzles makes it a useful gauge for developers. Grok 4.5’s lead suggests its training approach may have advantages in handling the messy, context-heavy problems that come up in everyday software work.
Whether the other models can close the gap in future updates remains an open question. For now, Grok 4.5 holds the top spot on VulcanBench — and does so without the usual premium price tag.



