xAI’s Grok 4.5 tops VulcanBench with 91.3% score

The model solved 21 of 23 multi-file coding tasks across five programming languages, outperforming Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 Sol on the new benchmark.

Summary

verifying reliability

Terms & Concepts

No specialized terms available for this topic.