The model solved 21 of 23 multi-file coding tasks across five programming languages, outperforming Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 Sol on the new benchmark.
37d ago
verifying reliability
No specialized terms available for this topic.