Claude Opus 5 tops Artificial Analysis benchmark with 61 score

The model ranked ahead of Fable 5, GPT-5.6 Sol and Kimi K3 across nine tests, while posting lower average task costs than Fable 5 but weaker factual accuracy in one measure.

Summary

verifying reliability

Terms & Concepts

No specialized terms available for this topic.