Artificial Analysis’s latest comparison put Claude Opus 5 at 61, ahead of Fable 5 at 60 and GPT-5.6 Sol at 59, while highlighting cost, token-efficiency and reliability trade-offs across frontier AI models.
Artificial Analysis’s latest model comparison ranked Claude Opus 5 first with a score of 61 across nine tests, followed by Fable 5 at 60, GPT-5.6 Sol at 59 and Kimi K3 at 57. In the token-efficiency view, GPT-5.6 Sol generated about 15,000 output tokens per task at a cost of $1.04 per task, while Fable 5 produced about 33,000 tokens; a separate cost comparison in the newer benchmark showed Claude Opus 5 averaging $2.03 per task versus $2.75 for Fable 5. The results also showed OpenAI occupying most of the Pareto frontier, while Terra and Luna did not appear. Artificial Analysis’s breakdown further indicated that Claude Opus 5 led on GDPval-AA v2 and AA-Briefcase, but Fable 5 performed better in factual knowledge and Claude Opus 5 recorded a 50% hallucination rate in AA-Omniscience.