ChainThink reports that on July 25, according to third-party evaluation firm Artificial Analysis, Claude Opus 5 scored 61 points in an intelligence index combining nine tests, narrowly outperforming Fable 5's 60 points.
GPT-5.6 Sol scored 59 points, and Kimi K3 scored 57 points. In terms of cost, Opus 5 averages $2.03 per task, which is 26% lower than Fable 5’s $2.75.
The model ranks first in both the GDPval-AA v2 and AA-Briefcase knowledge work evaluations, and ties for first place in the Programming Agent Index when paired with Claude Code.
Terminal-Bench v2.1 scored 89%, roughly matching GPT-5.6 Sol. Opus 5 offers five levels of reasoning intensity.
From low to max, the output tokens differ by approximately 8 times, and the GDPval-AA v2 score differs by 407 Elo, allowing users to balance performance and cost.
The evaluation also shows that Opus 5 still lags behind Fable 5 in factual knowledge; its hallucination rate increased to 50% on the AA-Omniscience test, a 14-percentage-point rise compared to Opus 4.8;
The cost-effectiveness of the low推理 tier is still slightly lower than that of the GPT-5.6 series.
