ChainThink reports that on July 22, Artificial Analysis updated the AA-Briefcase rankings, with Kimi K3 scoring 1543 Elo, just below Claude Fable 5’s 1574 and above GPT-5.6 Sol’s 1501, Claude Sonnet 5, and Claude Opus 4.8.
AA-Briefcase evaluation requires the model to locate information from nearly 2,000 emails, Slack messages, and company documents, and deliver tables, presentations, and interface prototypes, encompassing 4 long-term projects and 91 confidential tasks.
In detail, the objective pass rate for K3 is 51%, lower than Fable 5’s 56%; the analysis quality score is 1754, slightly higher than Fable 5’s 1744.
K3's main gap lies in final output quality, falling below GPT-5.6 Sol and Opus 4.8. In terms of cost, K3 averages $10.57 per task, approximately ten times that of Kimi K2.6;
Average of 83 rounds executed, outputting 120,000 tokens, taking 56.4 minutes—approximately 2.5 times longer than Fable 5 for similar tasks.
