ChainThink reports that on July 30, AI agent infrastructure company Composio connected the same Kimi K3 model to Kimi Code, Hermes, and Claude Code to execute 28 identical tasks.
The number of completed operational frameworks are 22, 21, and 20, respectively, with minimal differences, but significant variations in token usage and costs.
Test results show that the median token counts for single tasks for Kimi Code, Hermes, and Claude Code are 61,000, 67,000, and 340,000, respectively;
Composio estimates that the average cost per task is approximately $0.22, $0.28, and $2. Individual tasks can vary by up to 30 times in token usage.
In terms of speed, the median time for a single task was 179 seconds for Hermes, 297 seconds for Kimi Code, and 348 seconds for Claude Code.
Another study by Writer also showed that, after replacing only the execution framework across 22 enterprise tasks and six models, token usage decreased by 38%, per-task cost dropped by 41%, processing time was reduced by 44%, and completion quality remained largely unchanged.
