ChainThink reports that on August 1, AI agent infrastructure company Composio expanded its testing by integrating the same Kimi K3 model into six different agent frameworks, completing 26 identical tasks.
The comparison was conducted on the execution framework, with all underlying models being Kimi K3. Test results show that Kimi Code completed 21 tasks, ranking first; Hermes completed 20 tasks;
Pi Agent and Claude Code each completed 19 items; OpenCode completed 18; Codex completed 17.
According to the official pricing of Kimi K3, Hermes and Pi Agent have the lowest average cost per task at $0.39 and $0.40, respectively;
Claude Code costs $1.47, approximately 3.8 times that of Hermes. In terms of speed, the median time for a single task by Pi Agent is 161.7 seconds, while Claude Code takes 347.6 seconds.
Testing showed that, under the same model, switching frameworks resulted in a maximum difference of four successful tasks, an average cost difference of nearly four times, and a time difference of more than double.
