ChainThink reports that on August 5, security firm Aikido tested seven models using its code audit agent, including Qwen3.8-Max, Claude Opus 5, Kimi K3 Max, DeepSeek V4 Flash, GPT-5.6 Sol, Luna, and Terra.
The test included 32 recently disclosed vulnerabilities, with each model run three times. Qwen3.8-Max identified a total of 26 vulnerabilities across all three rounds, achieving a recall rate of 81.3%, tying for first place with Claude Opus 5.
Qwen3.8-Max has an F1 score of 83.2%, slightly lower than Opus 5, Kimi K3 Max, and GPT-5.6 Sol.
Its single-round performance stability is weaker, with only 10 out of 26 vulnerabilities found in all three consecutive rounds, while both Opus 5 and Sol found 19.
In terms of cost, Qwen3.8-Max has a total cost of approximately $821, about half that of Opus 5 and Sol, but five times that of DeepSeek V4 Flash.
