ChainThink reports that, according to testing by Artificial Analysis on September 11, DeepSeek V4.1 Flash achieved a comprehensive intelligence score of 40, surpassing its own V4 Pro 0813 score of 36, but still below Kimi K3’s 44 and GLM-5.3’s 45.
According to this third-party ranking, DeepSeek has not yet reclaimed the top spot among domestic models. There is a clear gap in cost and speed.
V4.1 Flash has an average cost of $0.27 per Intelligence Index task, while Kimi K3 and GLM-5.3 both cost around $2, a difference of approximately 7 times.
The output speed is 197 tokens/s, significantly higher than the approximate 36 tokens/s and 58 tokens/s of the two. Agent performance remains a strength.
On AutomationBench-AA, V4.1 Flash achieved 69%, tying with GPT-6 Astra and surpassing GLM-5.3's 62%; on the long-context test AA-LCR, it scored 84%.
However, its output length is relatively high, with AA testing showing an average output of approximately 89,000 tokens per task—62% higher than V4 Pro—partially offsetting its price advantage.
