ChainThink reports that on August 28, according to Tencent's official announcement, its Hy4 preview model scored higher than GLM-5.3 and Kimi K3 in internal engineering blind tests.
163 Tencent internal experts tested Hy4 using 203 real engineering tasks; Hy4 achieved an average score of 2.99/4, higher than GLM-5.3’s 2.92/4 and Kimi K3’s 2.94/4.
In direct comparisons, Hy4 has a win rate of 46.8%, a draw rate of 12.8%, and a loss rate of 40.4% against GLM-5.3; against Kimi K3, it has a win rate of 51.2%, a draw rate of 7.9%, and a loss rate of 40.9%.
However, Hy4 did not lead comprehensively in public benchmarks and still lags behind GLM-5.3 in code and cybersecurity tests such as DeepSWE and CyberGym.
Price is Hy4’s key advantage: at ¥6 per million tokens for input and ¥18 for output, it is 25% and approximately 36% cheaper than GLM-5.3, and 70% and 82% cheaper than Kimi K3.
Cache hit costs only $0.30, which is 85% cheaper than GLM-5.3 and Kimi K3 at $2.
