Zhipu has launched the new GLM-5.3-Flash model (codenamed Ox Alpha), scoring 57 on the Artificial Analysis intelligence index, tying with Claude Opus 4.8.Article author and source: LatePost
Zhipu's mysterious model, Ox Alpha, has officially been unveiled, driving a significant rise in the company's stock price and attracting widespread attention in the AI hardware sector due to the disclosure that it fully utilizes domestic chips to handle massive inference workloads.
On Thursday, Zhipu's stock rose over 9% intraday to HK$1,124. On the news front, Zhipu officially confirmed on August 26 that the anonymous model Ox Alpha, which sparked widespread discussion in the developer community and is known in Chinese circles as "Niu Lai," is its newly released GLM-5.3-Flash (320B-A18B). The model scored 57 on the Artificial Analysis Intelligence composite intelligence index, matching Anthropic’s Claude Opus 4.8 and surpassing DeepSeek’s flagship model V4 Pro, which scored 53.

Before its official release, GLM-5.3-Flash was made available for free testing on OpenRouter and OpenCode under anonymous conditions, accumulating over 50 trillion tokens in traffic within five days, setting new records for traffic growth on both platforms. Current usage has surpassed twice that of DeepSeek. Zhipu also disclosed that all this traffic was handled entirely by domestic chips, with over 100,000 domestic chip units deployed for inference services, stating that “hardware efficiency and cost per token have reached levels comparable to mainstream NVIDIA GPUs.”
Fully powered by domestic chips, the hardware narrative has drawn market attention. In this disclosure, the performance of domestic chips has been one of the key focuses for the market. According to Zhipu’s official technical documentation, the company has deployed over 100,000 domestic chips in a cluster to provide inference computing power for GLM-5.3-Flash, stating that its hardware efficiency and cost per token are now comparable to NVIDIA GPUs.
The semiconductor research firm SemiAnalysis immediately commented on X: “All traffic is carried by domestic chips, with hardware efficiency and cost per token now comparable to NVIDIA GPUs. Following yesterday’s announcement of Jalapeño (OpenAI’s proprietary inference chip), the CUDA moat is once again under scrutiny.” The firm specifically noted that it was previously widely believed that only top-tier frontier labs could achieve a computational scale of 100 trillion tokens per day—yet here, 100 trillion tokens of free traffic daily are all running on domestic chips.
According to LatePost, the suppliers of these chips may include Huawei, Moore Threads, and Hygon. Zhipu has not commented on this, and the technical documentation does not specify the exact chip models.
To address the bottlenecks of memory capacity and bandwidth per chip, Zhipu has built a dedicated inference engine on top of SGLang, incorporating technologies such as W8A8 quantization, mixed-cache quantization with INT8/FP8/BF16, and intra-node tensor parallelism, while introducing a production-grade Encode–Prefill–Decode (EPD) separated architecture. Zhipu reports that this approach achieves a threefold improvement in end-to-end service performance compared to the initial baseline on the same hardware.
Architecture Refactoring: Parameters and layers nearly halved—GLM-5.3-Flash differs significantly from its predecessor at the architectural level, which is the core reason it achieves higher performance at a lower cost.
Compared to GLM-4.5, GLM-5.3-Flash has a similar total parameter count (355B vs. 320B), but the activated parameters have decreased from 32B to 18B, and the number of layers has been reduced from 92 to 45—nearly halved. The parameter count is approximately 40% of the previous flagship GLM-5.3 and is nearly equivalent to DeepSeek V4 Flash.
Zhipu stated that GLM-5.3-Flash is the first open-source frontier model to adopt a hybrid architecture combining sparse attention and linear attention, reducing attention computation and KV cache size by 3.01x and 4.44x, respectively, compared to GLM-5.3. Additionally, this model is the first natively multimodal model in the GLM-5 series, supporting image and video inputs, and marks Zhipu’s first new model with multimodal capabilities since strategically focusing on coding.
Pricing Strategy: Targeting the Demand Gap Following the Price Surge The release of GLM-5.3-Flash comes immediately after a wave of collective price increases in China’s AI model market. Kimi K3’s output price has surged to 100 yuan per million tokens—more than three times higher than its predecessor, K2.6; DeepSeek V4 Pro has risen from 6 yuan to 13.5 yuan, reaching up to 27 yuan during peak hours; and DeepSeek V4 Flash’s output price has increased from 2 yuan to 4.5 yuan, peaking at 9 yuan during high-demand periods.
According to LatePost, after DeepSeek V4 Flash increased its price, the number of calls on the OpenCode platform dropped by half, and this spilled-over demand has become a new growth opportunity for domestic model providers.
The pricing strategy for GLM-5.3-Flash is specifically designed to address this gap. The model is priced at CNY 0.8 per million tokens for input, CNY 2.8 per million tokens for output, and CNY 0.23 for cache hits—just one-tenth the cost of GLM-5.3. During the first two weeks after launch, it will be offered at half price. Zhipu AI states that during this limited-time discount period, the pricing is 1/20 that of GLM-5.3 and 1/40 that of Claude Opus 4.8. Compared to DeepSeek V4 Flash after its price adjustment, GLM-5.3-Flash is more affordable in the vast majority of standard usage scenarios.
Zhipu plans to release its first detailed financial report next Monday, covering the first six months of performance since its listing in January. The launch of GLM-5.3-Flash and its market performance will provide investors with a key reference for evaluating the company’s commercialization progress.
