Anonymous AI model 'stealth/ox-alpha' demonstrates superior programming skills; technical clues suggest it is Zhipu AI’s next-generation model.

iconMetaEra
Share
AI summary iconSummary
A new anonymous AI model, 'stealth/ox-alpha,' has emerged on OpenRouter with programming capabilities exceeding those of GPT-5.6 and Claude. Technical indicators show it supports text, image, and video inputs with a 1,048,000-token context window. Analysis links it to Zhipu AI’s GLM-5.x series. Ox-alpha scored 80% on DeepSWE, outperforming leading models. Zhipu AI has not commented. The Fear and Greed Index in the AI market indicates rising interest in next-generation models.
An anonymous AI model, "stealth/ox-alpha," has recently appeared on the OpenRouter platform. Independent testing revealed that the model supports text, image, and video inputs, with a context window of 1.048 million tokens. Technical analysis indicates that multiple features—including its video encoder, tokenizer, and output style—closely align with those of the Zhipu GLM-5.x series. Researchers are "99% certain" that this model is Zhipu’s upcoming next-generation multimodal flagship. In programming capability tests, ox-alpha achieved an 80% Pass@1 score on the DeepSWE benchmark, surpassing mainstream models such as Claude Fable 5 and GPT-5.6. Zhipu has not yet issued an official response.

Article author and source: Wall Street Journal

An anonymous AI model has quietly appeared on the model distribution platform OpenRouter. Independent testing reveals that the model, named "stealth/ox-alpha," not only outperforms leading models like GPT and Claude on certain programming benchmarks, but its technical characteristics strongly point to Zhipu’s upcoming next-generation multimodal flagship model.

According to a test report released by technical researcher Ben Davis on August 21, ox-alpha launched on OpenRouter on August 20 and is currently offering one week of free access. The model supports text, image, and video inputs, features reasoning capabilities, and has a context window of 1.048 million tokens.

Davis stated on X that he is “99% certain” that ox-alpha belongs to the Zhipu GLM-5.x series, with evidence from the video encoder, tokenizer, output style, and audio rejection behavior all pointing to this conclusion. His tests also showed that ox-alpha outperformed some GPT and Claude models in comparative evaluations.

However, it should be emphasized that the above conclusions are based solely on personal testing and technical inference and have not yet been confirmed by Zhipu AI.

The video encoder became key evidence.

Davis's analysis conducted a technical溯源 of ox-alpha across multiple dimensions, including video encoder, tokenizer, audio interface, and output style, with the high degree of match in the video encoder considered the strongest evidence.

Testing showed that, across four controlled videos, ox-alpha and GLM-5V-Turbo consumed identical video tokens, both exhibiting the same three characteristics: frame-rate-independent frame sampling, a duration scaling ratio of approximately 147 tokens per second, and a per-frame resolution scaling mechanism.

In contrast, candidate models such as MiMo v2.5, Qwen 3.8 Max, and GLM-4.6V all exhibit distinctly different video encoding characteristics.

The tokenizer test also points to GLM. The report states that after testing 25 different prompts, ox-alpha's token count matched GLM-5.3 exactly, with only a fixed +75 token hidden overhead difference, suggesting that both may share the same vocabulary.

In addition, ox-alpha's rejection of audio input aligns with GLM-5V, while MiMo v2.5, listed as a primary competing candidate, supports audio input—further reducing its likelihood.

In terms of output style, ox-alpha uses approximately 1.3 emojis per thousand characters, which is closer to the GLM/Qwen series; in contrast, Claude, GPT-5.6, and Grok have an emoji usage rate close to zero under the same test conditions.

Programming capability has exceeded GLM-5.3

Regarding capability testing, the report cites DeepSWE benchmark data showing that ox-alpha passed 8 out of 10 deterministic tasks, achieving a Pass@1 score of 80%.

For comparison, Claude Fable 5 has a pass rate of 65%, GLM-5.3 and Grok 4.6 both have 62%, and GPT-5.6-sol has 52%. However, since the number of tests varies across models and the ox-alpha sample size is currently small, these results require further validation through additional independent testing.

Among these, in the "meriyah-explicit-resource-declarations" task, ox-alpha passed on the first attempt, while GLM-5.3, GPT-5.6-sol, and Grok 4.6 all previously failed with a score of 0/4. Meanwhile, ox-alpha maintained a perfect record with all 51,469 regression tests passing.

The report also documented an agent task involving 69 tool calls. Throughout the process, the model made only one error, experienced no retry loops, and exhibited low reasoning overhead. Based on this, Davis concluded that ox-alpha's performance significantly surpasses GLM-5.3, resembling a next-generation model checkpoint rather than a mere variant.

Why point to Zhipu?

In addition to technical fingerprints, Davis also made cross-judgments based on the model's release date and Zhipu's previous testing methods.

Zhipu released the text-only GLM-5.3 on August 14, and the unified vision flagship model has consistently been a focus of the community. More importantly, Zhipu has previously tested models through covert channels, and Pony Alpha was ultimately confirmed to be related to GLM-5.

In terms of model scale, the report states that the decoding speed of ox-alpha is approximately 6% different from that of GLM-5V-Turbo. The latter has a total of 744B parameters and 40B activated parameters, leading Davis to speculate that ox-alpha may employ a similar-scale MoE architecture.

The report further suggests that if the model indeed has approximately 40 billion activated parameters, the operator's claimed daily capacity of 100 trillion tokens would be more feasible from both technical and cost perspectives.

Meanwhile, the report systematically ruled out other potential sources. Xiaomi MiMo shows clear differences from ox-alpha in video encoding and audio interfaces; DeepSeek previously did not release video capabilities, and differs in tokenizer and model release methodology; Google, Qwen, xAI, OpenAI, and Anthropic were deemed inconsistent with ox-alpha in terms of tokenizer, output style, or video encoding.

Free trial or ongoing until August 27

Notably, the current free access window for ox-alpha may last until August 27.

Davis noted that previously, some similar stealth models were also officially claimed by relevant Chinese AI labs after the free testing period ended.

Zhipu has not yet made an official statement regarding the identity of ox-alpha. If this model is ultimately confirmed to be Zhipu’s next-generation multimodal flagship, this “stealth test” on OpenRouter could serve as a public preview prior to its official launch.

Before official confirmation, the origin of ox-alpha remains uncertain, but multiple independent tests—from video encoders and tokenizers to model behavior—are currently pointing to the same direction.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.