Alibaba Launches Qwen3.8-Max AI Model on Nvidia GB300, Hits 4,000 Tokens Per Second

iconCryptoBriefing
Share
AI summary iconSummary
Alibaba released Qwen3.8-Max on August 3, a 2.4 trillion parameter AI model running at over 4,000 tokens per second on Nvidia GB300 NVL72. The model supports multimodal input and a 1 million token context window. It scored 86.1 on the OSWorld-Verified benchmark, outperforming GPT-5.6 and Claude Fable 5 in specific tasks. Alibaba will open-source the model and a smaller version, Qwen3.8-27B, next week. This on-chain news highlights AI + crypto news momentum.

Alibaba just put the global AI leaderboard on notice. The company’s Qwen team launched Qwen3.8-Max on August 3, a massive 2.4 trillion parameter model that runs at over 4,000 tokens per second per GPU on Nvidia’s GB300 NVL72 hardware.

To put that speed in perspective, 4,000 tokens per second per GPU means the model can generate roughly 3,000 words of text every single second on a single chip.

What Qwen3.8-Max actually is

The model operates as a sparse Mixture-of-Experts system, a design pattern where only a fraction of the model’s total parameters activate for any given input. Of its 2.4 trillion total parameters, approximately 95 billion are active per token.

Advertisement

Qwen3.8-Max handles text, images, and video natively as a multimodal system. It also supports a context window of 1 million tokens, which translates to roughly 750,000 words of input.

On the OSWorld-Verified benchmark, a test designed to measure how well AI models can autonomously complete computer tasks, Qwen3.8-Max scored 86.1. That places it ahead of both Anthropic’s Claude Fable 5 (85.0) and OpenAI’s GPT-5.6 Sol Max. It ranks as the second highest-performing AI model overall behind Fable 5 on broader evaluations.

The model targets advanced coding, long-horizon planning tasks, and complex agentic workflows. Alibaba is pricing API access at $2 per million input tokens and $6 per million output tokens.

The hardware story matters just as much

The GB300 NVL72 represents Nvidia’s latest rack-scale inference platform, and early indications suggest it can deliver over 4,500 tokens per second per GPU for Qwen family models under optimized conditions.

Open weights and the competitive landscape

Alibaba plans to release open weights for both Qwen3.8-Max and a smaller variant, Qwen3.8-27B. Those were expected to become available in the week following launch, giving researchers, startups, and enterprise developers direct access to fine-tune and deploy the models.

At $2 per million input tokens, Qwen3.8-Max undercuts the premium pricing tiers of both Claude and GPT on a per-token basis while claiming comparable or superior performance on key benchmarks.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.