Alibaba has released and open-sourced the Qwen 3.8 Flash model, which employs a 125B MoE architecture with only about 6B parameters activated per token, natively supporting 262,144 context tokens and extendable to 1M tokens via YaRN. Its innovation lies in enabling deployment on consumer-grade 4090 GPUs, achieving a breakthrough in migrating data center-grade long-context models to consumer hardware. Even more competitively, its pricing stands at just 0.8 RMB per million input tokens and 2.7 RMB per million output tokens—only one-third the cost of DeepSeek-V4-Flash. Real-world tests show the model can draft content for four platforms in two minutes and extract meeting notes, organize tables, and generate emails within four minutes. Qwen 3.8 Flash is now available on the Qwen Office platform, with API access simultaneously open. On the same day, Zhipu also open-sourced GLM-5.3-Flash, whose performance matches Claude Opus 4.8 at one-fortieth the price.Author and source: Quantum Bit
Alibaba's Qwen new model has raised the bar once again.
They released and open-sourced Qwen 3.8 Flash-Next, which not only slashes the price of DeepSeek-V4-Flash but also topped the Hugging Face leaderboard upon release!

Netizens on Twitter were also excited, with some engineers even sharing real-world test videos exclaiming how impressive it was:
Qwen 3.8 Flash has lowered the VRAM barrier,enabling data center-grade long-text large models to run on consumer-grade 4090 GPUs!
Some people have already eagerly put it to use in their work.
In less than eight minutes, it completed an entire workflow covering Python programming, multi-table analysis, and report generation:

Even the Baota special effects are no problem:

Wow, is it really that cool? We definitely have to give it a try.

Only 0.8 yuan per million input tokens
Qwen3.8-Flash employs a 125B MoE architecture, activating only approximately 6B parameters per token, natively supporting up to 262,144 context tokens, and extendable to 1M tokens via YaRN.
At the same time, it also serves as a preview of Qwen4's new architecture.
Compared to Qwen 3.7 Plus, Qwen 3.8 Flash reduces activated parameters to one-third and cuts training costs to just one-ninth, while significantly enhancing capabilities in general tasks, mathematical reasoning, and programming.
Even more impressive is the pricing: Qwen 3.8 Flash costs only ¥0.8 per million tokens for input and ¥2.7 for output, with cache hits as low as ¥0.1—making it as cheap as one-third the price of DeepSeek-V4-Flash.
It's way more intense than Pinduoduo's "cut one more time" (just kidding)

Alright, since the price has dropped to this level, let's not hold back—we'll put Qwen 3.8 Flash to the test with two real-world trials to see how well it performs in office scenarios.
Test Case 1: Design four copy variations tailored to different platform styles for a camera product
We designed a “Copywriting Challenge” where Qwen 3.8 Flash rewrote a nearly 10,000-word camera product description into social media posts for Xiaohongshu, Weibo, Douyin, and Moments—all within a budget of just 1 yuan—to see how many tasks it could complete.
The prompt is as follows:

After sending the prompt, I went to the office kitchenette to get some water, taking less than two minutes total.
When I came back, I found that all four copies for different platforms had been completed, and the brands, styles, and hashtags were perfectly targeted (except that I wouldn’t post something that long on my Moments anyway hhh):




Finished four copies of copywriting in under two minutes—my cyber worker speed is definitely up to par.
But in the workplace, what’s truly frustrating isn’t writing copy—it’s figuring out who will clean up the mess after the meeting.
So, we gave it a second job to test its practical office capabilities.
Test Case 2: Turn a product weekly meeting into an actionable plan
We took the raw, 10,000-word meeting transcript—complete with homophone errors, casual side conversations, and the unpolished tone of a meeting just concluded—and titled it “Requirements and Technical Review Meeting for Smart Customer Service Agent V2.0.” Qwen3.8-Flash processed it in one go, extracting summaries, organizing tables, and drafting the post-meeting notification email.
△ The image was generated by AI; hearing "token" as "steal and nibble" is so life-like.
The prompt is as follows:

In default mode, Qwen3.8-Flash completed all tasks in under four minutes, with accurate summaries of the five core conclusions, correct task assignments in the table, proper email formatting, and recipients perfectly matched to the meeting content.



△ Partial test results chart
Even without being asked, it kindly organizes for us the items from the meeting that haven’t been resolved yet, making it easy for us to pick them up in our next meeting—be like:

And even at this point, I haven’t used up my quota yet!
Hahaha, this feeling of being unrestricted is so great.
Qwen 3.8 Flash is now officially launched on the Qianwen Office platform, and its API interface is also available—feel free to give it a try!

One More Thing
The price war among domestic large models has intensified even further.
Almost simultaneously with the open-sourcing of Qwen 3.8 Flash, Zhipu also open-sourced GLM-5.3-Flash, the first native multimodal model in the GLM-5 series, previously known online as the popular "Niu Lai" large model Ox Alpha.
In evaluations, this model's performance rivals Claude Opus 4.8, while its price has been reduced to one-tenth that of GLM-5.3.

Moreover, GLM-5.3-Flash is generously offering a two-week half-price promotion, meaning that during the promotion period, GLM-5.3-Flash is priced at just 1/20th of GLM-5.3 and 1/40th of Opus 4.8.

In addition, Qwen 3.8 Flash and GLM-5.3-Flash share another commonality: they both caught DeepSeek-V4-Flash off guard with their ultra-low pricing (DeepSeek walked away grumbling):
△ Price comparison chart for DeepSeek-V4-Flash, Qwen 3.8 Flash, and GLM-5.3-Flash
No doubt, it's too competitive.
