Qwen 3.8 Flash Reduces Costs to One-Third of DeepSeek-V4-Flash

iconMetaEra
Share
AI summary iconSummary
Alibaba’s Qwen 3.8 Flash, now available on the Qwen Office platform, offers 125B MoE with 6B active parameters per token. It natively supports 262,144 tokens and runs on 4090 GPUs. Priced at 0.8 RMB per million input tokens, it costs one-third of DeepSeek-V4-Flash. Traders monitoring altcoins to watch may find this a significant development. The model handles complex tasks in under four minutes and is API-ready. On the same day, Zhipu open-sourced GLM-5.3-Flash, priced at one-fortieth of Claude Opus 4.8. Market sentiment, including the Fear & Greed Index, could shift as such cost-efficient models enter the space.
Alibaba has released and open-sourced the Qwen 3.8 Flash model, which employs a 125B MoE architecture with only about 6B parameters activated per token, natively supporting 262,144 context tokens and extendable to 1M tokens via YaRN. Its innovation lies in enabling deployment on consumer-grade 4090 GPUs, achieving a breakthrough in migrating data center-grade long-context models to consumer hardware. Even more competitively, its pricing stands at just 0.8 RMB per million input tokens and 2.7 RMB per million output tokens—only one-third the cost of DeepSeek-V4-Flash. Real-world tests show the model can draft content for four platforms in two minutes and extract meeting notes, organize tables, and generate emails within four minutes. Qwen 3.8 Flash is now available on the Qwen Office platform, with API access simultaneously open. On the same day, Zhipu also open-sourced GLM-5.3-Flash, whose performance matches Claude Opus 4.8 at one-fortieth the price.

Author and source: Quantum Bit

Alibaba's Qwen new model has raised the bar once again.

They released and open-sourced Qwen 3.8 Flash-Next, which not only slashes the price of DeepSeek-V4-Flash but also topped the Hugging Face leaderboard upon release!

Image

Netizens on Twitter were also excited, with some engineers even sharing real-world test videos exclaiming how impressive it was:

Qwen 3.8 Flash has lowered the VRAM barrier,enabling data center-grade long-text large models to run on consumer-grade 4090 GPUs!

Image

Some people have already eagerly put it to use in their work.

In less than eight minutes, it completed an entire workflow covering Python programming, multi-table analysis, and report generation:

Image

Even the Baota special effects are no problem:

Image

Wow, is it really that cool? We definitely have to give it a try.

Image

Only 0.8 yuan per million input tokens

Qwen3.8-Flash employs a 125B MoE architecture, activating only approximately 6B parameters per token, natively supporting up to 262,144 context tokens, and extendable to 1M tokens via YaRN.

At the same time, it also serves as a preview of Qwen4's new architecture.

Compared to Qwen 3.7 Plus, Qwen 3.8 Flash reduces activated parameters to one-third and cuts training costs to just one-ninth, while significantly enhancing capabilities in general tasks, mathematical reasoning, and programming.

Even more impressive is the pricing: Qwen 3.8 Flash costs only ¥0.8 per million tokens for input and ¥2.7 for output, with cache hits as low as ¥0.1—making it as cheap as one-third the price of DeepSeek-V4-Flash.

It's way more intense than Pinduoduo's "cut one more time" (just kidding)

Image

Alright, since the price has dropped to this level, let's not hold back—we'll put Qwen 3.8 Flash to the test with two real-world trials to see how well it performs in office scenarios.

Test Case 1: Design four copy variations tailored to different platform styles for a camera product

We designed a “Copywriting Challenge” where Qwen 3.8 Flash rewrote a nearly 10,000-word camera product description into social media posts for Xiaohongshu, Weibo, Douyin, and Moments—all within a budget of just 1 yuan—to see how many tasks it could complete.

The prompt is as follows:

Image

After sending the prompt, I went to the office kitchenette to get some water, taking less than two minutes total.

When I came back, I found that all four copies for different platforms had been completed, and the brands, styles, and hashtags were perfectly targeted (except that I wouldn’t post something that long on my Moments anyway hhh):

Image

Image

Image

Image

Finished four copies of copywriting in under two minutes—my cyber worker speed is definitely up to par.

But in the workplace, what’s truly frustrating isn’t writing copy—it’s figuring out who will clean up the mess after the meeting.

So, we gave it a second job to test its practical office capabilities.

Test Case 2: Turn a product weekly meeting into an actionable plan

We took the raw, 10,000-word meeting transcript—complete with homophone errors, casual side conversations, and the unpolished tone of a meeting just concluded—and titled it “Requirements and Technical Review Meeting for Smart Customer Service Agent V2.0.” Qwen3.8-Flash processed it in one go, extracting summaries, organizing tables, and drafting the post-meeting notification email.

Image△ The image was generated by AI; hearing "token" as "steal and nibble" is so life-like.

The prompt is as follows:

Image

In default mode, Qwen3.8-Flash completed all tasks in under four minutes, with accurate summaries of the five core conclusions, correct task assignments in the table, proper email formatting, and recipients perfectly matched to the meeting content.

Image

Image

Image

Image△ Partial test results chart

Even without being asked, it kindly organizes for us the items from the meeting that haven’t been resolved yet, making it easy for us to pick them up in our next meeting—be like:

Image

And even at this point, I haven’t used up my quota yet!

Hahaha, this feeling of being unrestricted is so great.

Qwen 3.8 Flash is now officially launched on the Qianwen Office platform, and its API interface is also available—feel free to give it a try!

Image

One More Thing

The price war among domestic large models has intensified even further.

Almost simultaneously with the open-sourcing of Qwen 3.8 Flash, Zhipu also open-sourced GLM-5.3-Flash, the first native multimodal model in the GLM-5 series, previously known online as the popular "Niu Lai" large model Ox Alpha.

In evaluations, this model's performance rivals Claude Opus 4.8, while its price has been reduced to one-tenth that of GLM-5.3.

Image

Moreover, GLM-5.3-Flash is generously offering a two-week half-price promotion, meaning that during the promotion period, GLM-5.3-Flash is priced at just 1/20th of GLM-5.3 and 1/40th of Opus 4.8.

Image

In addition, Qwen 3.8 Flash and GLM-5.3-Flash share another commonality: they both caught DeepSeek-V4-Flash off guard with their ultra-low pricing (DeepSeek walked away grumbling):

Image△ Price comparison chart for DeepSeek-V4-Flash, Qwen 3.8 Flash, and GLM-5.3-Flash

No doubt, it's too competitive.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.