DeepSeek Launches V4.1-Flash Model with 552B Parameters and 1M-Token Context Window

iconCryptoBriefing
Share
AI summary iconSummary
DeepSeek released a new token listing on September 10 with the V4.1-Flash model, featuring 552B parameters and a 1M-token context window. The model uses MoE architecture, activating 8B for input and 16B for output. It adds an asymmetric Causal Encoder-Decoder design and cuts KV cache use by 75%. The model beats DeepSeek V4-Pro and is now accessible via API under updated pricing from September 14. Weights are open-sourced on Hugging Face under MIT license, offering fresh token launch news for developers and traders.

DeepSeek just dropped a model that processes a million tokens of context while activating fewer parameters than some open-source models released two years ago. The V4.1-Flash, launched on September 10, represents the Chinese AI startup’s latest bid to rewrite the economics of large language models.

The model packs 552 billion parameters into a Mixture-of-Experts (MoE) architecture, but only fires up about 8 billion of them for input tasks and 16 billion for output. The result is a model that punches well above what its active compute footprint would suggest.

The architecture that makes it work

V4.1-Flash introduces what DeepSeek calls an asymmetric Causal Encoder-Decoder architecture, processing input and generating output through different pathways optimized for each task, rather than running everything through a single pipeline.

Advertisement

The context window stretches to 1 million tokens. Supporting that massive context is a KV cache that consumes approximately 890 bytes per token, about one-quarter of what the prior V4-Flash model required.

That cache reduction matters more than it might sound. KV cache is the memory bottleneck that determines how many concurrent users a model can serve and how long their conversations can run. Cutting it by 75% means operators can serve roughly four times as many users on the same hardware, or handle contexts four times as long without upgrading their GPU clusters.

On benchmarks, DeepSeek claims V4.1-Flash outperforms the company’s own V4-Pro model and competes directly with GPT-5.6 and Kimi K3. The company is already routing requests from V4-Pro to the new model, with updated pricing taking effect on September 14.

Open weights, open strategy

DeepSeek released the model’s weights under an MIT license on Hugging Face. The new model is accessible through DeepSeek’s API under the endpoint “deepseek-flash.” The combination of open weights and API access creates a two-track adoption path: developers who want to run inference on their own infrastructure can download and deploy locally, while those who prefer managed services can call the API at DeepSeek’s new pricing tiers.

IPO implications and competitive positioning

The timing of the V4.1-Flash launch is not accidental. DeepSeek is expected to go public on Shanghai’s STAR Market, and demonstrating continued technical momentum is the kind of thing that makes roadshow presentations more convincing.

The multimodal understanding built into V4.1-Flash, spanning both visual and text data, narrows the differentiation opportunities available to rivals including OpenAI and Moonshot AI, which develops Kimi K3.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.