On September 10, 2026, DeepSeek officially released the DeepSeek V4.1 Flash model. It is the most compact representative of the new-generation architecture, yet it possesses native multimodal image understanding capabilities and outperforms flagship models—including the V4 Pro—in a number of key benchmarks.

New asymmetric architecture: low cost, high intelligence

V4.1 Flash employs a 552-billion-parameter MoE architecture featuring a fundamentally new Causal-Encoder-Decoder structure. Its key feature is input-output asymmetry: only 8 billion parameters are activated when processing input data, while 16 billion are activated during response generation, making it significantly less costly than models of comparable size.

The model consists of 40 Transformer layers: the first 20 form a causal encoder, and the last 20 form a decoder. The decoder's global KV cache is projected directly from the encoder's last layer, drastically reducing the computational load during the prefill stage.

Cash reserves are squeezed to the limit, and the cost has been significantly reduced.

The new generation has achieved a breakthrough in KV-cache efficiency: compared to the previous V4 Flash, the requirement for HBM has been reduced fourfold, and for SSDs , eightfold. According to the developer, the KV-cache size has shrunk 437-fold relative to the very first generation, V1.

This is particularly important for agent-based scenarios, where cache-related costs account for a significant share, and reducing the cache size directly lowers usage costs.

Performance surpasses the flagship.

In benchmarks, V4.1 Flash outperformed V4 Pro in terms of intelligence. In the Terminal-Bench 2.1 test, the new model scored 90.6, compared to 87.9 for V4 Pro and 82.7 for the previous V4-Flash. In the more challenging Terminal-Bench 3.0, the gap widened dramatically: 30.0 versus 11.8 and 7.6, respectively.

In DeepSWE v1.1 (software development), the result was 74.2 compared to 62.7 for V4 Pro. In the Codeforces ranking, the model achieved a score of 3471, surpassing V4 Pro’s 3348. In the field of cybersecurity (CyberGym), the score was 88.1 versus 83.3.

Thus, V4.1 Flash significantly outperforms its predecessors in terminal operations, code generation, and cybersecurity—key agent use cases.

API and pricing

V4.1 Flash is now available via the DeepSeek API; simply specify the model name `deepseek-flash`. Prices have also been reduced: during off-peak hours, the cost is approximately €0.0026 for input (cache hit), €0.13 for input (cache miss), and €0.51 for output. Prices are double during peak hours and half during off-peak hours.

The old V4 Flash and V4 Flash Vision Exp models have been disabled, but for compatibility reasons, the old model names are temporarily being redirected to V4.1 Flash.

What's going on with the V4 Pro?

According to the latest information, in response to user requests, DeepSeek has decided to continue providing the V4 Pro API after September 14, with pricing remaining unchanged. Previously, the plan was to automatically redirect V4 Pro requests to V4.1 Flash .

For readers following developments in construction and 3D printing , the native multimodal capability of V4.1 Flash is of particular interest: the model can directly recognize construction blueprints, equipment photos, or images of work progress, opening up new possibilities for processing engineering documentation.

Source: deepseek.com