DeepSeek has launched DeepSeek v4.1-Flash, a new open-weight flagship model featuring a novel causal encoder-decoder architecture designed to maximize inference efficiency and reduce costs.
The model utilizes a mixture-of-experts structure with 763 billion total parameters, but only activates 8B parameters for input processing and 16B for output generation. This design achieves a sparsity of 1-2% and reduces the KV cache footprint to one-eighth that of the previous V4 Flash model. It supports a 1M-token context window with text and image inputs under an MIT license.
Independent benchmarks from Artificial Analysis and Vals rank v4.1-Flash as the top open-weight model, citing its superior cost-performance ratio with pricing at $0.30 per million input tokens.