started · updated
DeepSeek launches V4.1-Flash AI model to reduce inference costs
DeepSeek has released V4.1-Flash, a multimodal AI model utilizing a Mixture-of-Experts architecture with 552 billion parameters. The model is designed to significantly reduce inference and memory costs by activating only 8 billion parameters for input and 16 billion for output. Notably, the new architecture reportedly reduces the KV-cache size to approximately one-quarter of the high-bandwidth memory used by the previous V4 Flash generation.
In benchmark tests, V4.1-Flash achieved a score of 90.6 on Terminal-Bench 2.1, outperforming competitors such as OpenAI’s GPT-5.6 Sol and Moonshot AI’s Kimi K3. The model also shows strong performance in cybersecurity and software engineering agent tasks.
The release has caused market reactions and developer debate. The technical efficiency of the model led to a decline in shares for Samsung Electronics and SK Hynix as investors weighed shifts in AI hardware requirements. Additionally, some developers have expressed concerns regarding DeepSeek’s decision to route V4 Pro requests to the new Flash model with minimal transition time, citing potential stability issues for production environments and research reproducibility.
Entities
Claude Opus 5 · DeepSeek · GPT-5.6 Sol · Hugging Face · OpenAI