DeepSeek Releases V4.1 Flash With Native Visual Understanding
DeepSeek released DeepSeek-V4.1-Flash on 10 September, the smallest model in a new architecture family, with native visual understanding and a new Causal Encoder Decoder architecture. The 552 billion parameter mixture of experts model activates 8 billion parameters for input and 16 billion for output.
DeepSeek says new pre training methods and larger scale reinforcement learning put its benchmark results ahead of flagship models including DeepSeek-V4-Pro, and that tests by multiple parties place it ahead of V4-Pro on performance, cost, speed and total runtime. Its key value cache requires one quarter of the high bandwidth memory and one eighth of the SSD storage used by the previous generation.
DeepSeek-V4.1-Flash is available through the DeepSeek API under the deepseek-flash identifier. DeepSeek retired V4-Flash and V4-Flash-Vision-Exp, while DeepSeek-V4-Pro requests began routing to the new model on 14 September at V4.1 Flash rates, an arrangement DeepSeek says will last until V4.1-Pro launches. DeepSeek also cut API prices from 10 September, with off peak rates at half of peak rates. The model and a technical report are available on Hugging Face.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Daily AI Brief.
Also, consider following us on social media:
Daily AI Brief
Daily report covering major AI developments and industry news, with both top stories and complete market updates
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
Abacus.AI Releases Three Smaug Open Weight Models
Deepdub Launches Phantom Z 3.4 Conversational Voice Model
Google Introduces Gemini 3.8 Flash and Flash Cyber
ModelBest Releases MiniCPM5-2B Model for Edge Devices
L7 Informatics Releases L7|ESP 2026.1 for Precision Sciences
Daily AI Brief: the AI news that matters, in your inbox.