Chinese AI startup DeepSeek launched DeepSeek-V4.1-Flash on September 10, calling it the smallest model in a new architecture family built for higher capability, faster inference, and greater throughput, Reuters reported from a company statement. The release lands as DeepSeek prepares a Shanghai STAR Market IPO, Reuters has previously reported.
DeepSeek’s own post fills in the product claim. V4.1-Flash is a 552-billion-parameter mixture-of-experts model with a new Causal Encoder–Decoder design that activates about 8 billion parameters on input and 16 billion on output, plus native multimodal support on the API under the model name deepseek-flash. The company says the new architecture cuts KV-cache HBM needs to roughly one-quarter of the prior generation and SSD storage to about one-eighth, which matters for agent workloads where cache hits drive a large share of cost. New API pricing took effect at 04:00 UTC on September 10, with off-peak rates set at half of peak.
The bigger operational shift is retirement. DeepSeek says third-party tests put V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime, and it is phasing the flagship out. Starting at 04:00 UTC on September 14, all deepseek-v4-pro requests will route to V4.1-Flash at Flash rates until a V4.1-Pro ships. Legacy V4-Flash endpoints temporarily route to the new model for compatibility. Official partners WorkBuddy (including CodeBuddy) and OpenCode already support V4.1-Flash, DeepSeek said. This brief covers the Reuters launch wire and DeepSeek’s September 10 product post; independent benchmarks are not yet public.