DeepSeek Retired V4 Flash the Day V4.1-Flash Shipped. Anything Calling deepseek-v4-flash Is Already on the New Model
- What happened
- DeepSeek released V4.1 Flash (552B MoE, open weights) on 10 September and retired V4 Flash, routing the deepseek-v4-flash id to the new model at $0.15/$0.60 per 1M off-peak.
- Why it matters
- Anyone calling deepseek-v4-flash is silently running a different model: V4 Flash's recommended rating does not transfer, and we rate V4.1 Flash conditional.
- What to do
- Set your model to deepseek-flash, re-run evals on anything that called the old id since 10 September, and leave V4 Pro workloads where they are.
If your code calls deepseek-v4-flash, you are no longer running the model we rated. DeepSeek retired DeepSeek V4 Flash on 10 September and pointed the old id at DeepSeek V4.1 Flash, a different and bigger model. Switch the model name to deepseek-flash and re-run your evals. V4 Flash's recommended rating does not carry over: we rate V4.1 Flash conditional. DeepSeek V4 Pro keeps its entry and its rating.
What happened
DeepSeek's launch post, dated 2026/09/10, is explicit:
V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility,
deepseek-v4-flashanddeepseek-v4-flash-vision-exptemporarily route to V4.1-Flash.
The pricing page says the same about billing: the legacy names are still accepted, their requests are served by DeepSeek-V4.1-Flash and billed at the Flash price.
"Temporarily" comes with no end date. Treat the alias as something that will stop working.
What DeepSeek's launch post says about the new model:
- Size. A "552B-parameter MoE" on a "new Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output." V4 Flash was 284B parameters with 13B active, per our V4 Flash entry.
- Memory. The KV cache needs "1/4 the HBM" and "1/8 the SSD storage" of the previous generation.
- Modality. Native multimodal support (visual understanding) on the API.
- Weights. Published at huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash(opens in new tab). The launch post names no licence, so read the model card before assuming V4 Flash's MIT terms carry over.
- Name. Set your model to
deepseek-flash.
Why it matters
Same id, different model. Every call to deepseek-v4-flash since 10 September was answered by V4.1 Flash. Evals, prompts and output budgets tuned on V4 Flash describe a model that no longer serves traffic.
The alias got cheaper. Per 1M tokens, off-peak / peak, from DeepSeek's pricing page, against the V4 Flash rates we re-verified on 9 September:
V4.1 Flash (deepseek-flash) | V4 Flash, retired | |
|---|---|---|
| Input, cache miss | $0.15 / $0.30 | $0.22 / $0.44 |
| Output | $0.60 / $1.20 | $0.66 / $1.32 |
That is about 32% off input and about 9% off output. Cache hits bill at $0.003 off-peak and $0.006 peak. Context is 1M tokens, max output 384K, concurrency 2,500. The peak window is unchanged: 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday.
Cheaper is not the same as proven. DeepSeek claims "tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime." It names no parties and publishes no figures in the launch post, and we have not run V4.1 Flash ourselves. That is why our V4.1 Flash entry rates it conditional: its standout coding-agent results are DeepSeek's own, and peak hours double the rate.
V4 Pro: we are not reporting a retirement. DeepSeek's two pages still disagree. The launch post says that from 04:00 UTC on 14 September all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates, "until V4.1-Pro launches". The pricing page says the opposite: "We have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged," and still lists deepseek-v4-pro at $0.66 / $1.98 off-peak. We follow the pricing page until DeepSeek reconciles its own documentation.
What changes for you
- Rename the model. Change
deepseek-v4-flashtodeepseek-flashnow, before the temporary alias ends without notice. - Re-run your evals on any pipeline that called
deepseek-v4-flashordeepseek-v4-flash-vision-expsince 10 September. - Re-price batch budgets at $0.15 input and $0.60 output per 1M off-peak, doubling in the weekday peak window.
- Self-hosting? Read the V4.1 Flash model card for licence and hardware needs; a 552B model does not size like a 284B one.
- Leave V4 Pro workloads where they are. The pricing page says its API service continues after 14 September.
Our V4 Flash entry keeps its recommended rating as the record of the retired model; its listed prices now show the V4.1 Flash rates the old id bills at.
FAQ
Does deepseek-v4-flash still work?
Yes, for now. DeepSeek routes it to V4.1 Flash and bills it at the Flash price, but calls the routing temporary and gives no end date.
Does the V4 Flash recommended rating apply to V4.1 Flash? No. We rate V4.1 Flash conditional, separately from the retired V4 Flash.
Is V4 Pro being retired? Not per DeepSeek's pricing page, which says V4 Pro API service continues after 14 September 2026 with unchanged billing. The launch post says otherwise, so watch for a correction.
What to do
- 1 Change the model name from deepseek-v4-flash to deepseek-flash; DeepSeek calls the old-id routing temporary and gives no end date.
- 2 Re-run your evals on any pipeline that called deepseek-v4-flash or deepseek-v4-flash-vision-exp since 10 September: the outputs came from a different model.
- 3 Re-price batch budgets at $0.15 input and $0.60 output per 1M off-peak, doubling 01:00-04:00 and 06:00-10:00 UTC on weekdays.
- 4 If you self-host, read the V4.1 Flash model card for the licence and hardware needs before assuming V4 Flash's terms or sizing carry over.
- 5 Do not migrate V4 Pro workloads on the strength of the launch post: DeepSeek's pricing page says V4 Pro API service continues after 14 September.
Affected tools & models
Never need to catch up again
The weekly delta — only verdict changes and act-now items. No digest filler.