D

DeepSeek V4 Flash

DeepSeek · Released Jul 2026

Recommended

DeepSeek's volume-tier open-weight model at 284B parameters (13B active MoE) with a 1M-token context window. SWE-bench Verified 79.0% at $0.22 input and $0.66 output per 1M tokens off-peak, doubling during the seven daily peak hours. The August 2026 price rise ended its run as the cheapest credible API (GPT-5.6 Luna now undercuts it on input at $0.20), but it remains a strong value pick, and the MIT license means self-hosting sidesteps API pricing entirely. Still the default workhorse for high-throughput agent pipelines that can run off-peak.

Is it right for you?

Good for

  • Self-hosted deployments of a frozen checkpoint: the MIT-licensed 284B-parameter (13B active) V4 Flash weights remain downloadable from Hugging Face, so pipelines already validated on V4 Flash can keep running it in-house
  • Reproducible evals and audit trails: self-hosted V4 weights never change, while the deepseek-v4-flash API id began returning V4.1-Flash output on 10 September 2026 without any client-side change
  • Classification, extraction and structured-output workloads on your own hardware, where SWE-bench Verified 79.0% capability comes with no per-token bill

Not good for

  • New API integrations: DeepSeek retired V4 Flash on 10 September 2026. The deepseek-v4-flash id now serves DeepSeek-V4.1-Flash, so calling it does not give you this model; use deepseek-flash directly
  • Relying on the legacy API id for stable behaviour: DeepSeek says the old name is only temporarily routed to V4.1-Flash, and requests bill at the Flash rate of $0.15 input / $0.60 output per 1M off-peak, doubling 01:00-04:00 and 06:00-10:00 UTC Monday to Friday
  • Hardest reasoning and multi-step agentic tasks: V4 Pro or Western flagships lead on sustained complex reasoning
  • Vision and multimodal workloads: text-only model

Pricing

Input

$0.15 / 1M tokens

Output

$0.60 / 1M tokens

Context

Retired Sep 10; id serves V4.1-Flash. Peak 2x M-F

View full pricing

Benchmarks

BenchmarkScoreSource
SWE-bench Verified79.0% Source
SWE Pro (Resolved)52.6% Source
GPQA Diamond89.4% Source
τ²-Bench95.0% Source
Intelligence Index40.3 Source
Coding Index56.2 Source

Verdict history

Aug 20, 2026
recommended -> recommended (pricing premise corrected) — Flagged for re-review on 2026-08-15 when DeepSeek announced the hike. The old summary priced Flash at $0.14/$0.28 and claimed it was "cheaper than any Western frontier-adjacent API by 7-107x". Vendor pricing confirmed 2026-08-20: off-peak $0.22/$0.66, peak $0.44/$1.32, effective 2026-08-16. That 7-107x claim is now false; GPT-5.6 Luna undercuts Flash on input at a vendor-confirmed $0.20. Rating held at recommended because the output rate off-peak still beats Luna and the MIT self-hosting path is unaffected by API pricing.

Verification log

  • Pricing— Updated

    Automated agent

    MATERIAL: the model this entity describes no longer serves API requests. Rate card footnote (1), quoted verbatim: 'Use `deepseek-flash` as the model name. The legacy names `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price.' Table column deepseek-flash (MODEL VERSION DeepSeek-V4.1-Flash): cache-miss input $0.15 off-peak / $0.3 peak, output $0.6 off-peak / $1.2 peak, cache hit $0.003 / $0.006. Launch post api-docs.deepseek.com/news/news260910 (2026/09/10): 'V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.' Weights: huggingface.co/deepseek-ai/DeepSeek-V4-Flash still live, MIT, 284B/13B active, safetensors present. REFRESHED: pricing_input $0.22 -> $0.15, pricing_output $0.66 -> $0.60 (the rate the legacy id now bills at); pricing_context -> 'Retired Sep 10; id serves V4.1-Flash. Peak 2x M-F'; strengths and audiences rewritten around retirement + self-hosting (removed the $0.22/$0.66 and $0.44/$1.32 rows, the now-false 'GPT-5.6 Luna undercuts Flash on input at $0.20' row, the API model-routing audience, and the unsourced 'fits on a single H100 at INT4' claim, which is arithmetically impossible for 284B params at 4-bit, ~142 GB); added news + HF sources. Benchmarks untouched (they describe the V4 checkpoint). VERDICT: live summary still says '$0.22/$0.66' and 'seven daily peak hours' and calls a retired API 'the default workhorse'. entity_verdict recommended -> conditional STAGED for the existing curator draft 'deepseek-v4-1-flash-launch-v4-flash-retired' (no new draft, no verdict_set). The 2026-09-09 draft 'deepseek-v4-flash-peak-window-weekdays-only' states $0.22/$0.66 as live and its staged payload is superseded: HOLD it. New-model row for V4.1-Flash is the researcher's call (candidate deepseek-v4-1-flash).

  • Pricing— Updated

    Automated agent

    Rates re-verified UNCHANGED against the vendor rate card: off-peak $0.22 in / $0.66 out, peak $0.44 / $1.32 per 1M; cache hits $0.007 off-peak / $0.014 peak. Peak rule quoted verbatim: 'Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak).' TWO defects fixed. (1) pricing_url was stored WITHOUT a trailing slash and resolves to a page titled 'Your First API Call | DeepSeek API Docs' carrying no prices at all; the rate card is the same path WITH the slash, title 'Models & Pricing | DeepSeek API Docs'. Corrected. Verified reproducible: both URLs fetched this run, different pages. Same fix applied to deepseek-v4-pro. (2) The verdict summary still reads 'doubling during the seven daily peak hours' while this entity's own pricing_context and strengths have said Mon-Fri since 2026-09-07 - the entity contradicts itself in public. The 2026-09-07 log for this slug states the wording fix was staged in that run's return with NO announcing post; it was never routed. Now announced: draft 'deepseek-v4-flash-peak-window-weekdays-only' created this run with the entity_verdict payload staged in the return. Rating unchanged at recommended.

  • Pricing— Updated

    Automated agent

    Per-token rates re-verified and UNCHANGED: off-peak $0.22 in / $0.66 out, peak $0.44 / $1.32 per 1M. Same peak-window scope correction as V4 Pro: the vendor page scopes the 2x window to Monday through Friday, while our stored guidance said the rise doubled rates 'for 7 hours a day'. Refreshed pricing_context to '1M tokens. Peak 2x: 01-04+06-10 UTC Mon-Fri', rewrote the peak avoid_for row to name weekdays and the doubled rates, and added a weekend off-peak throughput best_for row. The 2026-08-20 pass already removed the false '7-107x cheaper' multiple and the GPT-5.6 Luna $0.20 input undercut was re-confirmed against the live Luna entity, so no false price multiple remains. Rating unchanged at recommended. NOTE: the verdict summary still reads 'doubling during the seven daily peak hours' — that wording fix is staged in the run return but has NO announcing post of its own (the DeepSeek draft carries the V4 Pro payload; one entity_verdict per post). Flagged for routing, not silently left.

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

How we evaluate