D
DeepSeek V4 Flash
DeepSeek · Released Jul 2026
Recommended
DeepSeek's volume-tier open-weight model at 284B parameters (13B active MoE) with a 1M-token context window. Delivers SWE-bench Verified 79.0% at $0.14/$0.28 per 1M tokens — cheaper than any Western frontier-adjacent API by 7-107x. MIT-licensed for self-hosting on a single H100 or 2x A100s. The strongest value proposition in AI — near-frontier coding at commoditized pricing — use as the default workhorse for high-throughput agent pipelines; reserve Pro for the hardest reasoning tasks.
Is it right for you?
Good for
- High-throughput pipelines — classification, extraction, summarization, and structured-output tasks at the lowest cost in frontier-adjacent AI
- Cost-sensitive coding and agentic workflows — SWE-bench Verified 79.0% at $0.14/$0.28, beating models 10-100x its price
- Self-hosted deployments on modest hardware — fits on a single H100 or 2x A100s at INT4 quantization
Not good for
- Hardest reasoning and multi-step agentic tasks — V4 Pro ($0.435/$0.87) or Western flagships lead on sustained complex reasoning
- Vision/multimodal workloads — text-only model
Pricing
Input
$0.14 / 1M tokens
Output
$0.28 / 1M tokens
Context
1M tokens
Benchmarks
No verdict changes yet
The clock starts day one — changes land here as our verdict evolves.
Sources
- AI Pricing Guru — DeepSeek API PricingJul 2026
- Design for Online — V4 Flash ReviewJul 2026
- Requesty — DeepSeek V4 Flash Specs & BenchmarksJul 2026
- BenchLM — DeepSeek API Pricing: V4 Pro & Flash RatesJul 2026
- Tech Jacks Solutions — Running DeepSeek V4 Cost-EffectivelyJul 2026
- NVIDIA Build — deepseek-v4-pro Model CardJul 2026
Verification log
- Pricing— No changes
Automated agent