DeepSeek V4.1 Flash
DeepSeek · Veröffentlicht Sept. 2026
DeepSeek V4.1 Flash, the model now behind DeepSeek's deepseek-flash API, is a strong-value open-weight pick: MIT-licensed weights, image input, a 1M-token context and an Artificial Analysis score well above average for comparable models. Conditional because its standout coding-agent results are DeepSeek's own, Artificial Analysis measured it as very verbose, and peak hours double the rate. Right for cost-sensitive pipelines that can run off-peak or self-host; wrong for buyers who need independently verified agent performance before committing.
Ist es das Richtige für dich?
Gut für
- High-volume API workloads at $0.15 input and $0.60 output per 1M tokens off-peak, with off-peak cache-hit input at $0.003 per 1M
- Scheduling around peak pricing: peak rates, twice off-peak, apply only 01:00-04:00 and 06:00-10:00 UTC Monday through Friday, and all other hours bill off-peak
- Self-hosting under the MIT licence, with the weights published on Hugging Face
- Terminal-style coding agents: DeepSeek's own max-effort table reports 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, the top score on both rows, ahead of Opus-5.0's 89.1 and 74.0
- Document and image understanding: Artificial Analysis lists text and image input, and DeepSeek's base-model table reports 95.6 on DocVQA
- Long-context work within a 1M-token window and up to 384K output tokens
Nicht geeignet für
- Terminal-Bench 3.0 and 4.0: DeepSeek's own table reports 30.0 and 31.2, against Opus-5.0's 43.3 and 51.8
- Hard expert-knowledge questions: DeepSeek's table reports 36.8 on HLE, against Opus-5.0's 56.3 and GPT-5.6 Sol's 44.5
- Tight output-token budgets: Artificial Analysis measured 250M output tokens to run its Intelligence Index, against a 140M median for open-weight models of similar size
- Assuming the headline agent score carries over to your harness: DeepSeek's own scaffold table ranges from 65.5 under OpenCode to 74.2 under mini-SWE on DeepSWE v1.1
- Broad factual recall: DeepSeek's base-model table reports 42.3 on SimpleQA-Verified, against 55.2 for V4-Pro-Base
- Image generation: Artificial Analysis lists text as its only output modality
Leistung nach Aufgabe
Terminal-style agentic coding
Tops DeepSeek's own max-effort table on Terminal-Bench 2.1 (90.6) and DeepSWE v1.1 (74.2). Vendor-run figures.
Terminal-Bench 3.0 and 4.0
31.2 on Terminal-Bench 4.0 in DeepSeek's table, against 51.8 for Opus-5.0.
Cost-efficient high-volume inference
$0.15/$0.60 per 1M off-peak, but Artificial Analysis records 250M output tokens on its index run against a 140M similar-size median.
Visual document understanding
95.6 on DocVQA in DeepSeek's base-model table; image input, text output. No independent vision measurement found.
Expert-level reasoning (HLE)
36.8 on HLE, against Opus-5.0's 56.3 in DeepSeek's table.
Preise
Eingabe
$0.15 / 1M tokens
Ausgabe
$0.60 / 1M tokens
Kontext
1M ctx. Peak 2x: 01-04+06-10 UTC Mon-Fri
Benchmarks
| Benchmark | Wert | Quelle |
|---|---|---|
| Artificial Analysis Intelligence Index | 40 | Quelle |
| Terminal-Bench 2.1 (DeepSeek-run) | 90.6 | Quelle |
| Terminal-Bench 4.0 (DeepSeek-run) | 31.2 | Quelle |
| DeepSWE v1.1 (DeepSeek-run) | 74.2 | Quelle |
| GPQA Diamond (DeepSeek-run) | 90.9 | Quelle |
| HLE (DeepSeek-run) | 36.8 | Quelle |
| CyberGym (DeepSeek-run) | 88.1 | Quelle |
| Codeforces rating (DeepSeek-run) | 3471 | Quelle |
Noch keine Urteilsänderungen
Die Uhr läuft ab dem ersten Tag — Änderungen erscheinen hier, sobald sich unser Urteil weiterentwickelt.
Quellen
Prüfprotokoll
Noch keine Prüfungen
Wir haben für diesen Eintrag noch keine Prüfung erfasst. Sobald eine Prüfung läuft, erscheint ihr Verlauf hier.