D

DeepSeek V4.1 Flash

DeepSeek · Released Sep 2026

Conditional

DeepSeek V4.1 Flash, the model now behind DeepSeek's deepseek-flash API, is a strong-value open-weight pick: MIT-licensed weights, image input, a 1M-token context and an Artificial Analysis score well above average for comparable models. Conditional because its standout coding-agent results are DeepSeek's own, Artificial Analysis measured it as very verbose, and peak hours double the rate. Right for cost-sensitive pipelines that can run off-peak or self-host; wrong for buyers who need independently verified agent performance before committing.

Is it right for you?

Good for

  • High-volume API workloads at $0.15 input and $0.60 output per 1M tokens off-peak, with off-peak cache-hit input at $0.003 per 1M
  • Scheduling around peak pricing: peak rates, twice off-peak, apply only 01:00-04:00 and 06:00-10:00 UTC Monday through Friday, and all other hours bill off-peak
  • Self-hosting under the MIT licence, with the weights published on Hugging Face
  • Terminal-style coding agents: DeepSeek's own max-effort table reports 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, the top score on both rows, ahead of Opus-5.0's 89.1 and 74.0
  • Document and image understanding: Artificial Analysis lists text and image input, and DeepSeek's base-model table reports 95.6 on DocVQA
  • Long-context work within a 1M-token window and up to 384K output tokens

Not good for

  • Terminal-Bench 3.0 and 4.0: DeepSeek's own table reports 30.0 and 31.2, against Opus-5.0's 43.3 and 51.8
  • Hard expert-knowledge questions: DeepSeek's table reports 36.8 on HLE, against Opus-5.0's 56.3 and GPT-5.6 Sol's 44.5
  • Tight output-token budgets: Artificial Analysis measured 250M output tokens to run its Intelligence Index, against a 140M median for open-weight models of similar size
  • Assuming the headline agent score carries over to your harness: DeepSeek's own scaffold table ranges from 65.5 under OpenCode to 74.2 under mini-SWE on DeepSWE v1.1
  • Broad factual recall: DeepSeek's base-model table reports 42.3 on SimpleQA-Verified, against 55.2 for V4-Pro-Base
  • Image generation: Artificial Analysis lists text as its only output modality

How it performs by task

Terminal-style agentic coding

Very Good

Tops DeepSeek's own max-effort table on Terminal-Bench 2.1 (90.6) and DeepSWE v1.1 (74.2). Vendor-run figures.

Terminal-Bench 3.0 and 4.0

Fair

31.2 on Terminal-Bench 4.0 in DeepSeek's table, against 51.8 for Opus-5.0.

Cost-efficient high-volume inference

Very Good

$0.15/$0.60 per 1M off-peak, but Artificial Analysis records 250M output tokens on its index run against a 140M similar-size median.

Visual document understanding

Good

95.6 on DocVQA in DeepSeek's base-model table; image input, text output. No independent vision measurement found.

Expert-level reasoning (HLE)

Fair

36.8 on HLE, against Opus-5.0's 56.3 in DeepSeek's table.

Pricing

Input

$0.15 / 1M tokens

Output

$0.60 / 1M tokens

Context

1M ctx. Peak 2x: 01-04+06-10 UTC Mon-Fri

View full pricing

Benchmarks

BenchmarkScoreSource
Artificial Analysis Intelligence Index40 Source
Terminal-Bench 2.1 (DeepSeek-run)90.6 Source
Terminal-Bench 4.0 (DeepSeek-run)31.2 Source
DeepSWE v1.1 (DeepSeek-run)74.2 Source
GPQA Diamond (DeepSeek-run)90.9 Source
HLE (DeepSeek-run)36.8 Source
CyberGym (DeepSeek-run)88.1 Source
Codeforces rating (DeepSeek-run)3471 Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

No verification checks yet

We haven't logged a verification check for this entry. Once a check runs, its history shows here.

How we evaluate