DeepSeek V4.1 Flash

Open weights by DeepSeek

Tops our lineup on DeepSWE and Terminal-Bench, runs fast, and reads images. The model still has two rare quirks: a long run can loop, or a tool call can fail. If that happens, restart the run or switch to DeepSeek V4 Flash, which doesn't do this.

Price per 1M tokens
$0.15 in$0.60 out
$0.028 cache read
Context
1M
384K max output
Parameters
552B
8B to 16B active
Weights
FP4 + FP8
Official release
  • Vision
  • Tool calling
  • Reasoning optional, default high

How it ranks

Against every model we serve, best first. Starred scores are reported by the publisher on its model card; the rest are independent measurements. About the scores

Artificial Analysis Intelligence Indexby Artificial Analysis
  1. GLM 5.345
  2. Kimi K344
  3. GLM 5.3 Flash42
  4. DeepSeek V4.1 Flash39
  5. DeepSeek V4 Flash34
  6. Qwen3.6 35B A3B18
DeepSWEby Datacurve
  1. DeepSeek V4.1 Flash74.2%* reported by the publisher
  2. GLM 5.369%
  3. Kimi K368.5%
  4. GLM 5.3 Flash63.4%
  5. DeepSeek V4 Flash54.4%* reported by the publisher

Not published: Qwen3.6 35B A3B.

Terminal-Bench 2.1by tbench.ai
  1. DeepSeek V4.1 Flash90.6%* reported by the publisher
  2. Kimi K385%
  3. GLM 5.3 Flash84.3%
  4. GLM 5.383.9%
  5. DeepSeek V4 Flash78.7%
  6. Qwen3.6 35B A3B44.9%
  • Artificial Analysis Intelligence Index for DeepSeek V4.1 Flash: independent measurement. Source
  • DeepSWE for DeepSeek V4.1 Flash: reported by the publisher on its model card. Source
  • Terminal-Bench 2.1 for DeepSeek V4.1 Flash: reported by the publisher on its model card. Source

Pricing

Input
$0.15 per 1M tokens
Output
$0.60 per 1M tokens
Cache read
$0.028 per 1M tokens

See all prices

Specs

Context window
1M tokens
Max output
384K tokens
Parameters
552B backbone plus 196B of conditional memory; 8B active per token on input, 16B on output
Vision
Yes
Tool calling
Yes
Reasoning
Adjustable (none, low, high, max)
Default effort
high
In production since
September 18, 2026

Quickstart

Pick your tool; the command already carries this model's id. More tools (Pi, Copilot, Zed, omp) in the setup guides.

OpenCode
Built in
Launch OpenCode on this model with the umans CLI. Umans AI is also built into OpenCode: run /connect and pick it.
Terminal
curl -fsSL https://api.code.umans.ai/cli/install.sh | bash   # once
umans opencode --model umans-deepseek-v4.1-flash

Or consider

DeepSeek V4.1 Flash API: pricing, context and benchmarks · Umans AI