DeepSeek V4.1 Flash
Tops our lineup on DeepSWE and Terminal-Bench, runs fast, and reads images. The model still has two rare quirks: a long run can loop, or a tool call can fail. If that happens, restart the run or switch to DeepSeek V4 Flash, which doesn't do this.
- Price per 1M tokens
- $0.15 in$0.60 out
- $0.028 cache read
- Context
- 1M
- 384K max output
- Parameters
- 552B
- 8B to 16B active
- Vision
- Tool calling
- Reasoning optional, default high
How it ranks
Against every model we serve, best first. Starred scores are reported by the publisher on its model card; the rest are independent measurements. About the scores
- GLM 5.345
- Kimi K344
- GLM 5.3 Flash42
- DeepSeek V4.1 Flash39
- DeepSeek V4 Flash34
- Qwen3.6 35B A3B18
- DeepSeek V4.1 Flash74.2%* reported by the publisher
- GLM 5.369%
- Kimi K368.5%
- GLM 5.3 Flash63.4%
- DeepSeek V4 Flash54.4%* reported by the publisher
Not published: Qwen3.6 35B A3B.
- DeepSeek V4.1 Flash90.6%* reported by the publisher
- Kimi K385%
- GLM 5.3 Flash84.3%
- GLM 5.383.9%
- DeepSeek V4 Flash78.7%
- Qwen3.6 35B A3B44.9%
Pricing
- Input
- $0.15 per 1M tokens
- Output
- $0.60 per 1M tokens
- Cache read
- $0.028 per 1M tokens
Specs
- Context window
- 1M tokens
- Max output
- 384K tokens
- Parameters
- 552B backbone plus 196B of conditional memory; 8B active per token on input, 16B on output
- Vision
- Yes
- Tool calling
- Yes
- Reasoning
- Adjustable (none, low, high, max)
- Default effort
- high
- In production since
- September 18, 2026
Quickstart
Pick your tool; the command already carries this model's id. More tools (Pi, Copilot, Zed, omp) in the setup guides.
OpenCode
Built inLaunch OpenCode on this model with the umans CLI. Umans AI is also built into OpenCode: run /connect and pick it.
curl -fsSL https://api.code.umans.ai/cli/install.sh | bash # once
umans opencode --model umans-deepseek-v4.1-flash