API model id: deepseek-v4-flash (0731 update)
Release: July 31, 2026; type: Proprietary / Reasoning
Context window: 1M; input modality: text; output modality: text
Status: active; published Prompt caching price: $0.003 / million cached input tokens
| Category | Weight | Score | Ranking |
|---|---|---|---|
| Agentic | 22% | 51.9 | Unranked |
| Coding | 20% | 48.5 | Unranked |
| Reasoning | 17% | Pending | Unranked |
| Knowledge | 12% | 61.1 | #42 of 57 (27th percentile) |
| Math | 5% | 81.4 | Unranked |
| Multilingual | 7% | Not tested | - |
| Multimodal | 12% | Not tested | - |
| Inst. Following | 5% | Not tested | - |
Coding
SWE-bench Verified: 79% (best 96%, Claude Opus 5)
SWE-bench Pro: 52.6%
LiveCodeBench Pass@1-COT: 91.6% (just 1.9 points behind V4 Pro 0813's 93.5%)
Codeforces: 3052.0
SWE Multilingual: 73.3%
Terminal-Bench 2.0: 56.9%; Terminal-Bench 2.1: 82.7%
NL2Repo: 54.2%; deepSwe: 54.4%; DSBench-FullStack: 68.7%; DSBench-Hard: 59.6%
Agentic
Terminal-Bench 2.0: 56.9%; Terminal-Bench 2.1: 82.7%
BrowseComp: 73.2%
HLE w/ tools: 45.1%
MCP Atlas: 69%; Toolathlon: 47.8%; Toolathlon-Verified: 70.3%
CyberGym: 76.7%
Agents' Last Exam: 25.2%
AutomationBench: 25.1%
Reasoning
MRCR 1M: 78.7%; CorpusQA 1M: 60.5%
Knowledge
HLE (Humanity's Last Exam): 34.8%
There is still a clear gap versus the strongest benchmark records (for example, it trails Claude Opus 5 by 17 points on SWE-bench Verified), but with a far lower price, its results in coding and agentic scenarios are usable
The 0731 update shows a significant improvement over the earlier version (deepSwe 7.3 → 54.4 on comparable data)
BenchLM note: Missing fields remain "not publicly disclosed" rather than being hidden; scores are shown only when supported by public evidence that can be displayed
DeepSeek V4 Flash