DeepSeek V4 Flash · Media / benchmark · Platform telemetry
BenchLM’s 0731 snapshot lists a 1M context window, Agentic 51.9, Coding 48.5, and Knowledge 61.1, with many scores attributed back to the official report.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
API model id: deepseek-v4-flash (0731 update)
Release: July 31, 2026; type: Proprietary / Reasoning
Context window: 1M; input modality: text; output modality: text
Status: active; published Prompt caching price: $0.003 / million cached input tokens
| Category | Weight | Score | Ranking |
|---|---|---|---|
| Agentic | 22% | 51.9 | Unranked |
| Coding | 20% | 48.5 | Unranked |
| Reasoning | 17% | Pending | Unranked |
| Knowledge | 12% | 61.1 | #42 of 57 (27th percentile) |
| Math | 5% | 81.4 | Unranked |
| Multilingual | 7% | Not tested | - |
| Multimodal | 12% | Not tested | - |
| Inst. Following | 5% | Not tested | - |
Coding
SWE-bench Verified: 79% (best 96%, Claude Opus 5)
SWE-bench Pro: 52.6%
LiveCodeBench Pass@1-COT: 91.6% (just 1.9 points behind V4 Pro 0813's 93.5%)
Codeforces: 3052.0
SWE Multilingual: 73.3%
Terminal-Bench 2.0: 56.9%; Terminal-Bench 2.1: 82.7%
NL2Repo: 54.2%; deepSwe: 54.4%; DSBench-FullStack: 68.7%; DSBench-Hard: 59.6%
Agentic
Terminal-Bench 2.0: 56.9%; Terminal-Bench 2.1: 82.7%
BrowseComp: 73.2%
HLE w/ tools: 45.1%
MCP Atlas: 69%; Toolathlon: 47.8%; Toolathlon-Verified: 70.3%
CyberGym: 76.7%
Agents' Last Exam: 25.2%
AutomationBench: 25.1%
Reasoning
MRCR 1M: 78.7%; CorpusQA 1M: 60.5%
Knowledge
HLE (Humanity's Last Exam): 34.8%
There is still a clear gap versus the strongest benchmark records (for example, it trails Claude Opus 5 by 17 points on SWE-bench Verified), but with a far lower price, its results in coding and agentic scenarios are usable
The 0731 update shows a significant improvement over the earlier version (deepSwe 7.3 → 54.4 on comparable data)
BenchLM note: Missing fields remain "not publicly disclosed" rather than being hidden; scores are shown only when supported by public evidence that can be displayed
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
BenchLM.ai (model benchmarking and pricing tracking site) · Author not disclosed · Original publication date 2026-07-31 · Site edit date 2026-09-20
Open original sourceDeepSeek V4 Flash
Download the Tabbit client to check model access