DeepSeek API DocsVendor report
DeepSeek-V4-Flash: 0731 Benchmark Update and Harness Conditions
The official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.
- Evidence
- Vendor report
- Boundary
- Does not support treating vendor scores as independent retests or extending them to chat, other harnesses, or undisclosed Pro comparisons.
MindStudioIndependent measurement
DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent Testing
MindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.
- Evidence
- Independent measurement
- Boundary
- Does not support treating one author’s hardware observation as universal throughput, minimum hardware, or production success rate.
BenchLM.ai (model benchmarking and pricing tracking site)Platform telemetry
DeepSeek V4 Flash 0731 Benchmarks, Pricing & Speed (BenchLM)
BenchLM’s 0731 snapshot lists a 1M context window, Agentic 51.9, Coding 48.5, and Knowledge 61.1, with many scores attributed back to the official report.
- Evidence
- Platform telemetry
- Boundary
- Does not support treating an aggregator ledger as independent measurement or turning snapshot pricing or rank into a current fact.