The official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.
DeepSeek API Docs · Read evidenceDeepSeek V4 Flash · Reviews and evidence
Which DeepSeek V4 Flash conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
MindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.
MindStudio · Read evidenceBenchLM’s 0731 snapshot lists a 1M context window, Agentic 51.9, Coding 48.5, and Knowledge 61.1, with many scores attributed back to the official report.
BenchLM.ai (model benchmarking and pricing tracking site) · Read evidenceFull reviews and related reading
Selected evidence
DeepSeek-V4-Flash: 0731 Benchmark Update and Harness Conditions
The official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.
- Test/source conditions
- 0731 API; DeepSeek Harness minimal mode; max; top_p=0.95; temperature=1.0
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent Testing
MindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.
- Test/source conditions
- Author field report; complete sample, harness, and repetition count were not disclosed
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Flash 0731 Benchmarks, Pricing & Speed (BenchLM)
BenchLM’s 0731 snapshot lists a 1M context window, Agentic 51.9, Coding 48.5, and Knowledge 61.1, with many scores attributed back to the official report.
Unverified: the original source could not be rechecked.
- Test/source conditions
- Dated tracking snapshot; independent harness was not disclosed
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
I Ran DeepSeek V4 Flash on 8 Agent Harnesses (Reddit r/DeepSeek)
A Reddit author compares eight harnesses on OpenRouter across 25 automation tasks: Pi Agent passes 66.7% versus OpenCode 46.7%, with about $0.028 versus $0.073 per successful task.
- Test/source conditions
- OpenRouter; eight harnesses; 25 automation tasks; method follows the linked post
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
All sources
All sources
DeepSeek-V4-Flash: 0731 Benchmark Update and Harness Conditions
The official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.
- Test/source conditions
- 0731 API; DeepSeek Harness minimal mode; max; top_p=0.95; temperature=1.0
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent Testing
MindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.
- Test/source conditions
- Author field report; complete sample, harness, and repetition count were not disclosed
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Flash 0731 Benchmarks, Pricing & Speed (BenchLM)
BenchLM’s 0731 snapshot lists a 1M context window, Agentic 51.9, Coding 48.5, and Knowledge 61.1, with many scores attributed back to the official report.
Unverified: the original source could not be rechecked.
- Test/source conditions
- Dated tracking snapshot; independent harness was not disclosed
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
I Ran DeepSeek V4 Flash on 8 Agent Harnesses (Reddit r/DeepSeek)
A Reddit author compares eight harnesses on OpenRouter across 25 automation tasks: Pi Agent passes 66.7% versus OpenCode 46.7%, with about $0.028 versus $0.073 per successful task.
- Test/source conditions
- OpenRouter; eight harnesses; 25 automation tasks; method follows the linked post
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
302.AI Benchmark Laboratory | Breaking the “Lightweight” Label: A Hands-on Test of DeepSeek-V4-Flash, a Low-Cost Challenger to Top-Tier Agents
The 302.AI lab positions Flash as a low-cost Agent candidate while mixing vendor figures with its own test material; the two evidence types must stay separate.
- Test/source conditions
- Independent lab article; full harness, sample, and repetitions remain source-specific
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Alters Everything We Knew About Price-Performance Math (Lightning AI)
Lightning AI reports an April 24, 2026 snapshot of about 60+ tokens/s and 79.0% SWE-bench Verified for Flash, with 1M context and persistent tool-loop reasoning as architectural context.
- Test/source conditions
- Lightning AI internal use; Pro/Flash and April 24, 2026 pricing are separate
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Flash at $0.112/M Output After the Price Increase (Reddit r/DeepSeek)
After the increase, a Reddit post relays InferX’s third-party quote of $0.056/M input and about $0.112/M output; this is a provider offer, not a controlled measurement.
Unverified: the original source could not be rechecked.
- Test/source conditions
- Source-specific client, task, and sample conditions; no unified controlled sample
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Flash doesn't like us? (Reddit r/opencodeCLI)
A user reports that changing 10 lines consumed 28% of quota after the increase, and a commit message raised it to 32%; no token ledger is supplied.
Unverified: the original source could not be rechecked.
- Test/source conditions
- OpenCode account experience after a price increase; no token ledger was provided
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Flash is a monster! Cheap & Good, and so fast (Reddit r/opencodeCLI)
An OpenCode user subjectively finds Flash cheap and fast enough to replace parts of a Claude/GPT workflow, while describing Pro as slower and costlier.
- Test/source conditions
- OpenCode personal workflow; task set, parameters, and repetitions were not disclosed
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Flash Just Drew a Pretty Brutal "Kill Line" on This Chart (Reddit r/LocalLLM)
Reddit interprets a historical Artificial Analysis chart as roughly a 50 index and about $0.03 per weighted task for Flash 0731.
- Test/source conditions
- Community chart interpretation based on an Artificial Analysis snapshot
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Flash Review (2026) — Specs, Tests & Speed
An unaffiliated site lists 284B/13B, 1M context, and prices as of July 25, 2026, and frames real-world consistency as “benchmark maxed.”
Unverified: the original source could not be rechecked.
- Test/source conditions
- Public review article; test configuration and freshness require rechecking
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Flash Third-Party Aggregator Data (Command Code / flaq.ai / OpenRouter / BenchLM Supplement)
Command Code, flaq, OpenRouter, and BenchLM aggregate model IDs, prices, and capability fields whose collection windows may differ.
Unverified: the original source could not be rechecked.
- Test/source conditions
- Aggregated model-card snapshot; source fields may mix collection windows
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Logic Evaluation
The Zhihu article focuses mainly on V4 Pro/family logic results; Flash appears only as family context and cannot be treated as a Flash test.
- Test/source conditions
- Public article focused mainly on V4 family/Pro logic results; Flash applicability is limited
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4-Flash/Pro Field Report: From Purchase to Practice, an Exceptional Price-to-Performance Experience with Chinese LLMs
The CSDN author documents credit purchase, API setup, code generation, and debugging while using V4-Flash and V4-Pro in one personal workflow.
- Test/source conditions
- Individual API purchase and coding report; task set, parameters, and repetitions were not fully disclosed
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek-V4-Flash Hands-on Experience: How a Powerful Model Can Actually Help You Get Work Done
The Cnblogs article frames stronger reasoning, coding, and Agent work and repeats a DeepSWE 7.3→54.4 change without publishing a reproduction protocol.
- Test/source conditions
- Source-specific client, task, and sample conditions; no unified controlled sample
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
Hands-on Test of DeepSeek V4's Agent Capabilities
The CowAgent author examines tool calls, long context, long-term memory, browser automation, and knowledge organization across six real scenarios.
- Test/source conditions
- CowAgent six-scenario field test; full sample and repetitions remain source-specific
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
Tested DeepSeek V4 Flash with Some Large Code-Change Evaluations (Reddit r/LocalLLaMA)
The author reports strong tool use and context management on large code-change evaluations, but publishes no task set, scores, version, or full traces.
- Test/source conditions
- Personal large-code-change observation; no unified controlled sample
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
Thank you, OpenCode! DeepSeek v4 Flash (Free) is just too good. (Reddit r/opencodeCLI)
An OpenCode user says the free Flash tier lasted about 5–6 hours before exhaustion and reset after roughly 7 hours.
Unverified: the original source could not be rechecked.
- Test/source conditions
- OpenCode free-tier account experience; quota rules were not independently checked
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
The Official DeepSeek V4-Flash Is Here! AI Developers Test It: "Fantastic" Pricing, Agent Capabilities Close to Top-Tier Models
Eastmoney relays developer cases: Hermes + Flash took about 40 seconds versus GPT + Codex at about 1:47, and another task used about 510K tokens and CNY 0.53.
- Test/source conditions
- Hermes/Pi Agent developer cases; tasks and timings were not controlled repeats
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
What Is Currently the Cheapest Good Alternative to DeepSeek V4 Flash? (Reddit r/hermesagent)
A Hermes Agent user seeks a cheaper alternative after the V4 price increase, focusing on quality, speed, caching, and free quotas; the post contains no uniform candidate test.
Unverified: the original source could not be rechecked.
- Test/source conditions
- Hermes Agent community discussion; no unified quality or cost experiment
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
What Is the DeepSeek “Kill Line”? A First-hand Review of DeepSeek-V4-Flash
Programmer Xiaohui explains Flash’s “kill line” with a price–capability framing; it is personal interpretation and source recap.
- Test/source conditions
- Personal price-performance interpretation; chart window and version follow the source
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
X (Twitter) Field-Test Feedback Summary: Real User Voices Before and After the Price Increase
This X roundup reuses the 0424/0731 pricing feedback and adds no independent task, sample, or billing data.
Unverified: the original source could not be rechecked.
- Test/source conditions
- X feedback summary citing 0424/0731; it is not an independent test
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
X user field test: Comparing the 0424 and 0731 versions, and the felt cost after the price increase (Wyq / ch1lam)
An X user prefers 0424 and says the 0731 increase worsened the cost experience for non-coding work; this is version-specific personal feedback.
Unverified: the original source could not be rechecked.
- Test/source conditions
- X field report comparing 0424 and 0731; no unified task or token ledger
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
X User Reports: One Prompt Fixes Hypervisor Scheduling, Local GPUs Reverse-Engineer a Chrome Extension (jtregunna / 0xRaghuboi)
X posts report one prompt fixing hypervisor guest scheduling and another reverse-engineering a Chrome extension locally; neither provides a diff, tests, or full input.
Unverified: the original source could not be rechecked.
- Test/source conditions
- X single-case reports; full inputs and tests were not provided
- Model version
- DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
- Collection boundary
- Existing source note collected August 17–18, 2026; dynamic facts require refresh
DeepSeek V4 Flash
Compare DeepSeek V4 Flash in Tabbit
Model access, features, and permissions depend on your current client account.