DeepSeek V4 Flash

DeepSeek V4 Flash · Reviews and evidence

Which DeepSeek V4 Flash conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

The official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.

DeepSeek API Docs · Read evidence

MindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.

MindStudio · Read evidence

Full reviews and related reading

Read the full analysis

Pricing · English

DeepSeek V4 Flash Pricing: What You Pay in 2026

DeepSeek V4 Flash pricing changed with the V4.1 migration. See the current cache, peak-hour, output, and workload cost math before you budget.

Selected evidence

OfficialVendor report

DeepSeek-V4-Flash: 0731 Benchmark Update and Harness Conditions

The official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.

SourceDeepSeek API Docs
Published2026-07-31
Collected2026-08-18
Test/source conditions
0731 API; DeepSeek Harness minimal mode; max; top_p=0.95; temperature=1.0
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCodingReasoning
Media / benchmarkIndependent measurement

DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent Testing

MindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.

SourceMindStudio
Published2026-08-01
Collected2026-08-18
Test/source conditions
Author field report; complete sample, harness, and repetition count were not disclosed
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCodingSpeed & latency
Media / benchmarkPlatform telemetry

DeepSeek V4 Flash 0731 Benchmarks, Pricing & Speed (BenchLM)

BenchLM’s 0731 snapshot lists a 1M context window, Agentic 51.9, Coding 48.5, and Knowledge 61.1, with many scores attributed back to the official report.

SourceBenchLM.ai (model benchmarking and pricing tracking site)
Published2026-07-31
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
Dated tracking snapshot; independent harness was not disclosed
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
CostReasoningSpeed & latency
CommunityIndependent measurement

I Ran DeepSeek V4 Flash on 8 Agent Harnesses (Reddit r/DeepSeek)

A Reddit author compares eight harnesses on OpenRouter across 25 automation tasks: Pi Agent passes 66.7% versus OpenCode 46.7%, with about $0.028 versus $0.073 per successful task.

SourceReddit r/DeepSeek
PublishedUnknown
Collected2026-08-17
Test/source conditions
OpenRouter; eight harnesses; 25 automation tasks; method follows the linked post
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCoding

All sources

All sources

24 / 24
OfficialVendor report

DeepSeek-V4-Flash: 0731 Benchmark Update and Harness Conditions

The official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.

SourceDeepSeek API Docs
Published2026-07-31
Collected2026-08-18
Test/source conditions
0731 API; DeepSeek Harness minimal mode; max; top_p=0.95; temperature=1.0
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCodingReasoning
Media / benchmarkIndependent measurement

DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent Testing

MindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.

SourceMindStudio
Published2026-08-01
Collected2026-08-18
Test/source conditions
Author field report; complete sample, harness, and repetition count were not disclosed
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCodingSpeed & latency
Media / benchmarkPlatform telemetry

DeepSeek V4 Flash 0731 Benchmarks, Pricing & Speed (BenchLM)

BenchLM’s 0731 snapshot lists a 1M context window, Agentic 51.9, Coding 48.5, and Knowledge 61.1, with many scores attributed back to the official report.

SourceBenchLM.ai (model benchmarking and pricing tracking site)
Published2026-07-31
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
Dated tracking snapshot; independent harness was not disclosed
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
CostReasoningSpeed & latency
CommunityIndependent measurement

I Ran DeepSeek V4 Flash on 8 Agent Harnesses (Reddit r/DeepSeek)

A Reddit author compares eight harnesses on OpenRouter across 25 automation tasks: Pi Agent passes 66.7% versus OpenCode 46.7%, with about $0.028 versus $0.073 per successful task.

SourceReddit r/DeepSeek
PublishedUnknown
Collected2026-08-17
Test/source conditions
OpenRouter; eight harnesses; 25 automation tasks; method follows the linked post
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCoding
CommunityIndependent measurement

302.AI Benchmark Laboratory | Breaking the “Lightweight” Label: A Hands-on Test of DeepSeek-V4-Flash, a Low-Cost Challenger to Top-Tier Agents

The 302.AI lab positions Flash as a low-cost Agent candidate while mixing vendor figures with its own test material; the two evidence types must stay separate.

SourceZhihu column (zhuanlan.zhihu.com)
PublishedUnknown
Collected2026-08-17
Test/source conditions
Independent lab article; full harness, sample, and repetitions remain source-specific
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
Speed & latency
Media / benchmarkCustomer case

DeepSeek V4 Alters Everything We Knew About Price-Performance Math (Lightning AI)

Lightning AI reports an April 24, 2026 snapshot of about 60+ tokens/s and 79.0% SWE-bench Verified for Flash, with 1M context and persistent tool-loop reasoning as architectural context.

SourceLightning AI blog
Published2026-04-27
Collected2026-08-17
Test/source conditions
Lightning AI internal use; Pro/Flash and April 24, 2026 pricing are separate
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
Cost
CommunityPersonal experience

DeepSeek V4 Flash at $0.112/M Output After the Price Increase (Reddit r/DeepSeek)

After the increase, a Reddit post relays InferX’s third-party quote of $0.056/M input and about $0.112/M output; this is a provider offer, not a controlled measurement.

SourceReddit r/DeepSeek
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
Source-specific client, task, and sample conditions; no unified controlled sample
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
Cost
CommunityPersonal experience

DeepSeek V4 Flash doesn't like us? (Reddit r/opencodeCLI)

A user reports that changing 10 lines consumed 28% of quota after the increase, and a commit message raised it to 32%; no token ledger is supplied.

SourceReddit r/opencodeCLI
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
OpenCode account experience after a price increase; no token ledger was provided
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCodingCost
CommunityPersonal experience

DeepSeek V4 Flash is a monster! Cheap & Good, and so fast (Reddit r/opencodeCLI)

An OpenCode user subjectively finds Flash cheap and fast enough to replace parts of a Claude/GPT workflow, while describing Pro as slower and costlier.

SourceReddit r/opencodeCLI
PublishedUnknown
Collected2026-08-17
Test/source conditions
OpenCode personal workflow; task set, parameters, and repetitions were not disclosed
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCoding
CommunityEditorial analysis

DeepSeek V4 Flash Just Drew a Pretty Brutal "Kill Line" on This Chart (Reddit r/LocalLLM)

Reddit interprets a historical Artificial Analysis chart as roughly a 50 index and about $0.03 per weighted task for Flash 0731.

SourceReddit r/LocalLLM
PublishedUnknown
Collected2026-08-17
Test/source conditions
Community chart interpretation based on an Artificial Analysis snapshot
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
Cost
Media / benchmarkEditorial analysis

DeepSeek V4 Flash Review (2026) — Specs, Tests & Speed

An unaffiliated site lists 284B/13B, 1M context, and prices as of July 25, 2026, and frames real-world consistency as “benchmark maxed.”

Sourcedeepseek.ai (an independent website, not affiliated with DeepSeek officially)
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
Public review article; test configuration and freshness require rechecking
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
Capability
Media / benchmarkPlatform telemetry

DeepSeek V4 Flash Third-Party Aggregator Data (Command Code / flaq.ai / OpenRouter / BenchLM Supplement)

Command Code, flaq, OpenRouter, and BenchLM aggregate model IDs, prices, and capability fields whose collection windows may differ.

SourceCommand Code — https://commandcode.ai/models/deepseek-v4-flash (collected on 2026-08-17)
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
Aggregated model-card snapshot; source fields may mix collection windows
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
Cost
CommunityEditorial analysis

DeepSeek V4 Logic Evaluation

The Zhihu article focuses mainly on V4 Pro/family logic results; Flash appears only as family context and cannot be treated as a Flash test.

SourceZhihu Column (zhuanlan.zhihu.com)
PublishedUnknown
Collected2026-08-17
Test/source conditions
Public article focused mainly on V4 family/Pro logic results; Flash applicability is limited
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
Reasoning
CommunityPersonal experience

DeepSeek V4-Flash/Pro Field Report: From Purchase to Practice, an Exceptional Price-to-Performance Experience with Chinese LLMs

The CSDN author documents credit purchase, API setup, code generation, and debugging while using V4-Flash and V4-Pro in one personal workflow.

SourceCSDN DeepSeek Technology Community (deepseek.csdn.net)
Published2026-05-06
Collected2026-08-17
Test/source conditions
Individual API purchase and coding report; task set, parameters, and repetitions were not fully disclosed
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCodingSpeed & latency
CommunityPersonal experience

DeepSeek-V4-Flash Hands-on Experience: How a Powerful Model Can Actually Help You Get Work Done

The Cnblogs article frames stronger reasoning, coding, and Agent work and repeats a DeepSWE 7.3→54.4 change without publishing a reproduction protocol.

Source博客园 (cnblogs.com)
Published2026-08-04
Collected2026-08-17
Test/source conditions
Source-specific client, task, and sample conditions; no unified controlled sample
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCodingSpeed & latency
CommunityIndependent measurement

Hands-on Test of DeepSeek V4's Agent Capabilities

The CowAgent author examines tool calls, long context, long-term memory, browser automation, and knowledge organization across six real scenarios.

Source博客园 (cnblogs.com)
PublishedUnknown
Collected2026-08-17
Test/source conditions
CowAgent six-scenario field test; full sample and repetitions remain source-specific
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCodingSpeed & latency
CommunityPersonal experience

Tested DeepSeek V4 Flash with Some Large Code-Change Evaluations (Reddit r/LocalLLaMA)

The author reports strong tool use and context management on large code-change evaluations, but publishes no task set, scores, version, or full traces.

SourceReddit r/LocalLLaMA
PublishedUnknown
Collected2026-08-17
Test/source conditions
Personal large-code-change observation; no unified controlled sample
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCodingReasoning
CommunityPersonal experience

Thank you, OpenCode! DeepSeek v4 Flash (Free) is just too good. (Reddit r/opencodeCLI)

An OpenCode user says the free Flash tier lasted about 5–6 hours before exhaustion and reset after roughly 7 hours.

SourceReddit r/opencodeCLI
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
OpenCode free-tier account experience; quota rules were not independently checked
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCoding
Media / benchmarkCustomer case

The Official DeepSeek V4-Flash Is Here! AI Developers Test It: "Fantastic" Pricing, Agent Capabilities Close to Top-Tier Models

Eastmoney relays developer cases: Hermes + Flash took about 40 seconds versus GPT + Codex at about 1:47, and another task used about 510K tokens and CNY 0.53.

SourceEastmoney.com (finance.eastmoney.com), source: National Business Daily
PublishedUnknown
Collected2026-08-17
Test/source conditions
Hermes/Pi Agent developer cases; tasks and timings were not controlled repeats
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
CostSpeed & latency
CommunityPersonal experience

What Is Currently the Cheapest Good Alternative to DeepSeek V4 Flash? (Reddit r/hermesagent)

A Hermes Agent user seeks a cheaper alternative after the V4 price increase, focusing on quality, speed, caching, and free quotas; the post contains no uniform candidate test.

SourceReddit r/hermesagent
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
Hermes Agent community discussion; no unified quality or cost experiment
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCoding
CommunityPersonal experience

What Is the DeepSeek “Kill Line”? A First-hand Review of DeepSeek-V4-Flash

Programmer Xiaohui explains Flash’s “kill line” with a price–capability framing; it is personal interpretation and source recap.

SourceZhihu column (zhuanlan.zhihu.com)
Published2026-08-06
Collected2026-08-17
Test/source conditions
Personal price-performance interpretation; chart window and version follow the source
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
CostReasoning
CommunityPersonal experience

X (Twitter) Field-Test Feedback Summary: Real User Voices Before and After the Price Increase

This X roundup reuses the 0424/0731 pricing feedback and adds no independent task, sample, or billing data.

SourceX (Twitter)
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
X feedback summary citing 0424/0731; it is not an independent test
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
Speed & latency
CommunityPersonal experience

X user field test: Comparing the 0424 and 0731 versions, and the felt cost after the price increase (Wyq / ch1lam)

An X user prefers 0424 and says the 0731 increase worsened the cost experience for non-coding work; this is version-specific personal feedback.

SourceX (Twitter)
Published2026-08-17
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
X field report comparing 0424 and 0731; no unified task or token ledger
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
Cost
CommunityPersonal experience

X User Reports: One Prompt Fixes Hypervisor Scheduling, Local GPUs Reverse-Engineer a Chrome Extension (jtregunna / 0xRaghuboi)

X posts report one prompt fixing hypervisor guest scheduling and another reverse-engineering a Chrome extension locally; neither provides a diff, tests, or full input.

SourceX (Twitter)
Published2026-08-17
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test/source conditions
X single-case reports; full inputs and tests were not provided
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh
AgentCoding

DeepSeek V4 Flash

Compare DeepSeek V4 Flash in Tabbit

Model access, features, and permissions depend on your current client account.