Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

DeepSeek V4 Flash · Media / benchmark · Platform telemetry

DeepSeek V4 Flash 0731 Benchmarks, Pricing & Speed (BenchLM)

BenchLM’s 0731 snapshot lists a 1M context window, Agentic 51.9, Coding 48.5, and Knowledge 61.1, with many scores attributed back to the official report.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkPlatform telemetryEdited 2026-09-20

Test conditions

Test/source conditions
Dated tracking snapshot; independent harness was not disclosed
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh

Key data and applicable tasks

Model information

  • API model id: deepseek-v4-flash (0731 update)

  • Release: July 31, 2026; type: Proprietary / Reasoning

  • Context window: 1M; input modality: text; output modality: text

  • Status: active; published Prompt caching price: $0.003 / million cached input tokens

Category scores and rankings

CategoryWeightScoreRanking
Agentic22%51.9Unranked
Coding20%48.5Unranked
Reasoning17%PendingUnranked
Knowledge12%61.1#42 of 57 (27th percentile)
Math5%81.4Unranked
Multilingual7%Not tested-
Multimodal12%Not tested-
Inst. Following5%Not tested-

Key benchmark results (official technical report / 0731 update)

Coding

  • SWE-bench Verified: 79% (best 96%, Claude Opus 5)

  • SWE-bench Pro: 52.6%

  • LiveCodeBench Pass@1-COT: 91.6% (just 1.9 points behind V4 Pro 0813's 93.5%)

  • Codeforces: 3052.0

  • SWE Multilingual: 73.3%

  • Terminal-Bench 2.0: 56.9%; Terminal-Bench 2.1: 82.7%

  • NL2Repo: 54.2%; deepSwe: 54.4%; DSBench-FullStack: 68.7%; DSBench-Hard: 59.6%

Agentic

  • Terminal-Bench 2.0: 56.9%; Terminal-Bench 2.1: 82.7%

  • BrowseComp: 73.2%

  • HLE w/ tools: 45.1%

  • MCP Atlas: 69%; Toolathlon: 47.8%; Toolathlon-Verified: 70.3%

  • CyberGym: 76.7%

  • Agents' Last Exam: 25.2%

  • AutomationBench: 25.1%

Reasoning

  • MRCR 1M: 78.7%; CorpusQA 1M: 60.5%

Knowledge

  • HLE (Humanity's Last Exam): 34.8%

Key takeaways

  • There is still a clear gap versus the strongest benchmark records (for example, it trails Claude Opus 5 by 17 points on SWE-bench Verified), but with a far lower price, its results in coding and agentic scenarios are usable

  • The 0731 update shows a significant improvement over the earlier version (deepSwe 7.3 → 54.4 on comparable data)

  • BenchLM note: Missing fields remain "not publicly disclosed" rather than being hidden; scores are shown only when supported by public evidence that can be displayed

What this supports

  • Supports tracing the model ID, category scores, and missing fields in the August 17, 2026 snapshot.

What this does not support

  • Does not support treating an aggregator ledger as independent measurement or turning snapshot pricing or rank into a current fact.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

BenchLM.ai (model benchmarking and pricing tracking site) · Author not disclosed · Original publication date 2026-07-31 · Site edit date 2026-09-20

Open original source

DeepSeek V4 Flash

Compare DeepSeek V4 Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

DeepSeek V4 Flash Pricing: What You Pay in 2026

DeepSeek V4 Flash pricing changed with the V4.1 migration. See the current cache, peak-hour, output, and workload cost math before you budget.

Related reviews

The Official DeepSeek V4-Flash Is Here! AI Developers Test It: "Fantastic" Pricing, Agent Capabilities Close to Top-Tier ModelsEastmoney relays developer cases: Hermes + Flash took about 40 seconds versus GPT + Codex at about 1:47, and another task used about 510K tokens and CNY 0.53.What Is the DeepSeek “Kill Line”? A First-hand Review of DeepSeek-V4-FlashProgrammer Xiaohui explains Flash’s “kill line” with a price–capability framing; it is personal interpretation and source recap.DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent TestingMindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.DeepSeek-V4-Flash: 0731 Benchmark Update and Harness ConditionsThe official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.Configure reasoning tiers and continue tool callsTurn low/high/max, tool results, and reasoning_content handoff into a checkable integration path.Connect DeepSeek to Codex with the official configurationBack up local configuration, use the official script or minimal provider fields, and verify with a reversible task.Delegate in layers and synthesize a monograph with DSHThe source publishes a complete starting prompt with research, pushback, editing, and final synthesis targeting one cited Markdown artifact.Structure DeepSeek tasks with the CRISPE frameworkTurn role, request, context, constraints, style, and experiments into an explicit task contract.