Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

DeepSeek V4 Flash · Official source · Vendor report

DeepSeek-V4-Flash: 0731 Benchmark Update and Harness Conditions

The official 0731 table reports Terminal Bench 2.1 82.7, DeepSWE 54.4, and Toolathlon Verified 70.3 under DeepSeek Harness minimal mode, max effort, top_p 0.95, and temperature 1.0.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Test/source conditions
0731 API; DeepSeek Harness minimal mode; max; top_p=0.95; temperature=1.0
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh

Key data and applicable tasks

One-sentence takeaway

The official 0731 update shows that V4-Flash scores significantly higher than V4-Pro-Preview on agent benchmarks, but the results depend on DeepSeek Harness minimal mode and max effort; the harness and parameters must be reproduced together for a meaningful comparison.

Use cases

  • Suitable tasks: Official positioning of the version's capabilities for code agents, terminal operations, tool calling, automation, and full-stack development.

  • Unsuitable tasks: Extrapolating scores from a single official harness directly to chat, vision, multimodal use, or another agent framework.

  • Applicable model version: The public beta deepseek-v4-flash with the 0731 API update; the update page says the model architecture and size remain unchanged, with only post-training redone.

  • Applicable client, agent, or API: DeepSeek API; official DeepSeek Harness minimal mode.

  • Recommended reasoning level and parameters: To reproduce the official Code Agent conditions, use max, top_p=0.95, and temperature=1.0; ordinary applications should conduct their own parameter evaluation again.

Test environment, input/configuration

  • Model: DeepSeek-V4-Flash 0731 API.

  • Harness: DeepSeek Harness minimal mode (the official explanation said it was about to be released at the time).

  • Reasoning level: max.

  • Sampling configuration: top_p=0.95, temperature=1.0.

  • Task sets: Terminal Bench 2.1, NL2Repo, CyberGym, DeepSWE, Toolathlon Verified, Agents' Last Exam, Automation Bench Public, as well as the internal DSBench-FullStack and DSBench-Hard.

  • Original inputs: The complete task inputs and harness implementation for each benchmark were not made public on the update page, so an individual test cannot be reconstructed from this page alone.

Result data

BenchmarkV4-Flash 0731
Terminal Bench 2.182.7
NL2Repo54.2
CyberGym76.7
DeepSWE54.4
Toolathlon Verified70.3
Agents' Last Exam25.2
Automation Bench (Public)25.1
DSBench-FullStack (internal)68.7
DSBench-Hard (internal)59.6

Conclusion

This update supports treating V4-Flash as a low-cost code-agent candidate, particularly for terminal work, repository modifications, and tool orchestration. However, it demonstrates the performance of the combination “V4-Flash + official harness + specified parameters,” not a uniform capability of the bare model across all clients.

Limitations

  • The official update page does not provide each question's inputs, outputs, failure samples, number of repetitions, confidence intervals, or complete harness code, so this cannot be called an independently reproducible evaluation.

  • DSBench-FullStack and DSBench-Hard are internal collections and cannot be rerun directly by external parties.

  • The update page says the comparison target is V4-Pro-Preview, but does not provide a complete comparison table in the same paragraph; do not fill in the missing comparison scores yourself.

  • The 0731 API upgraded only V4-Flash; the V4-Pro API and the APP/Web models had not changed at the time. Product-side results cannot be treated directly as API results.

Reproduction steps

  1. In the DeepSeek API, pin the model version and the available 0731 endpoint.

  2. Record the API response protocol, tool descriptions, timeouts, retries, context trimming, and workspace snapshot.

  3. Have your own agent replicate the official minimal harness as closely as possible, and fix max, top_p=0.95, and temperature=1.0.

  4. Start with multiple repetitions of publicly available tasks such as Terminal Bench 2.1 and NL2Repo, saving each task's patch, test results, number of tool calls, and token usage.

  5. Present the results alongside the official table, while also running the target production harness separately; do not compare only a single overall score.

Original evidence and data

The public figures given in the original update page are 82.7 / 54.2 / 76.7 / 54.4 / 70.3 / 25.2 / 25.1, and the two DSBench figures are marked as internal tests; the same page explicitly states DeepSeek Harness minimal mode, max effort, topp=0.95, and temperature=1.0.

Source excerpt or observation (compliance short quote only)

The official announcement calls this update “significantly enhanced agent capabilities,” while also noting that Flash 0731 “keeps the same model architecture and size ... and was only re-post-trained.”

What this supports

  • Supports repeating the 0731 Agent benchmark claims under the stated official harness and separating public tasks from internal DSBench.

What this does not support

  • Does not support treating vendor scores as independent retests or extending them to chat, other harnesses, or undisclosed Pro comparisons.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

DeepSeek API Docs · DeepSeek official · Original publication date 2026-07-31 · Site edit date 2026-09-20

Open original source

DeepSeek V4 Flash

Compare DeepSeek V4 Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

DeepSeek V4 Flash Pricing: What You Pay in 2026

DeepSeek V4 Flash pricing changed with the V4.1 migration. See the current cache, peak-hour, output, and workload cost math before you budget.

Related reviews

Tested DeepSeek V4 Flash with Some Large Code-Change Evaluations (Reddit r/LocalLLaMA)The author reports strong tool use and context management on large code-change evaluations, but publishes no task set, scores, version, or full traces.DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent TestingMindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.I Ran DeepSeek V4 Flash on 8 Agent Harnesses (Reddit r/DeepSeek)A Reddit author compares eight harnesses on OpenRouter across 25 automation tasks: Pi Agent passes 66.7% versus OpenCode 46.7%, with about $0.028 versus $0.073 per successful task.DeepSeek V4 Flash doesn't like us? (Reddit r/opencodeCLI)A user reports that changing 10 lines consumed 28% of quota after the increase, and a commit message raised it to 32%; no token ledger is supplied.Connect DeepSeek to Codex with the official configurationBack up local configuration, use the official script or minimal provider fields, and verify with a reversible task.Configure reasoning tiers and continue tool callsTurn low/high/max, tool results, and reasoning_content handoff into a checkable integration path.Delegate in layers and synthesize a monograph with DSHThe source publishes a complete starting prompt with research, pushback, editing, and final synthesis targeting one cited Markdown artifact.Choose a DeepSeek-to-Codex integration pathA case roundup comparing integration routes; it is not one complete copyable template.