Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

DeepSeek V4 Flash · Media / benchmark · Customer case

The Official DeepSeek V4-Flash Is Here! AI Developers Test It: "Fantastic" Pricing, Agent Capabilities Close to Top-Tier Models

Eastmoney relays developer cases: Hermes + Flash took about 40 seconds versus GPT + Codex at about 1:47, and another task used about 510K tokens and CNY 0.53.

Media / benchmarkCustomer caseEdited 2026-09-20

Test conditions

Test/source conditions
Hermes/Pi Agent developer cases; tasks and timings were not controlled repeats
Model version
DeepSeek V4 Flash; do not merge V4 Pro, 0424, 0731, or reasoning tiers unless the source explicitly does so
Collection boundary
Existing source note collected August 17–18, 2026; dynamic facts require refresh

Key data and applicable tasks

Core facts

On July 31, DeepSeek announced that the official DeepSeek-V4-Flash API had entered public beta, with a major upgrade to its agent capabilities. The official version's benchmark scores were far higher than those of the V4-Pro preview. In the ultimate agent test, the official V4-Flash scored 25.2 points, close to Opus-4.8's 25.7 and well above the V4-Pro preview's 15.8.

Developer field-test feedback (interviewees using pseudonyms)

Zhang Ze (an AI developer who uses codex and hermes extensively):

  • Integrated V4-Flash with Hermes: on the same task with the same prompt, the official V4-Flash + Hermes finished in about 40 seconds, while GPT-5.6 Sol + Codex took 1 minute 47 seconds

  • Acknowledged that Codex's output quality is indeed higher, but said the official V4-Flash delivers work at a usable level, "above the passing line"

  • Its accuracy in selecting and calling skills is generally good. It can adjust subsequent steps based on tool-returned results, and its combinations can already smoothly meet his personal needs

Sun Tao (feedback from August 1):

  • "Among Chinese models, it is somewhat better than GLM 5.2, but it is still somewhat behind GPT 5.6"

  • His primary tools are Codex and Pi Agent; integrating V4-Flash with Pi Agent provided a good experience

  • Because it is not multimodal, its main uses are reasoning-heavy work, code, and long documents

Cost comparison

  • GPT-5.6 Sol: $5 per million input tokens, $0.5 for cached input, and $30 for output

  • Official V4-Flash: 1 yuan for input, 0.02 yuan for cached input, and 2 yuan for output (RMB)

  • The cost gap ranges from dozens to hundreds of times

  • Zhang Ze's field test: V4-Flash connected to Hermes performed a task involving local file reading and aggregation of information retrieved from across the web, using about 510,000 tokens and consuming 0.53 yuan

Peak/off-peak pricing mechanism

In an email sent at the end of June, DeepSeek previewed the API's first introduction of a peak/off-peak pricing mechanism:

  • V4-Pro: for one million input tokens (cache hit), the price drops from 1 yuan in the April preview to 0.025 yuan during regular periods and 0.05 yuan during peak periods in the official version

  • Official V4-Flash: during peak periods, one million cache-miss input tokens and output tokens are actually twice as expensive as in the April preview, at 2 yuan and 4 yuan, respectively

  • The specific implementation time for the peak/off-peak pricing mechanism has not yet been formally announced

Market impact

  • OpenRouter data from July 28 showed Chinese AI models leading the world in weekly call volume: Xiaomi MiMo-V2.5, DeepSeek V4-Flash, Tencent HY3, Zhipu GLM-5.2, and DeepSeek V4-Pro (two DeepSeek models made the top five)

  • In the 28 days through July 26: Chinese models held about 63.5% of the market, while US models held 35.5%

  • DeepSeek Harness to be announced soon: DeepSeek increased its investment in Harness during the first half of the year (on June 21, Cui Tianyi posted that the Harness team was still short-staffed). CITIC Securities believes Harness's core value is solving the pain points of long-horizon tasks and connecting models with broad workplace scenarios. AI competition is shifting from models to task delivery

What this supports

  • Supports the observation that harness and pricing affect practical delivery experience.

What this does not support

  • Does not support controlled latency, quality, or current-price comparisons; the cases mix models and tools.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Eastmoney.com (finance.eastmoney.com), source: National Business Daily · Author not disclosed · Original publication date Unknown · Site edit date 2026-09-20

Open original source

DeepSeek V4 Flash

Compare DeepSeek V4 Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Pricing · English

DeepSeek V4 Flash Pricing: What You Pay in 2026

DeepSeek V4 Flash pricing changed with the V4.1 migration. See the current cache, peak-hour, output, and workload cost math before you budget.

Related reviews

DeepSeek V4 Flash 0731 Benchmarks, Pricing & Speed (BenchLM)BenchLM’s 0731 snapshot lists a 1M context window, Agentic 51.9, Coding 48.5, and Knowledge 61.1, with many scores attributed back to the official report.DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent TestingMindStudio reports about 25–30 tokens/s on dual DGX Spark, estimates 168GB for 4-bit and 110GB for 3-bit, and still records live-data errors in two small projects.DeepSeek V4 Alters Everything We Knew About Price-Performance Math (Lightning AI)Lightning AI reports an April 24, 2026 snapshot of about 60+ tokens/s and 79.0% SWE-bench Verified for Flash, with 1M context and persistent tool-loop reasoning as architectural context.DeepSeek V4 Flash Third-Party Aggregator Data (Command Code / flaq.ai / OpenRouter / BenchLM Supplement)Command Code, flaq, OpenRouter, and BenchLM aggregate model IDs, prices, and capability fields whose collection windows may differ.Configure reasoning tiers and continue tool callsTurn low/high/max, tool results, and reasoning_content handoff into a checkable integration path.Connect DeepSeek to Codex with the official configurationBack up local configuration, use the official script or minimal provider fields, and verify with a reversible task.Delegate in layers and synthesize a monograph with DSHThe source publishes a complete starting prompt with research, pushback, editing, and final synthesis targeting one cited Markdown artifact.Structure DeepSeek tasks with the CRISPE frameworkTurn role, request, context, constraints, style, and experiments into an explicit task contract.