DeepSeek V4 Flash · Media / benchmark · Customer case
Eastmoney relays developer cases: Hermes + Flash took about 40 seconds versus GPT + Codex at about 1:47, and another task used about 510K tokens and CNY 0.53.
On July 31, DeepSeek announced that the official DeepSeek-V4-Flash API had entered public beta, with a major upgrade to its agent capabilities. The official version's benchmark scores were far higher than those of the V4-Pro preview. In the ultimate agent test, the official V4-Flash scored 25.2 points, close to Opus-4.8's 25.7 and well above the V4-Pro preview's 15.8.
Zhang Ze (an AI developer who uses codex and hermes extensively):
Integrated V4-Flash with Hermes: on the same task with the same prompt, the official V4-Flash + Hermes finished in about 40 seconds, while GPT-5.6 Sol + Codex took 1 minute 47 seconds
Acknowledged that Codex's output quality is indeed higher, but said the official V4-Flash delivers work at a usable level, "above the passing line"
Its accuracy in selecting and calling skills is generally good. It can adjust subsequent steps based on tool-returned results, and its combinations can already smoothly meet his personal needs
Sun Tao (feedback from August 1):
"Among Chinese models, it is somewhat better than GLM 5.2, but it is still somewhat behind GPT 5.6"
His primary tools are Codex and Pi Agent; integrating V4-Flash with Pi Agent provided a good experience
Because it is not multimodal, its main uses are reasoning-heavy work, code, and long documents
GPT-5.6 Sol: $5 per million input tokens, $0.5 for cached input, and $30 for output
Official V4-Flash: 1 yuan for input, 0.02 yuan for cached input, and 2 yuan for output (RMB)
The cost gap ranges from dozens to hundreds of times
Zhang Ze's field test: V4-Flash connected to Hermes performed a task involving local file reading and aggregation of information retrieved from across the web, using about 510,000 tokens and consuming 0.53 yuan
In an email sent at the end of June, DeepSeek previewed the API's first introduction of a peak/off-peak pricing mechanism:
V4-Pro: for one million input tokens (cache hit), the price drops from 1 yuan in the April preview to 0.025 yuan during regular periods and 0.05 yuan during peak periods in the official version
Official V4-Flash: during peak periods, one million cache-miss input tokens and output tokens are actually twice as expensive as in the April preview, at 2 yuan and 4 yuan, respectively
The specific implementation time for the peak/off-peak pricing mechanism has not yet been formally announced
OpenRouter data from July 28 showed Chinese AI models leading the world in weekly call volume: Xiaomi MiMo-V2.5, DeepSeek V4-Flash, Tencent HY3, Zhipu GLM-5.2, and DeepSeek V4-Pro (two DeepSeek models made the top five)
In the 28 days through July 26: Chinese models held about 63.5% of the market, while US models held 35.5%
DeepSeek Harness to be announced soon: DeepSeek increased its investment in Harness during the first half of the year (on June 21, Cui Tianyi posted that the Harness team was still short-staffed). CITIC Securities believes Harness's core value is solving the pain points of long-horizon tasks and connecting models with broad workplace scenarios. AI competition is shifting from models to task delivery
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Eastmoney.com (finance.eastmoney.com), source: National Business Daily · Author not disclosed · Original publication date Unknown · Site edit date 2026-09-20
Open original sourceDeepSeek V4 Flash
Download the Tabbit client to check model access