On July 31, DeepSeek announced that the official DeepSeek-V4-Flash API had entered public beta, with a major upgrade to its agent capabilities. The official version's benchmark scores were far higher than those of the V4-Pro preview. In the ultimate agent test, the official V4-Flash scored 25.2 points, close to Opus-4.8's 25.7 and well above the V4-Pro preview's 15.8.
Zhang Ze (an AI developer who uses codex and hermes extensively):
Integrated V4-Flash with Hermes: on the same task with the same prompt, the official V4-Flash + Hermes finished in about 40 seconds, while GPT-5.6 Sol + Codex took 1 minute 47 seconds
Acknowledged that Codex's output quality is indeed higher, but said the official V4-Flash delivers work at a usable level, "above the passing line"
Its accuracy in selecting and calling skills is generally good. It can adjust subsequent steps based on tool-returned results, and its combinations can already smoothly meet his personal needs
Sun Tao (feedback from August 1):
"Among Chinese models, it is somewhat better than GLM 5.2, but it is still somewhat behind GPT 5.6"
His primary tools are Codex and Pi Agent; integrating V4-Flash with Pi Agent provided a good experience
Because it is not multimodal, its main uses are reasoning-heavy work, code, and long documents
GPT-5.6 Sol: $5 per million input tokens, $0.5 for cached input, and $30 for output
Official V4-Flash: 1 yuan for input, 0.02 yuan for cached input, and 2 yuan for output (RMB)
The cost gap ranges from dozens to hundreds of times
Zhang Ze's field test: V4-Flash connected to Hermes performed a task involving local file reading and aggregation of information retrieved from across the web, using about 510,000 tokens and consuming 0.53 yuan
In an email sent at the end of June, DeepSeek previewed the API's first introduction of a peak/off-peak pricing mechanism:
V4-Pro: for one million input tokens (cache hit), the price drops from 1 yuan in the April preview to 0.025 yuan during regular periods and 0.05 yuan during peak periods in the official version
Official V4-Flash: during peak periods, one million cache-miss input tokens and output tokens are actually twice as expensive as in the April preview, at 2 yuan and 4 yuan, respectively
The specific implementation time for the peak/off-peak pricing mechanism has not yet been formally announced
OpenRouter data from July 28 showed Chinese AI models leading the world in weekly call volume: Xiaomi MiMo-V2.5, DeepSeek V4-Flash, Tencent HY3, Zhipu GLM-5.2, and DeepSeek V4-Pro (two DeepSeek models made the top five)
In the 28 days through July 26: Chinese models held about 63.5% of the market, while US models held 35.5%
DeepSeek Harness to be announced soon: DeepSeek increased its investment in Harness during the first half of the year (on June 21, Cui Tianyi posted that the Harness team was still short-staffed). CITIC Securities believes Harness's core value is solving the pain points of long-horizon tasks and connecting models with broad workplace scenarios. AI competition is shifting from models to task delivery
DeepSeek V4 Flash