Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaDeepSeek V4 Flash

The Official DeepSeek V4-Flash Is Here! AI Developers Test It: "Fantastic" Pricing, Agent Capabilities Close to Top-Tier Models

Original source

Eastmoney.com (finance.eastmoney.com), source: National Business Daily

Tabbit curation2026-08-19

Read original

Core facts

On July 31, DeepSeek announced that the official DeepSeek-V4-Flash API had entered public beta, with a major upgrade to its agent capabilities. The official version's benchmark scores were far higher than those of the V4-Pro preview. In the ultimate agent test, the official V4-Flash scored 25.2 points, close to Opus-4.8's 25.7 and well above the V4-Pro preview's 15.8.

Developer field-test feedback (interviewees using pseudonyms)

Zhang Ze (an AI developer who uses codex and hermes extensively):

  • Integrated V4-Flash with Hermes: on the same task with the same prompt, the official V4-Flash + Hermes finished in about 40 seconds, while GPT-5.6 Sol + Codex took 1 minute 47 seconds

  • Acknowledged that Codex's output quality is indeed higher, but said the official V4-Flash delivers work at a usable level, "above the passing line"

  • Its accuracy in selecting and calling skills is generally good. It can adjust subsequent steps based on tool-returned results, and its combinations can already smoothly meet his personal needs

Sun Tao (feedback from August 1):

  • "Among Chinese models, it is somewhat better than GLM 5.2, but it is still somewhat behind GPT 5.6"

  • His primary tools are Codex and Pi Agent; integrating V4-Flash with Pi Agent provided a good experience

  • Because it is not multimodal, its main uses are reasoning-heavy work, code, and long documents

Cost comparison

  • GPT-5.6 Sol: $5 per million input tokens, $0.5 for cached input, and $30 for output

  • Official V4-Flash: 1 yuan for input, 0.02 yuan for cached input, and 2 yuan for output (RMB)

  • The cost gap ranges from dozens to hundreds of times

  • Zhang Ze's field test: V4-Flash connected to Hermes performed a task involving local file reading and aggregation of information retrieved from across the web, using about 510,000 tokens and consuming 0.53 yuan

Peak/off-peak pricing mechanism

In an email sent at the end of June, DeepSeek previewed the API's first introduction of a peak/off-peak pricing mechanism:

  • V4-Pro: for one million input tokens (cache hit), the price drops from 1 yuan in the April preview to 0.025 yuan during regular periods and 0.05 yuan during peak periods in the official version

  • Official V4-Flash: during peak periods, one million cache-miss input tokens and output tokens are actually twice as expensive as in the April preview, at 2 yuan and 4 yuan, respectively

  • The specific implementation time for the peak/off-peak pricing mechanism has not yet been formally announced

Market impact

  • OpenRouter data from July 28 showed Chinese AI models leading the world in weekly call volume: Xiaomi MiMo-V2.5, DeepSeek V4-Flash, Tencent HY3, Zhipu GLM-5.2, and DeepSeek V4-Pro (two DeepSeek models made the top five)

  • In the 28 days through July 26: Chinese models held about 63.5% of the market, while US models held 35.5%

  • DeepSeek Harness to be announced soon: DeepSeek increased its investment in Harness during the first half of the year (on June 21, Cui Tianyi posted that the Harness team was still short-staffed). CITIC Securities believes Harness's core value is solving the pain points of long-horizon tasks and connecting models with broad workplace scenarios. AI competition is shifting from models to task delivery

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

DeepSeek V4 Flash

Use and compare models in Tabbit

DeepSeek V4 Flash

Related reviews

OfficialDeepSeek API Docs2026-07-31

DeepSeek-V4-Flash: 0731 Benchmark Update and Harness Conditions

MediaMindStudio2026-08-01

DeepSeek-V4-Flash: Local Deployment, Quantization, and Agent Testing

MediaLightning AI blog2026-04-27

DeepSeek V4 Alters Everything We Knew About Price-Performance Math (Lightning AI)

MediaBenchLM.ai (model benchmarking and pricing tracking site)2026-07-31

DeepSeek V4 Flash 0731 Benchmarks, Pricing & Speed (BenchLM)

DeepSeek V4 Flash

Related prompts

OfficialDeepSeek API Docs2026-07-31

DeepSeek-V4-Flash: 0731 Thinking Levels and Tool-Calling Configuration

OfficialDeepSeek API Docs2026-07-31

DeepSeek-V4-Flash: Codex Responses API integration workflow

CommunityGitHub repository victorchen96/deepseekv4rolepalyinstruct

A Guide to Special Control Instructions for DeepSeek-V4 Role-Playing (Thinking-Mode Switching Guide)

CommunityReddit r/SillyTavernAI

DeepSeek V4 RP Guide — How to Switch Between Character Immersion & Pure Analysis Thinking Modes (Reddit r/SillyTavernAI)