The official release of DeepSeek-V4-Flash is live! Post-training has sent its Agent and coding capabilities soaring, with overall capabilities surpassing V4 Pro. A “floor price” of just RMB 1 for one million input tokens completely resets the barrier to entry for Agent tasks.
A major boost in Agent capabilities: Extensive post-training optimization for Agent scenarios has significantly improved complex task decomposition, tool use, and code execution. The official scorecard comprehensively outperforms the V4-Pro-Preview released in April; on Agent Last Exam, the scores are 25.2 versus 25.7, approaching Opus 4.8-level performance
Native Responses API support, compatible with Codex workflows: You can connect to the Codex CLI, VS Code extensions, or the ChatGPT desktop app without changing the base_url. Official one-click setup scripts are available for macOS/Linux and Windows
The price slasher: One million input tokens (cache miss) cost RMB 1, output costs RMB 2, and cache-hit input costs RMB 0.02. Every item is cheaper than GPT-5.6 Luna
The Pro version is on the way: This upgrade only targets the Flash API; the app and web versions remain unchanged for now
Ranked 21st on the Artificial Analysis leaderboard
Independent testing with the question bank collected by 302.AI: logic and mathematics (10 questions), human intuition (7 questions), and programming simulation (12 questions)
All models were tested in the 302.AI Studio client with the same prompts, using the first generated result
Scoring: Scores are averaged after applying the corresponding deduction criteria, with 10 points as the maximum; grades are S/A/B/C
Case 1: Complex logical reasoning (find a 10-digit number that satisfies the self-descriptive conditions)
DeepSeek-V4-Flash reasoned correctly, producing the unique solution 6210001000
Claude Opus 4.8 was broadly logically consistent, but its reasoning result did not match the problem’s requirements
Case 2: Programmatic SVG graphic generation (a kangaroo jumping through a desert and an animated F1 race car)
V4-Flash was slightly better than Opus 4.8 in the complexity of its graphic composition, but the quality of its dynamic implementation was relatively rudimentary
Benchmark scores do not equal real-world experience; Agent tasks involve planning, execution, error correction, and other stages
This evaluation focuses on logic, mathematics, programming, multimodality, human intuition, and related tests. It is not an authoritative test of specialized cutting-edge fields; its purpose is to observe the model’s evolutionary trend and provide a reference for model selection
DeepSeek V4 Flash