V4 Pro 0813 GA announcement dated 2026-08-13; effort, Responses API and Codex positioning; pricing effective 2026-08-16; no unified benchmark or sample.
DeepSeek API Docs · Read evidenceDeepSeek V4 Pro · Reviews and evidence
Which DeepSeek V4 Pro conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
V4 Pro 0813, MindStudio eight-task test on 2026-08-13; 61/80 (76.25%), frontend/planning/math/long-horizon; full prompts, repeats and blind review undisclosed.
MindStudio · Read evidenceV4 Pro, XSCT Bench two cases collected 2026-08-21; autonomous planning 98.0/92.6 versus ambiguous clarification 68.5; prompts, repeats, and harness undisclosed.
XSCT Bench (xsctbench.com, scenario-based model-selection evaluation) · Read evidenceFull reviews and related reading
Selected evidence
DeepSeek-V4-Pro Official Release: Reasoning and Agent Upgrades
V4 Pro 0813 GA announcement dated 2026-08-13; effort, Responses API and Codex positioning; pricing effective 2026-08-16; no unified benchmark or sample.
Unverified: the original source could not be rechecked.
- Conditions
- V4 Pro 0813 GA announcement dated 2026-08-13; effort, Responses API and Codex positioning; pricing effective 2026-08-16; no unified benchmark or sample
DeepSeek-V4-Pro-0813: MindStudio's Eight-Task Coding and Agent Hands-on Comparison
V4 Pro 0813, MindStudio eight-task test on 2026-08-13; 61/80 (76.25%), frontend/planning/math/long-horizon; full prompts, repeats and blind review undisclosed.
Unverified: the original source could not be rechecked.
- Conditions
- V4 Pro 0813, MindStudio eight-task test on 2026-08-13; 61/80 (76.25%), frontend/planning/math/long-horizon; full prompts, repeats and blind review undisclosed
DeepSeek-V4-Pro XSCT Bench Two-Case Comparison: Strong Planning, Weak Clarification
V4 Pro, XSCT Bench two cases collected 2026-08-21; autonomous planning 98.0/92.6 versus ambiguous clarification 68.5; prompts, repeats, and harness undisclosed.
Unverified: the original source could not be rechecked.
- Conditions
- V4 Pro, XSCT Bench two cases collected 2026-08-21; autonomous planning 98.0/92.6 versus ambiguous clarification 68.5; prompts, repeats, and harness undisclosed
Artificial Analysis: DeepSeek V4 Pro 0813 (Max Effort) Intelligence Index, Cost, and Positioning
The 2026-08-21 Artificial Analysis snapshot recorded V4 Pro 0813 max effort at index 53, 80.3 tok/s, $3.96/1M output, 1M context, and 1.6T/49B; the page reopened on 2026-09-20 shows index 36, so the snapshots must not be mixed.
Unverified: the original source could not be rechecked.
- Conditions
- Version/tier V4 Pro 0813, Reasoning Max Effort; the 2026-08-21 snapshot showed index 53 and the page reopened on 2026-09-20 shows 36; speed, price, and specs belong to their respective snapshots, with items, repeats, and hardware incomplete.
All sources
All sources
DeepSeek-V4-Pro Official Release: Reasoning and Agent Upgrades
V4 Pro 0813 GA announcement dated 2026-08-13; effort, Responses API and Codex positioning; pricing effective 2026-08-16; no unified benchmark or sample.
Unverified: the original source could not be rechecked.
- Conditions
- V4 Pro 0813 GA announcement dated 2026-08-13; effort, Responses API and Codex positioning; pricing effective 2026-08-16; no unified benchmark or sample
DeepSeek-V4-Pro-0813: MindStudio's Eight-Task Coding and Agent Hands-on Comparison
V4 Pro 0813, MindStudio eight-task test on 2026-08-13; 61/80 (76.25%), frontend/planning/math/long-horizon; full prompts, repeats and blind review undisclosed.
Unverified: the original source could not be rechecked.
- Conditions
- V4 Pro 0813, MindStudio eight-task test on 2026-08-13; 61/80 (76.25%), frontend/planning/math/long-horizon; full prompts, repeats and blind review undisclosed
DeepSeek-V4-Pro XSCT Bench Two-Case Comparison: Strong Planning, Weak Clarification
V4 Pro, XSCT Bench two cases collected 2026-08-21; autonomous planning 98.0/92.6 versus ambiguous clarification 68.5; prompts, repeats, and harness undisclosed.
Unverified: the original source could not be rechecked.
- Conditions
- V4 Pro, XSCT Bench two cases collected 2026-08-21; autonomous planning 98.0/92.6 versus ambiguous clarification 68.5; prompts, repeats, and harness undisclosed
Artificial Analysis: DeepSeek V4 Pro 0813 (Max Effort) Intelligence Index, Cost, and Positioning
The 2026-08-21 Artificial Analysis snapshot recorded V4 Pro 0813 max effort at index 53, 80.3 tok/s, $3.96/1M output, 1M context, and 1.6T/49B; the page reopened on 2026-09-20 shows index 36, so the snapshots must not be mixed.
Unverified: the original source could not be rechecked.
- Conditions
- Version/tier V4 Pro 0813, Reasoning Max Effort; the 2026-08-21 snapshot showed index 53 and the page reopened on 2026-09-20 shows 36; speed, price, and specs belong to their respective snapshots, with items, repeats, and hardware incomplete.
Reuters DeepSeek-V4-Pro-0813: Official Pricing vs. Independent Index
Reuters cites Artificial Analysis's independent pricing and index data: V4-Pro-0813 scores 53 on the reasoning Intelligence Index, versus 40 for V4 Flash, but Pro's input and output prices are roughly 9 and 14 times those of Flash, respectively. Model selection must account for both quality and cost.
Unverified: the original source could not be rechecked.
- Model/version
- DeepSeek-V4-Pro; source title “Reuters DeepSeek-V4-Pro-0813: Official Pricing vs. Independent Index”, with no cross-version merge.
- Task/harness
- One-sentence takeaway Reuters cites Artificial Analysis's independent pricing and index data: V4-Pro-0813 scores 53 on the reasoning Intelligence Index, versus 40 for V4 Flash, but Pro's input and output prices are rough The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
DeepSeek-V4-Pro Reddit Field Report: Long Context and Prompt Precision
The community's on-the-ground view is that V4-Pro is useful for large amounts of context, messy coding prompts, and low-cost personal projects. Larger architecture tasks, however, depend more on precise specifications, documentation, testing, and safeguards against irreversible changes; a single prompt is not enough.
Unverified: the original source could not be rechecked.
- Model/version
- DeepSeek-V4-Pro; source title “DeepSeek-V4-Pro Reddit Field Report: Long Context and Prompt Precision”, with no cross-version merge.
- Task/harness
- One-sentence takeaway The community's on-the-ground view is that V4-Pro is useful for large amounts of context, messy coding prompts, and low-cost personal projects. Larger architecture tasks, however, depend more on pre The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
ChatBench: DeepSeek V4 Pro (high) Task Leaderboard and Spec Aggregation
ChatBench's aggregation page shows V4 Pro (high) ranks best on the coding leaderboard (Coding #22 · 76.9), mid-pack on agent tasks (Agent tasks #38 · 60.9), and clearly far back on Browser/Computer use (#57/#58) — the same model ranking very differently across task types is the key basis for model selection.
Unverified: the original source could not be rechecked.
- Model/version
- DeepSeek-V4-Pro; source title “ChatBench: DeepSeek V4 Pro (high) Task Leaderboard and Spec Aggregation”, with no cross-version merge.
- Task/harness
- One-sentence takeaway ChatBench's aggregation page shows V4 Pro (high) ranks best on the coding leaderboard (Coding 22 · 76.9), mid-pack on agent tasks (Agent tasks 38 · 60.9), and clearly far back on Browser/Computer us The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
DeepSeek-V4-Pro Reddit Field Report: Pricing Hike and Model Migration
The community's general reaction to the new pricing effective 2026-08-16 (which introduces peak/off-peak billing, with cache hit up as much as 1,114%) is "the model is good but no longer cheap." Some users have already migrated their main workloads to Luna/Claude/Codex and others, and use V4-Pro only for high-value planning/review tasks.
Unverified: the original source could not be rechecked.
- Model/version
- DeepSeek-V4-Pro; source title “DeepSeek-V4-Pro Reddit Field Report: Pricing Hike and Model Migration”, with no cross-version merge.
- Task/harness
- One-sentence takeaway The community's general reaction to the new pricing effective 2026-08-16 (which introduces peak/off-peak billing, with cache hit up as much as 1,114%) is "the model is good but no longer cheap." Som The complete task set, runtime parameters, and review procedure are not fully public.
- Sample/date
- Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.
DeepSeek V4 Pro
Compare DeepSeek V4 Pro in Tabbit
Model access, features, and permissions depend on your current client account.