ChatBench's aggregation page shows V4 Pro (high) ranks best on the coding leaderboard (Coding #22 · 76.9), mid-pack on agent tasks (Agent tasks #38 · 60.9), and clearly far back on Browser/Computer use (#57/#58) — the same model ranking very differently across task types is the key basis for model selection.
Tasks it suits: Quickly reviewing V4-Pro's relative position on task leaderboards such as coding, agent, browser, and instruction-following; cross-checking price, speed, and specs against the raw AA page; judging that "Pro suits coding and routine text tasks and does not suit browser/computer-automation tasks."
Tasks it does not suit: Treating task-leaderboard scores as absolute success rates; treating the 2026-08-12 snapshot as the latest data; treating ChatBench as an independent hands-on test (it has no own scores yet).
Applicable model version: DeepSeek V4 Pro (high) — the AA snapshot linked on the page is the 2026-04-24 release (not the 0813 GA; see boundaries).
Applicable client, agent, or API: The API service behind the AA leaderboard; the prices listed are the old prices (see below).
Recommended reasoning tier and parameters: This page corresponds to AA's high tier (page title "(high)"); compared with Document 06's max-tier data (Index 53), the tiers differ and the values cannot be mixed directly.
Data source: The public Artificial Analysis LLM Leaderboard (stated on the page), retrieved 2026-08-12 16:00 GMT+8; ChatBench's own evals: 0; third-party scores: 12.
Specs: DeepSeek family; 1.6T parameters; 1M context; released 2026-04-24; open weights; tool support; no vision/audio/video listed.
Price and speed (snapshot values): Input $0.43/1M, output $0.87/1M, blended $0.18/1M (the old price before the 2026-08-16 adjustment); speed 68.4 tok/s; first-token latency 1.75s; total response 38.17s.
Leaderboard ranks and scores: Coding #22 · 76.9; Agent tasks #38 · 60.9; Blog writing #39 · 74.0; Trivia knowledge #39 · 72.8; Retrieval #38 · 69.0; Instruction following #38 · 67.6; Browser use #57 · 61.8; Computer use #58 · 58.9; Speed #29 · 49.2; Open source #6 · 68.6.
Pending evals (PENDING): Airwolf role/line recall, retry micro-tasks after tool errors, context-based citation extraction — ChatBench has not yet scored these items.
What is disclosed: the page aggregates AA public-leaderboard data, and the EVAL EVIDENCE section notes the data carries raw-value provenance; the page tracks the AA "high" tier of a 2026-04-24 release. What is not disclosed: AA's sample sets and weights (controlled by AA and not fully public), and ChatBench has no own evals to independently reproduce the scores.
| Task category / metric | Rank · Score |
|---|---|
| Coding | #22 · 76.9 |
| Agent tasks | #38 · 60.9 |
| Blog writing | #39 · 74.0 |
| Trivia knowledge | #39 · 72.8 |
| Retrieval | #38 · 69.0 |
| Instruction following | #38 · 67.6 |
| Browser use | #57 · 61.8 |
| Computer use | #58 · 58.9 |
| Speed | #29 · 49.2 |
| Open source | #6 · 68.6 |
| Input price | $0.43 / 1M tokens |
| Output price | $0.87 / 1M tokens |
| Blended price | $0.18 / 1M tokens (old price) |
| Speed | 68.4 tok/s |
| First-token latency | 1.75s |
| Total response | 38.17s |
Coding is the strongest category for V4 Pro (high) (Coding #22 · 76.9), while browser and computer use are the weakest (#57/#58). The snapshot predates the 0813 GA and the price change, so treat all numbers as directional: the leaderboards reflect an older release and the prices are the pre-adjustment ones, so re-check the live AA page and official pricing before making selection decisions.
Limitations: The snapshot date 2026-08-12 predates the V4-Pro-0813 GA (08-13) and the price adjustment (08-16); the leaderboards correspond to the April preview or an older snapshot, and the prices are the old ones ($0.43/$0.87). The ranks come from AA-normalized data whose samples and weights are controlled by AA and not fully public, and ChatBench has no own evals, so the scores cannot be independently reproduced. The "high" tier is only AA's naming for this configuration and its semantics may differ from the official "high".
Reproduction steps: Re-check the live Artificial Analysis page and the official DeepSeek pricing page for current ranks and prices; treat the values on this page as an old snapshot and do not use them as the basis for absolute success rates or cross-model conclusions.
The page states "DeepSeek V4 Pro (high) is tracked from the live Artificial Analysis public leaderboard with ChatBench-normalized intelligence, coding, agentic, price, latency, speed, and context metadata".
The "EVAL EVIDENCE" section explicitly marks the only PASS item as "Artificial Analysis public leaderboard ingestion" and notes the data carries raw-value provenance; the other 3 items are PENDING and were not counted in the conclusions.
The leaderboard values differ in basis from the AA page in Document 06 (that page is max effort, 0813; this page is high, an old snapshot); the difference comes from version and tier, not from the same data transcribed in two places.
The snapshot date 2026-08-12 predates the V4-Pro-0813 GA (08-13) and the price adjustment (08-16): the leaderboards correspond to the April preview or an older snapshot, and the prices are the old ones ($0.43/$0.87); for selection, defer to the live AA page and the official pricing page.
The leaderboard ranks come from AA-normalized data; the samples and weights are controlled by AA and not fully public; ChatBench itself has no own evaluations, so these scores cannot be independently reproduced.
The "high" tier only reflects AA's naming for this configuration; its semantics may differ from the official "high", so follow the official API parameters when integrating.
The low Browser/Computer use ranks are directional evidence (a common pattern for text-only models on such tasks), but the exact values are influenced by the AA harness.
The page's original text (EVAL EVIDENCE): the PASS item is "Artificial Analysis public leaderboard ingestion — Imports public intelligence, price, speed, latency, and response-time metrics with raw-value provenance."
DeepSeek V4 Pro