Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

DeepSeek V4 Pro · Media / benchmark · Editorial analysis

ChatBench: DeepSeek V4 Pro (high) Task Leaderboard and Spec Aggregation

ChatBench's aggregation page shows V4 Pro (high) ranks best on the coding leaderboard (Coding #22 · 76.9), mid-pack on agent tasks (Agent tasks #38 · 60.9), and clearly far back on Browser/Computer use (#57/#58) — the same model ranking very differently across task types is the key basis for model selection.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Model/version
DeepSeek-V4-Pro; source title “ChatBench: DeepSeek V4 Pro (high) Task Leaderboard and Spec Aggregation”, with no cross-version merge.
Task/harness
One-sentence takeaway ChatBench's aggregation page shows V4 Pro (high) ranks best on the coding leaderboard (Coding 22 · 76.9), mid-pack on agent tasks (Agent tasks 38 · 60.9), and clearly far back on Browser/Computer us The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.

Key data and applicable tasks

One-sentence takeaway

ChatBench's aggregation page shows V4 Pro (high) ranks best on the coding leaderboard (Coding #22 · 76.9), mid-pack on agent tasks (Agent tasks #38 · 60.9), and clearly far back on Browser/Computer use (#57/#58) — the same model ranking very differently across task types is the key basis for model selection.

Use cases

  • Tasks it suits: Quickly reviewing V4-Pro's relative position on task leaderboards such as coding, agent, browser, and instruction-following; cross-checking price, speed, and specs against the raw AA page; judging that "Pro suits coding and routine text tasks and does not suit browser/computer-automation tasks."

  • Tasks it does not suit: Treating task-leaderboard scores as absolute success rates; treating the 2026-08-12 snapshot as the latest data; treating ChatBench as an independent hands-on test (it has no own scores yet).

  • Applicable model version: DeepSeek V4 Pro (high) — the AA snapshot linked on the page is the 2026-04-24 release (not the 0813 GA; see boundaries).

  • Applicable client, agent, or API: The API service behind the AA leaderboard; the prices listed are the old prices (see below).

  • Recommended reasoning tier and parameters: This page corresponds to AA's high tier (page title "(high)"); compared with Document 06's max-tier data (Index 53), the tiers differ and the values cannot be mixed directly.

Test environment

  • Data source: The public Artificial Analysis LLM Leaderboard (stated on the page), retrieved 2026-08-12 16:00 GMT+8; ChatBench's own evals: 0; third-party scores: 12.

  • Specs: DeepSeek family; 1.6T parameters; 1M context; released 2026-04-24; open weights; tool support; no vision/audio/video listed.

  • Price and speed (snapshot values): Input $0.43/1M, output $0.87/1M, blended $0.18/1M (the old price before the 2026-08-16 adjustment); speed 68.4 tok/s; first-token latency 1.75s; total response 38.17s.

  • Leaderboard ranks and scores: Coding #22 · 76.9; Agent tasks #38 · 60.9; Blog writing #39 · 74.0; Trivia knowledge #39 · 72.8; Retrieval #38 · 69.0; Instruction following #38 · 67.6; Browser use #57 · 61.8; Computer use #58 · 58.9; Speed #29 · 49.2; Open source #6 · 68.6.

  • Pending evals (PENDING): Airwolf role/line recall, retry micro-tasks after tool errors, context-based citation extraction — ChatBench has not yet scored these items.

Input/configuration

What is disclosed: the page aggregates AA public-leaderboard data, and the EVAL EVIDENCE section notes the data carries raw-value provenance; the page tracks the AA "high" tier of a 2026-04-24 release. What is not disclosed: AA's sample sets and weights (controlled by AA and not fully public), and ChatBench has no own evals to independently reproduce the scores.

Results data

Task category / metricRank · Score
Coding#22 · 76.9
Agent tasks#38 · 60.9
Blog writing#39 · 74.0
Trivia knowledge#39 · 72.8
Retrieval#38 · 69.0
Instruction following#38 · 67.6
Browser use#57 · 61.8
Computer use#58 · 58.9
Speed#29 · 49.2
Open source#6 · 68.6
Input price$0.43 / 1M tokens
Output price$0.87 / 1M tokens
Blended price$0.18 / 1M tokens (old price)
Speed68.4 tok/s
First-token latency1.75s
Total response38.17s

Conclusion

Coding is the strongest category for V4 Pro (high) (Coding #22 · 76.9), while browser and computer use are the weakest (#57/#58). The snapshot predates the 0813 GA and the price change, so treat all numbers as directional: the leaderboards reflect an older release and the prices are the pre-adjustment ones, so re-check the live AA page and official pricing before making selection decisions.

Limitations and reproduction steps

  • Limitations: The snapshot date 2026-08-12 predates the V4-Pro-0813 GA (08-13) and the price adjustment (08-16); the leaderboards correspond to the April preview or an older snapshot, and the prices are the old ones ($0.43/$0.87). The ranks come from AA-normalized data whose samples and weights are controlled by AA and not fully public, and ChatBench has no own evals, so the scores cannot be independently reproduced. The "high" tier is only AA's naming for this configuration and its semantics may differ from the official "high".

  • Reproduction steps: Re-check the live Artificial Analysis page and the official DeepSeek pricing page for current ranks and prices; treat the values on this page as an old snapshot and do not use them as the basis for absolute success rates or cross-model conclusions.

Original evidence and data

  • The page states "DeepSeek V4 Pro (high) is tracked from the live Artificial Analysis public leaderboard with ChatBench-normalized intelligence, coding, agentic, price, latency, speed, and context metadata".

  • The "EVAL EVIDENCE" section explicitly marks the only PASS item as "Artificial Analysis public leaderboard ingestion" and notes the data carries raw-value provenance; the other 3 items are PENDING and were not counted in the conclusions.

  • The leaderboard values differ in basis from the AA page in Document 06 (that page is max effort, 0813; this page is high, an old snapshot); the difference comes from version and tier, not from the same data transcribed in two places.

Applicability boundaries

  • The snapshot date 2026-08-12 predates the V4-Pro-0813 GA (08-13) and the price adjustment (08-16): the leaderboards correspond to the April preview or an older snapshot, and the prices are the old ones ($0.43/$0.87); for selection, defer to the live AA page and the official pricing page.

  • The leaderboard ranks come from AA-normalized data; the samples and weights are controlled by AA and not fully public; ChatBench itself has no own evaluations, so these scores cannot be independently reproduced.

  • The "high" tier only reflects AA's naming for this configuration; its semantics may differ from the official "high", so follow the official API parameters when integrating.

  • The low Browser/Computer use ranks are directional evidence (a common pattern for text-only models on such tasks), but the exact values are influenced by the AA harness.

Source excerpt or observation (compliance short quote only)

The page's original text (EVAL EVIDENCE): the PASS item is "Artificial Analysis public leaderboard ingestion — Imports public intelligence, price, speed, latency, and response-time metrics with raw-value provenance."

What this supports

  • Supports the source-specific observation in “ChatBench: DeepSeek V4 Pro (high) Task Leaderboard and Spec Aggregation”: One-sentence takeaway ChatBench's aggregation page shows V4 Pro (high) ranks best on the coding leaderboard (Coding 22 · 76.9), mid-pack on agent tasks (Agent tasks 38 · 60.9), and clearly f

What this does not support

  • Does not support a general capability or production-rate claim from “ChatBench: DeepSeek V4 Pro (high) Task Leaderboard and Spec Aggregation”; the source lacks a controlled task set, provider snapshot, and repeated independent retest.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

ChatBench (benchmarks.chatbench.org) · ChatBench project (independent third-party aggregation platform) · Original publication date Unknown · Site edit date 2026-09-20

Open original source

DeepSeek V4 Pro

Compare DeepSeek V4 Pro in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

DeepSeek V4 Pro: What Changed, What It Costs, and Who Should Use It

A sourced guide to DeepSeek V4 Pro 0813: the agent upgrades, live API limits, price boundary, independent evidence and a safer pilot plan.

Related reviews

DeepSeek-V4-Pro-0813: MindStudio's Eight-Task Coding and Agent Hands-on ComparisonV4 Pro 0813, MindStudio eight-task test on 2026-08-13; 61/80 (76.25%), frontend/planning/math/long-horizon; full prompts, repeats and blind review undisclosed.DeepSeek-V4-Pro XSCT Bench Two-Case Comparison: Strong Planning, Weak ClarificationV4 Pro, XSCT Bench two cases collected 2026-08-21; autonomous planning 98.0/92.6 versus ambiguous clarification 68.5; prompts, repeats, and harness undisclosed.Artificial Analysis: DeepSeek V4 Pro 0813 (Max Effort) Intelligence Index, Cost, and PositioningThe 2026-08-21 Artificial Analysis snapshot recorded V4 Pro 0813 max effort at index 53, 80.3 tok/s, $3.96/1M output, 1M context, and 1.6T/49B; the page reopened on 2026-09-20 shows index 36, so the snapshots must not be mixed.DeepSeek-V4-Pro Official Release: Reasoning and Agent UpgradesV4 Pro 0813 GA announcement dated 2026-08-13; effort, Responses API and Codex positioning; pricing effective 2026-08-16; no unified benchmark or sample.DeepSeek-V4-Pro Thinking Levels and Tool-Calling WorkflowV4-Pro enables thinking by default and uses high as the default effort level; use low for simple tasks, high for day-to-day Agents, and max for complex tasks, and pass the complete `reasoning_content` back on every round of a tool call.DeepSeek-V4-Pro Responses Configuration Workflow in CodexDeepSeek-V4-Pro can be connected to the Codex CLI, the ChatGPT desktop app, and the VS Code extension through the native Responses API; a single configuration is shared across them, but you should back up and validate `config.toml`/`models.json` first.XSCT Bench “Autonomous Planning and Execution” Case: Agent Tool-Calling Prompt and Generated Result for deepseek-v4-proThe platform publishes the complete system prompt, user prompt, the model's actual generated output, and scores at two difficulty levels (Basic 98.0 / Advanced 92.6): a directly reusable Agent execution prompt that says “plan with `<plan>` first, call tools via JSON, review with `<observation>`, and wrap up with `<summary>`.”.DeepSeek-V4-Pro 1M Context Environment Variable Configuration Workflow in Claude CodeWith 8 environment variables, you can point Claude Code (and Claude Desktop Developer Mode) to DeepSeek, unlock a 1M context window with `deepseek-v4-pro[1m]`, use `deepseek-v4-flash` for subagents, set the main model's effort to `max`, and set the automatic compaction window to 786432.