Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

DeepSeek V4 Pro · Media / benchmark · Editorial analysis

Artificial Analysis: DeepSeek V4 Pro 0813 (Max Effort) Intelligence Index, Cost, and Positioning

The 2026-08-21 Artificial Analysis snapshot recorded V4 Pro 0813 max effort at index 53, 80.3 tok/s, $3.96/1M output, 1M context, and 1.6T/49B; the page reopened on 2026-09-20 shows index 36, so the snapshots must not be mixed.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Conditions
Version/tier V4 Pro 0813, Reasoning Max Effort; the 2026-08-21 snapshot showed index 53 and the page reopened on 2026-09-20 shows 36; speed, price, and specs belong to their respective snapshots, with items, repeats, and hardware incomplete.

Key data and applicable tasks

One-sentence takeaway

AA's measurement of DeepSeek V4 Pro 0813 (Reasoning, Max Effort): Intelligence Index 53, ranked 3rd in its pool of 107 models; speed 80.3 tok/s; output $3.96/1M (peak); 1M context; 1600B total parameters with 49B active — top-tier intelligence, but on the expensive and verbose side (a single index evaluation generated 130M tokens).

Use cases

  • Suitable tasks: Comparing V4-Pro with other frontier models on the same index scale (intelligence/speed/cost/verbosity); providing third-party evidence for “what Pro is good at” (high intelligence + long context + text-only).

  • Unsuitable tasks: Treating the Intelligence Index as the success rate for a specific codebase; mixing data from the 2026-04-24 older version with the new 0813 version; using peak prices to replace all cost calculations (off-peak and cache discounts differ substantially).

  • Applicable model versions: DeepSeek V4 Pro 0813 (the page measures it under a max effort reasoning configuration; the official non-reasoning variant also exists but is not tested on this page).

  • Applicable clients, Agents, or APIs: AA's unified API Provider Benchmark harness; the model outputs text, and the official API supports the Responses/Anthropic formats.

  • Recommended reasoning levels and parameters: This page is the max effort tier; AA notes that lower tiers (e.g., high/low) produce different intelligence, token, and cost figures, so each tier must be queried separately.

Test environment

  • Metrics: Artificial Analysis Intelligence Index (composite intelligence), cost per task, output speed, and Verbosity (total tokens generated during the index evaluation).

  • Measurement results (page text): Intelligence Index 53, ranked 3/107 (peer median 27); Speed 80.3 tok/s, ranked 22/107 (median 67); Cost $1.32/1M input, $3.96/1M output (peak pricing, 97% cache discount), $0.25 per Intelligence Index task, ranked 24/107 (input median $0.30, output median $1.20); Verbosity 130M tokens (median 100M, on the verbose side); total cost of the full index evaluation $604.51.

  • Technical specs: 1M context (roughly 1,500 A4 pages); 1600B total parameters, 49B active; text input and output; MIT license; weights on Hugging Face.

  • Comparison baseline: A pool of 107 peer models (open weights, Large >150B, or reasoning class); the “Highlights” comparison chart lists Claude Opus 5 (max), Claude Fable 5, GPT-5.6 Sol (max), Grok 4.6 (high), Kimi K3 (max), GLM-5.3 (max), Muse Spark 1.2 (xhigh), and others.

Input/configuration

The page's measurements were run through AA's unified API Provider Benchmark harness with a fixed max effort configuration. What is disclosed: the Intelligence Index methodology, per-task cost, output speed, and verbosity figures. What is not disclosed: the full sample sets, weighting, and harness details, which are controlled by AA and not fully public (the Reuters report makes the same point).

Results data

MetricValue
Intelligence Index53, ranked 3/107 (peer median 27)
Speed80.3 tok/s, ranked 22/107 (median 67)
Cost$1.32/1M in / $3.96/1M out (peak pricing, 97% cache discount); $0.25 per Intelligence Index task; ranked 24/107 (input median $0.30, output median $1.20)
Verbosity130M tokens per index evaluation (median 100M)
Total evaluation cost$604.51
Context1M (roughly 1,500 A4 pages)
Parameters1600B total / 49B active
ModalityText in / text out
LicenseMIT (weights on Hugging Face)

Conclusion

Under AA's measurement at max effort, DeepSeek V4 Pro 0813 delivers top-tier intelligence (Index 53, 3rd of 107) but is expensive and verbose: peak prices of $1.32/$3.96 per 1M tokens, $0.25 per task, and 130M tokens generated per evaluation are all above the peer medians. The index score of 53 is consistent with the Reuters 2026-08-13 report, so the AA data cross-checks. Note the version distinction: the 0813 snapshot on this page is not the same as the older 2026-04-24 preview snapshot (ChatBench AA Index 43.75), so the two must not be mixed when comparing.

Limitations and reproduction steps

  • Limitation: AA is a third-party index; the samples, weights, and harness are controlled by AA and not fully disclosed (the Reuters report makes the same point). The index is a composite score, not a task success rate.

  • Limitation: “max effort” is the page's fixed configuration; the max tier's token consumption and cost are significantly higher than the official default high tier. Production model selection should re-check data at the actual tier in use.

  • Limitation: Verbosity of 130M tokens shows that max-tier output is wordy; long conversations and batch tasks must factor token costs into the budget, using low/high tier comparisons when needed.

  • Limitation: Page prices are point-in-time data as collected; DeepSeek moved to peak/off-peak billing effective 2026-08-16, so later price changes require re-verification.

  • Reproduction steps: Re-check the live Artificial Analysis page for DeepSeek V4 Pro per effort tier (max effort as shown here, plus high/low for comparison), and cross-check current prices against the official pricing page for the effective peak/off-peak rates.

Original evidence and data

  • Page positioning statement (verbatim): “DeepSeek V4 Pro 0813 (Reasoning, Max Effort) is amongst the leading models in intelligence, but somewhat expensive when comparing to other open weight models of similar size. It's also faster than average, however somewhat verbose.”

  • The same index score of 53 matches the Reuters 2026-08-13 report (see review 03), showing the AA data can be cross-checked; the ChatBench 2026-08-12 snapshot shows the older version (Released 2026-04-24, AA Index 43.75), which differs from this page's 0813 — mind the version distinction.

  • Cost basis: page prices are the peak tier ($1.32/$3.96), consistent with the peak prices on the official pricing page effective 2026-08-16; off-peak is half ($0.66/$1.98), and cache-hit input is $0.022/$0.044.

Applicability boundaries

  • AA is a third-party index: the samples, weights, and harness are controlled by AA and not fully disclosed (the Reuters report makes the same point); the index is a composite score, not a task success rate.

  • “max effort” is the page's fixed configuration: the max tier's token consumption and cost are significantly higher than the official default high tier; production model selection should re-check data at the actual tier in use.

  • Verbosity of 130M tokens shows that max-tier output is wordy: long conversations and batch tasks must factor token costs into the budget, using low/high tier comparisons when needed.

  • Page prices are point-in-time data as collected: DeepSeek implemented peak/off-peak billing effective 2026-08-16, and later price adjustments require re-verification.

Source excerpt or observation (compliance short quote only)

Page text: “In total, it cost $604.51 to evaluate DeepSeek V4 Pro 0813 (Reasoning, Max Effort) on the Intelligence Index.”

What this supports

  • supports comparing index, speed, price, context and verbosity within AA method

What this does not support

  • does not make aggregate values fixed for your provider, effort, or task set

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis (independent model intelligence platform) · The Artificial Analysis team (independent third party) · Original publication date Unknown · Site edit date 2026-09-20

Open original source

DeepSeek V4 Pro

Compare DeepSeek V4 Pro in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

DeepSeek V4 Pro: What Changed, What It Costs, and Who Should Use It

A sourced guide to DeepSeek V4 Pro 0813: the agent upgrades, live API limits, price boundary, independent evidence and a safer pilot plan.

Related reviews

DeepSeek-V4-Pro-0813: MindStudio's Eight-Task Coding and Agent Hands-on ComparisonV4 Pro 0813, MindStudio eight-task test on 2026-08-13; 61/80 (76.25%), frontend/planning/math/long-horizon; full prompts, repeats and blind review undisclosed.DeepSeek-V4-Pro XSCT Bench Two-Case Comparison: Strong Planning, Weak ClarificationV4 Pro, XSCT Bench two cases collected 2026-08-21; autonomous planning 98.0/92.6 versus ambiguous clarification 68.5; prompts, repeats, and harness undisclosed.DeepSeek-V4-Pro Official Release: Reasoning and Agent UpgradesV4 Pro 0813 GA announcement dated 2026-08-13; effort, Responses API and Codex positioning; pricing effective 2026-08-16; no unified benchmark or sample.Reuters DeepSeek-V4-Pro-0813: Official Pricing vs. Independent IndexReuters cites Artificial Analysis's independent pricing and index data: V4-Pro-0813 scores 53 on the reasoning Intelligence Index, versus 40 for V4 Flash, but Pro's input and output prices are roughly 9 and 14 times those of Flash, respectively. Model selection must account for both quality and cost.DeepSeek-V4-Pro Thinking Levels and Tool-Calling WorkflowV4-Pro enables thinking by default and uses high as the default effort level; use low for simple tasks, high for day-to-day Agents, and max for complex tasks, and pass the complete `reasoning_content` back on every round of a tool call.DeepSeek-V4-Pro Responses Configuration Workflow in CodexDeepSeek-V4-Pro can be connected to the Codex CLI, the ChatGPT desktop app, and the VS Code extension through the native Responses API; a single configuration is shared across them, but you should back up and validate `config.toml`/`models.json` first.XSCT Bench “Autonomous Planning and Execution” Case: Agent Tool-Calling Prompt and Generated Result for deepseek-v4-proThe platform publishes the complete system prompt, user prompt, the model's actual generated output, and scores at two difficulty levels (Basic 98.0 / Advanced 92.6): a directly reusable Agent execution prompt that says “plan with `<plan>` first, call tools via JSON, review with `<observation>`, and wrap up with `<summary>`.”.DeepSeek-V4-Pro 1M Context Environment Variable Configuration Workflow in Claude CodeWith 8 environment variables, you can point Claude Code (and Claude Desktop Developer Mode) to DeepSeek, unlock a 1M context window with `deepseek-v4-pro[1m]`, use `deepseek-v4-flash` for subagents, set the main model's effort to `max`, and set the automatic compaction window to 786432.