Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

DeepSeek V4 Pro · Media / benchmark · Editorial analysis

DeepSeek-V4-Pro XSCT Bench Two-Case Comparison: Strong Planning, Weak Clarification

V4 Pro, XSCT Bench two cases collected 2026-08-21; autonomous planning 98.0/92.6 versus ambiguous clarification 68.5; prompts, repeats, and harness undisclosed.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Conditions
V4 Pro, XSCT Bench two cases collected 2026-08-21; autonomous planning 98.0/92.6 versus ambiguous clarification 68.5; prompts, repeats, and harness undisclosed

Key data and applicable tasks

One-sentence takeaway

A two-testcase comparison on the same platform with the same model under test shows that V4-Pro is close to a perfect score on "autonomous planning and execution" (自主规划执行) (base 98.0 / advanced 92.6), yet manages only 68.5 on the persona testcase for "ambiguous-requirement clarification" (模糊需求澄清) — structured tool planning is a strength; persona-driven dialogue that must proactively probe ambiguous requirements is a weak spot.

Use cases

  • Suitable tasks: Judging whether V4-Pro fits "agent tool tasks that plan first and then execute" (evidence: strong) and "customer-service/assistant-style role-play dialogue that needs to clarify ambiguity" (evidence: weak); serving as a structural reference for writing agent system prompts.

  • Unsuitable tasks: Treating the two testcase scores as an overall capability ranking; treating LLM-as-a-Judge comments as human review conclusions.

  • Applicable model versions: deepseek-v4-pro (the platform's model name; snapshot and effort are not disclosed — see boundaries).

  • Applicable clients, Agents, or APIs: The platform's unified harness; reproduction on the user side requires your own tool-calling environment.

  • Recommended reasoning levels and parameters: Not disclosed; when reproducing, you must record and fix effort/temperature yourself, otherwise cross-comparison is impossible.

Test environment

  • Platform and method: XSCT Bench — LLM-as-a-Judge with multiple judges (CLAUDE/GEMINI/KIMI), evidence anchoring, difficulty tiers, and scoring separated from the model under test; the methodology page states that it "lacks Ground Truth validation" and that "testcase coverage has blind spots".

  • Testcase 1 (l_agent_008, autonomous planning and execution): Agent MCP-type text generation. The system prompt requires leading with <plan>, JSON tool calls, reviewing in <observation>, and closing with <summary>; the user prompt asks the model to continue a multi-step task that involves reading the README, checking config/, and not reading secrets.env. Score: base 98.0 / advanced 92.6, both passed; the advanced item requires fault tolerance for read_file failures and compiling a list of failures.

  • Testcase 2 (l_alice_persona_002, ambiguous-requirement clarification): Persona-driven conversational text generation. The system prompt is the complete persona of Alice, a 26-year-old remote assistant (addressing rules, speaking style, don't guess / don't act rashly, WLB boundaries); the user task contains a double ambiguity around "organizing files" (文件整理). Score: overall 68.5 / base 68.5 — passed, but the review identified core defects.

  • Review summary (testcase 2): CLAUDE — "only asked where the file was, completely missing the ambiguity in the 'organize' action itself"; GEMINI — excellent adherence to persona and taboos, but incomplete clarification; KIMI — "appears to have asked, but didn't get to the point".

  • Additional context: On the platform leaderboard, deepseek-v4-pro scores 87.7 overall (base 89.1 / advanced 87.4 / hard 86.5), tied with kimi-for-coding and GLM-5v-turbo in that tier.

Input/configuration

Both testcase pages disclose the complete system prompt, user prompt, task requirements (scoring anchors), the model's generated output, the multi-judge review comments, and the base/advanced scores. Not disclosed: the tested model snapshot, effort level, temperature, repeat count, and the calling harness.

Results data

TestcaseScoresKey observations
l_agent_008 — autonomous planning and executionBase 98.0 / advanced 92.6; both passedThe model's collected output declares a safety constraint in <plan> (not reading secrets.env), then runs read_file on the README and list_directory on config/ in sequence; on the advanced item it fully completed the "skip on failure + record the reason + compile and write the summary" flow, and the review called it a "high-quality agent execution example"
l_alice_persona_002 — ambiguous-requirement clarificationOverall 68.5 / base 68.5; passed, but with core defects flaggedActual collected output: "Boss, I don't see that file you mentioned on my end — could you send it over? Also, what should I call you?" — it clarified only the missing file, never asking about the ambiguity of "organize". CLAUDE: "only asked where the file was, completely missing the ambiguity in the 'organize' action itself"; GEMINI: excellent persona and taboo adherence, but incomplete clarification; KIMI: "appears to have asked, but didn't get to the point"

Conclusion

V4-Pro is strong at structured tool planning and execution, but weak at persona-driven dialogue that requires proactively clarifying ambiguous requirements; the single-platform, two-testcase evidence is directional, not a general ranking.

Limitations and reproduction steps

  • Limitations: The tested model's snapshot, effort level, temperature, repeat count, and calling harness are not disclosed, and scores change as the platform updates, so reproduction must first lock these variables. The evidence comes from only one platform and two testcases. LLM-as-a-Judge comments may carry systematic bias relative to human judgment (a preference for format completeness and normative tone).

  • Reproduction steps: Rerun the two testcases with your own tool-calling environment while locking effort and temperature (and recording the snapshot) so results can be compared horizontally; treat the page as collected as authoritative, since the platform updates results (the "96" and "83" figures in search summaries differ from the current page values of 98.0/68.5). Single-case scores cannot be generalized into an overall capability ranking.

Original evidence and data

  • Testcase 1 — actual model output (visible at collection): The <plan> declares safety constraints (not reading secrets.env); the model then runs read_file on the README and list_directory on config/ in sequence; on the advanced item it fully completed the "skip on failure + record the reason + compile and write the summary" flow, and the review called it a "high-quality agent execution example".

  • Testcase 2 — full actual model output (visible at collection): "Boss, I don't see that file you mentioned on my end — could you send it over? Also, what should I call you?" — it clarified only that the file was missing, and never asked about the ambiguity of "organize".

  • Both testcase pages disclose: the complete system prompt, user prompt, task requirements (scoring anchors), the generated output, the multi-judge review comments, and the base/advanced scores.

  • The "96" and "83" scores in the search summary do not match the current page values (98.0/68.5), indicating that the platform updates its results; treat the page as collected as authoritative.

Judge review quotes (testcase 2): CLAUDE — "only asked where the file was, completely missing the ambiguity in the 'organize' action itself"; GEMINI — persona and taboo adherence were excellent, but clarification was incomplete; KIMI — "appears to have asked, but didn't get to the point".

Applicability boundaries

  • The tested model's snapshot, effort level, temperature, repeat count, and calling harness are not disclosed; scores change as the platform updates, and reproduction requires locking these variables first.

  • The evidence from one platform and two testcases is limited in strength: it can serve as directional evidence for "strong planning, weak clarification", but cannot replace official benchmarks or self-built evaluations.

  • Testcase 1 has low personalization and heavily constrained output formatting, so high scores come easily; testcase 2 emphasizes persona tone and proactive clarification, exposing a "prompt requirements pulled off track by format" problem rather than pure reasoning ability.

  • LLM-as-a-Judge comments may differ systematically from human judgment (a preference for format completeness and normative tone).

Source excerpt or observation (compliance short quote only)

Testcase 2 KIMI review, verbatim: "The candidate output seriously lost focus on the core task: it handled only a single ambiguity, and did so inappropriately (assuming the file was not received), completely missing the action ambiguity of 'organize', violating the principle of 'don't guess, don't act rashly — ask the most critical question first'."

What this supports

  • supports using the contrast to design planning and clarification prompts

What this does not support

  • does not make two cases overall Agent capability or communication success

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

XSCT Bench (xsctbench.com, scenario-based model-selection evaluation) · XSCT Bench (question writing and platform operation by xsct.ai) · Original publication date Unknown · Site edit date 2026-09-20

Open original source

DeepSeek V4 Pro

Compare DeepSeek V4 Pro in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

DeepSeek V4 Pro: What Changed, What It Costs, and Who Should Use It

A sourced guide to DeepSeek V4 Pro 0813: the agent upgrades, live API limits, price boundary, independent evidence and a safer pilot plan.

Related reviews

DeepSeek-V4-Pro-0813: MindStudio's Eight-Task Coding and Agent Hands-on ComparisonV4 Pro 0813, MindStudio eight-task test on 2026-08-13; 61/80 (76.25%), frontend/planning/math/long-horizon; full prompts, repeats and blind review undisclosed.Artificial Analysis: DeepSeek V4 Pro 0813 (Max Effort) Intelligence Index, Cost, and PositioningThe 2026-08-21 Artificial Analysis snapshot recorded V4 Pro 0813 max effort at index 53, 80.3 tok/s, $3.96/1M output, 1M context, and 1.6T/49B; the page reopened on 2026-09-20 shows index 36, so the snapshots must not be mixed.DeepSeek-V4-Pro Official Release: Reasoning and Agent UpgradesV4 Pro 0813 GA announcement dated 2026-08-13; effort, Responses API and Codex positioning; pricing effective 2026-08-16; no unified benchmark or sample.Reuters DeepSeek-V4-Pro-0813: Official Pricing vs. Independent IndexReuters cites Artificial Analysis's independent pricing and index data: V4-Pro-0813 scores 53 on the reasoning Intelligence Index, versus 40 for V4 Flash, but Pro's input and output prices are roughly 9 and 14 times those of Flash, respectively. Model selection must account for both quality and cost.DeepSeek-V4-Pro Thinking Levels and Tool-Calling WorkflowV4-Pro enables thinking by default and uses high as the default effort level; use low for simple tasks, high for day-to-day Agents, and max for complex tasks, and pass the complete `reasoning_content` back on every round of a tool call.DeepSeek-V4-Pro Responses Configuration Workflow in CodexDeepSeek-V4-Pro can be connected to the Codex CLI, the ChatGPT desktop app, and the VS Code extension through the native Responses API; a single configuration is shared across them, but you should back up and validate `config.toml`/`models.json` first.XSCT Bench “Autonomous Planning and Execution” Case: Agent Tool-Calling Prompt and Generated Result for deepseek-v4-proThe platform publishes the complete system prompt, user prompt, the model's actual generated output, and scores at two difficulty levels (Basic 98.0 / Advanced 92.6): a directly reusable Agent execution prompt that says “plan with `<plan>` first, call tools via JSON, review with `<observation>`, and wrap up with `<summary>`.”.DeepSeek-V4-Pro 1M Context Environment Variable Configuration Workflow in Claude CodeWith 8 environment variables, you can point Claude Code (and Claude Desktop Developer Mode) to DeepSeek, unlock a 1M context window with `deepseek-v4-pro[1m]`, use `deepseek-v4-flash` for subagents, set the main model's effort to `max`, and set the automatic compaction window to 786432.