Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

DeepSeek V3.2 · Media / benchmark · Editorial analysis

DeepSeek V3.2 Coding Agent Results on the SWE-bench Leaderboard

V3.2 high and Reasoner on mini-SWE-agent; entries dated 2026-02-17/2025-12-01, 70.00%/$0.45 and 60.00%/$0.03; Verified uses 500 instances, full logs undisclosed.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Conditions
V3.2 high and Reasoner on mini-SWE-agent; entries dated 2026-02-17/2025-12-01, 70.00%/$0.45 and 60.00%/$0.03; Verified uses 500 instances, full logs undisclosed

Key data and applicable tasks

One-sentence takeaway

The official SWE-bench leaderboard's mini-SWE-agent entries show DeepSeek V3.2 high at 70.00% resolved and $0.45 per task, while V3.2 Reasoner reaches 60.00% at $0.03 per task, demonstrating that the Agent harness/version and reasoning configuration can significantly change the result.

Test environment

  • Task set: Real GitHub issues; the original benchmark covers 12 Python repositories, while the expanded page states that multilingual tasks cover 42 repositories across 9 languages.

  • Agent: mini-SWE-agent; different entries use different releases.

  • Metrics: % Resolved, average cost, model, date, and release.

  • Team verification: The page indicates that some entries were run or verified directly by the SWE-bench team.

Input/configuration

  • DeepSeek V3.2 high: mini-SWE-agent 2.0.0, dated 2026-02-17, 70.00%, $0.45/task.

  • DeepSeek V3.2 Reasoner: mini-SWE-agent 1.17.1, dated 2025-12-01, 60.00%, $0.03/task.

Results

In the same table, DeepSeek V3.2 high ranks 14th with 70.00%; Claude 4.5 Sonnet high at 71.40%, Kimi K2.5 high at 70.80%, and GPT 5.2 high at 72.80% provide relative reference points under the same harness. V3.2 Reasoner's 60.00% comes from an earlier release/harness and should not be treated as a pure model difference from 2.0.0.

Conclusion

DeepSeek V3.2 is a strong open-model baseline for Agents that fix real issues; when deploying it, prioritize reproducing the quality/cost of high with the current harness, then evaluate the low-cost Reasoner configuration separately.

Limitations

  • The leaderboard's “DeepSeek V3.2 high” is not equivalent to every API alias or self-hosted checkpoint; Chat/Thinking/tool protocols may differ.

  • The two rows use different dates and mini-SWE-agent versions, so 70 vs. 60 cannot be explained as a pure difference in reasoning effort.

  • The page does not provide per-task prompts, failure traces, token distributions, or confidence intervals; average cost is not the same as your provider bill.

  • Resolved measures only issue fixing and does not cover research, conversation, or long-context question answering.

Reproduction steps

  1. Fix the SWE-bench split, task version, mini-SWE-agent release, model checkpoint, thinking/tool configuration, and timeout.

  2. Run V3.2 high and Reasoner separately, saving patches, test logs, reasoning_content, call traces, tokens, and costs.

  3. Run a closed- or open-model reference under the same harness, and report resolved, timeout, error, and average cost.

  4. Treat the leaderboard row as a verifiable baseline, and state whether your provider/harness matches it.

What this supports

  • supports comparing resolved rate, task cost and reasoning tier

What this does not support

  • does not establish current pricing, general code quality, or tool-free success

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

SWE-bench Leaderboards · SWE-bench team · Original publication date 2026-02-17 · Site edit date 2026-09-20

Open original source

DeepSeek V3.2

Compare DeepSeek V3.2 in Tabbit

Download the Tabbit client to check model access

Related reviews

DeepSeek-V3.2 Technical Report: DSA, Agent Synthetic Data, and Reasoning BaselinesV3.2 technical report v1 dated 2025-12-03; DSA, scalable RL, 1,800+ environments and 85k+ instructions; paper benchmarks differ from an API harness.DeepSeek-V3.2 Official Release: Reasoning and Agent Positioning of V3.2 and SpecialeV3.2 and Speciale, official release 2025-12-01; V3.2 App/Web/API with thinking and non-thinking tools, Speciale API-only and no tools at launch; full harness undisclosed.Reddit LocalLLaMA: Experience Boundaries for DeepSeek V3.2 Agent CodingV3.2 in a personal Claude Code discussion around Dec 2025; provider, snapshot, tasks, logs and repeats were not pinned; SWE-bench was called approximate.DeepSeek-V3.2's Long-Context and Agent Evidence-Anchoring WorkflowDeepSeek's technical report shows that V3.2 uses large-scale environment and complex-instruction synthesis to train Agent generalization; when using it, organize tool results, task constraints, and verifiable outcomes into a trajectory instead of relying on a single “please think autonomously” instruction.DeepSeek V3.2 Thinking Tool Calls and Multi-turn State ConfigurationV3.2 integrates thinking directly into tool use. Multi-turn tool requests must fully pass back the previous turn's `reasoning_content`; otherwise, the API may return an error or lose the reasoning state. The thinking mode should not be configured with temperature/top_p.