DeepSeek V3.2 · Media / benchmark · Editorial analysis
V3.2 high and Reasoner on mini-SWE-agent; entries dated 2026-02-17/2025-12-01, 70.00%/$0.45 and 60.00%/$0.03; Verified uses 500 instances, full logs undisclosed.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
The official SWE-bench leaderboard's mini-SWE-agent entries show DeepSeek V3.2 high at 70.00% resolved and $0.45 per task, while V3.2 Reasoner reaches 60.00% at $0.03 per task, demonstrating that the Agent harness/version and reasoning configuration can significantly change the result.
Task set: Real GitHub issues; the original benchmark covers 12 Python repositories, while the expanded page states that multilingual tasks cover 42 repositories across 9 languages.
Agent: mini-SWE-agent; different entries use different releases.
Metrics: % Resolved, average cost, model, date, and release.
Team verification: The page indicates that some entries were run or verified directly by the SWE-bench team.
DeepSeek V3.2 high: mini-SWE-agent 2.0.0, dated 2026-02-17, 70.00%, $0.45/task.
DeepSeek V3.2 Reasoner: mini-SWE-agent 1.17.1, dated 2025-12-01, 60.00%, $0.03/task.
In the same table, DeepSeek V3.2 high ranks 14th with 70.00%; Claude 4.5 Sonnet high at 71.40%, Kimi K2.5 high at 70.80%, and GPT 5.2 high at 72.80% provide relative reference points under the same harness. V3.2 Reasoner's 60.00% comes from an earlier release/harness and should not be treated as a pure model difference from 2.0.0.
DeepSeek V3.2 is a strong open-model baseline for Agents that fix real issues; when deploying it, prioritize reproducing the quality/cost of high with the current harness, then evaluate the low-cost Reasoner configuration separately.
The leaderboard's “DeepSeek V3.2 high” is not equivalent to every API alias or self-hosted checkpoint; Chat/Thinking/tool protocols may differ.
The two rows use different dates and mini-SWE-agent versions, so 70 vs. 60 cannot be explained as a pure difference in reasoning effort.
The page does not provide per-task prompts, failure traces, token distributions, or confidence intervals; average cost is not the same as your provider bill.
Resolved measures only issue fixing and does not cover research, conversation, or long-context question answering.
Fix the SWE-bench split, task version, mini-SWE-agent release, model checkpoint, thinking/tool configuration, and timeout.
Run V3.2 high and Reasoner separately, saving patches, test logs, reasoning_content, call traces, tokens, and costs.
Run a closed- or open-model reference under the same harness, and report resolved, timeout, error, and average cost.
Treat the leaderboard row as a verifiable baseline, and state whether your provider/harness matches it.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
SWE-bench Leaderboards · SWE-bench team · Original publication date 2026-02-17 · Site edit date 2026-09-20
Open original sourceDeepSeek V3.2
Download the Tabbit client to check model access