The official SWE-bench leaderboard's mini-SWE-agent entries show DeepSeek V3.2 high at 70.00% resolved and $0.45 per task, while V3.2 Reasoner reaches 60.00% at $0.03 per task, demonstrating that the Agent harness/version and reasoning configuration can significantly change the result.
Task set: Real GitHub issues; the original benchmark covers 12 Python repositories, while the expanded page states that multilingual tasks cover 42 repositories across 9 languages.
Agent: mini-SWE-agent; different entries use different releases.
Metrics: % Resolved, average cost, model, date, and release.
Team verification: The page indicates that some entries were run or verified directly by the SWE-bench team.
DeepSeek V3.2 high: mini-SWE-agent 2.0.0, dated 2026-02-17, 70.00%, $0.45/task.
DeepSeek V3.2 Reasoner: mini-SWE-agent 1.17.1, dated 2025-12-01, 60.00%, $0.03/task.
In the same table, DeepSeek V3.2 high ranks 14th with 70.00%; Claude 4.5 Sonnet high at 71.40%, Kimi K2.5 high at 70.80%, and GPT 5.2 high at 72.80% provide relative reference points under the same harness. V3.2 Reasoner's 60.00% comes from an earlier release/harness and should not be treated as a pure model difference from 2.0.0.
DeepSeek V3.2 is a strong open-model baseline for Agents that fix real issues; when deploying it, prioritize reproducing the quality/cost of high with the current harness, then evaluate the low-cost Reasoner configuration separately.
The leaderboard's “DeepSeek V3.2 high” is not equivalent to every API alias or self-hosted checkpoint; Chat/Thinking/tool protocols may differ.
The two rows use different dates and mini-SWE-agent versions, so 70 vs. 60 cannot be explained as a pure difference in reasoning effort.
The page does not provide per-task prompts, failure traces, token distributions, or confidence intervals; average cost is not the same as your provider bill.
Resolved measures only issue fixing and does not cover research, conversation, or long-context question answering.
Fix the SWE-bench split, task version, mini-SWE-agent release, model checkpoint, thinking/tool configuration, and timeout.
Run V3.2 high and Reasoner separately, saving patches, test logs, reasoning_content, call traces, tokens, and costs.
Run a closed- or open-model reference under the same harness, and report resolved, timeout, error, and average cost.
Treat the leaderboard row as a verifiable baseline, and state whether your provider/harness matches it.
DeepSeek V3.2