V3.2 and Speciale, official release 2025-12-01; V3.2 App/Web/API with thinking and non-thinking tools, Speciale API-only and no tools at launch; full harness undisclosed.
DeepSeek API Docs / DeepSeek-V3.2 Release · Read evidenceDeepSeek V3.2 · Reviews and evidence
Which DeepSeek V3.2 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
V3.2 technical report v1 dated 2025-12-03; DSA, scalable RL, 1,800+ environments and 85k+ instructions; paper benchmarks differ from an API harness.
arXiv / DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models · Read evidenceV3.2 high and Reasoner on mini-SWE-agent; entries dated 2026-02-17/2025-12-01, 70.00%/$0.45 and 60.00%/$0.03; Verified uses 500 instances, full logs undisclosed.
SWE-bench Leaderboards · Read evidenceSelected evidence
DeepSeek-V3.2 Official Release: Reasoning and Agent Positioning of V3.2 and Speciale
V3.2 and Speciale, official release 2025-12-01; V3.2 App/Web/API with thinking and non-thinking tools, Speciale API-only and no tools at launch; full harness undisclosed.
Unverified: the original source could not be rechecked.
- Conditions
- V3.2 and Speciale, official release 2025-12-01; V3.2 App/Web/API with thinking and non-thinking tools, Speciale API-only and no tools at launch; full harness undisclosed
DeepSeek-V3.2 Technical Report: DSA, Agent Synthetic Data, and Reasoning Baselines
V3.2 technical report v1 dated 2025-12-03; DSA, scalable RL, 1,800+ environments and 85k+ instructions; paper benchmarks differ from an API harness.
Unverified: the original source could not be rechecked.
- Conditions
- V3.2 technical report v1 dated 2025-12-03; DSA, scalable RL, 1,800+ environments and 85k+ instructions; paper benchmarks differ from an API harness
DeepSeek V3.2 Coding Agent Results on the SWE-bench Leaderboard
V3.2 high and Reasoner on mini-SWE-agent; entries dated 2026-02-17/2025-12-01, 70.00%/$0.45 and 60.00%/$0.03; Verified uses 500 instances, full logs undisclosed.
Unverified: the original source could not be rechecked.
- Conditions
- V3.2 high and Reasoner on mini-SWE-agent; entries dated 2026-02-17/2025-12-01, 70.00%/$0.45 and 60.00%/$0.03; Verified uses 500 instances, full logs undisclosed
Reddit LocalLLaMA: Experience Boundaries for DeepSeek V3.2 Agent Coding
V3.2 in a personal Claude Code discussion around Dec 2025; provider, snapshot, tasks, logs and repeats were not pinned; SWE-bench was called approximate.
Unverified: the original source could not be rechecked.
- Conditions
- V3.2 in a personal Claude Code discussion around Dec 2025; provider, snapshot, tasks, logs and repeats were not pinned; SWE-bench was called approximate
All sources
All sources
DeepSeek-V3.2 Official Release: Reasoning and Agent Positioning of V3.2 and Speciale
V3.2 and Speciale, official release 2025-12-01; V3.2 App/Web/API with thinking and non-thinking tools, Speciale API-only and no tools at launch; full harness undisclosed.
Unverified: the original source could not be rechecked.
- Conditions
- V3.2 and Speciale, official release 2025-12-01; V3.2 App/Web/API with thinking and non-thinking tools, Speciale API-only and no tools at launch; full harness undisclosed
DeepSeek-V3.2 Technical Report: DSA, Agent Synthetic Data, and Reasoning Baselines
V3.2 technical report v1 dated 2025-12-03; DSA, scalable RL, 1,800+ environments and 85k+ instructions; paper benchmarks differ from an API harness.
Unverified: the original source could not be rechecked.
- Conditions
- V3.2 technical report v1 dated 2025-12-03; DSA, scalable RL, 1,800+ environments and 85k+ instructions; paper benchmarks differ from an API harness
DeepSeek V3.2 Coding Agent Results on the SWE-bench Leaderboard
V3.2 high and Reasoner on mini-SWE-agent; entries dated 2026-02-17/2025-12-01, 70.00%/$0.45 and 60.00%/$0.03; Verified uses 500 instances, full logs undisclosed.
Unverified: the original source could not be rechecked.
- Conditions
- V3.2 high and Reasoner on mini-SWE-agent; entries dated 2026-02-17/2025-12-01, 70.00%/$0.45 and 60.00%/$0.03; Verified uses 500 instances, full logs undisclosed
Reddit LocalLLaMA: Experience Boundaries for DeepSeek V3.2 Agent Coding
V3.2 in a personal Claude Code discussion around Dec 2025; provider, snapshot, tasks, logs and repeats were not pinned; SWE-bench was called approximate.
Unverified: the original source could not be rechecked.
- Conditions
- V3.2 in a personal Claude Code discussion around Dec 2025; provider, snapshot, tasks, logs and repeats were not pinned; SWE-bench was called approximate
DeepSeek V3.2
Compare DeepSeek V3.2 in Tabbit
Model access, features, and permissions depend on your current client account.