DeepSeek positions V3.2 as a balanced, everyday Agent model that can call tools in both thinking and non-thinking modes, while positioning V3.2-Speciale as a top-tier reasoning/competition model that did not support tools at launch. They should not be treated as the same model.
Models: DeepSeek-V3.2 and DeepSeek-V3.2-Speciale; V3.2 is available on App/Web/API, while Speciale is API-only.
Capabilities: Reasoning, tool use, Agent data synthesis, and competition mathematics/programming; the official release page does not disclose a complete itemized harness.
Context/cost: The release notes emphasize V3.2's balanced inference vs length; they do not list a complete token/pricing table on that page.
DeepSeek says V3.2 inherits the usage pattern of V3.2-Exp. At the time, V3.2-Speciale was available through a temporary endpoint that ended on 2025-12-15, with the same pricing and no tool calls. Details on V3.2 thinking/tool use point to the Thinking Mode documentation.
V3.2: DeepSeek describes it as offering balanced inference vs length and serving as an everyday driver, with performance at GPT-5 level (official positioning statement).
V3.2-Speciale: DeepSeek describes it as offering maxed-out reasoning, rivaling Gemini-3.0-Pro, and achieving gold-level results in the IMO, CMO, ICPC World Finals, and IOI 2025.
Agent training: Covers 1,800+ environments and 85k+ complex instructions; V3.2 is the first to integrate thinking directly into tool use and supports tool-use modes with and without thinking.
If a workflow needs API tool calls, everyday coding, and a research Agent, choose V3.2 and keep the thinking state fixed. If the comparison is limited to mathematics or extreme reasoning, Speciale can be studied separately, but its competition results should not be treated as evidence of V3.2's Agent capabilities.
The release page presents vendor positioning and selective results, without complete original inputs, failure samples, costs, or variance.
“GPT-5 level” and “rivals Gemini-3.0-Pro” are official descriptions, not independent proof from the same harness.
The Speciale temporary endpoint has expired (according to the release page's timeline), so it should not be used to design a current production integration.
The scale of Agent data synthesis is not a benchmark score.
Lock the currently available model ID; do not use the expired Speciale endpoint.
Run the same tasks separately for V3.2 in thinking and non-thinking modes, recording tool calls, reasoning_content, tokens, latency, and final correctness.
Create a separate competition split for mathematics/programming tasks, explicitly stating whether tools and additional tokens are allowed.
Report V3.2 Agent and Speciale reasoning as two result lines; do not combine them into a single overall score.
DeepSeek V3.2