Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

DeepSeek V3.2 · Media / benchmark · Editorial analysis

DeepSeek-V3.2 Technical Report: DSA, Agent Synthetic Data, and Reasoning Baselines

V3.2 technical report v1 dated 2025-12-03; DSA, scalable RL, 1,800+ environments and 85k+ instructions; paper benchmarks differ from an API harness.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Conditions
V3.2 technical report v1 dated 2025-12-03; DSA, scalable RL, 1,800+ environments and 85k+ instructions; paper benchmarks differ from an API harness

Key data and applicable tasks

One-sentence takeaway

The technical report attributes V3.2's advantages to sparse attention, scalable RL, and large-scale Agent task synthesis. The goal is to reduce costs and improve tool generalization in long contexts, but the report's benchmark and API production performance still require independent reproduction.

Test environment

  • Models: DeepSeek-V3.2, V3.2-Exp, and V3.2-Speciale; the report covers standard reasoning, long-context, Agent, and human-preference evaluations.

  • Architecture: DSA (lightning indexer + fine-grained top-k selection), followed by specialist distillation and mixed RL after continued pretraining.

  • Training/evaluation data: The Agent synthesis pipeline covers 1,800+ environments and 85,000+ complex instructions; the report also mentions 128K long-context training and H800 inference cost estimates.

Inputs/configuration

The report's detailed benchmark tables/figures are not fully available as reproducible text. Its method description includes specialized thinking/non-thinking tracks, six domains (mathematics, programming, logic, general Agent, Agentic coding, and search), and training factors such as outcome reward, length penalty, and language consistency reward.

Results

  • The report's abstract states that V3.2 is close to GPT-5 on multiple reasoning benchmarks, while V3.2-Speciale surpasses GPT-5 and approaches Gemini-3.0-Pro; Speciale achieved gold-medal performance at IMO/IOI 2025.

  • DSA reduces the main attention complexity from O(L²) to O(Lk), with each query selecting 2,048 key-value tokens (the report's sparse-training configuration).

  • Continued pretraining's sparse stage: 1,500 steps, 480×128K sequences per step, and approximately 943.7B tokens; these are training settings, not parameters required for deployment.

  • The report says V3.2-Exp scored 4 points higher than V3.1-Terminus on Artificial Analysis Long Context Reasoning; it also cites multiple areas of leadership on Fiction.liveBench, but these are external evaluations cited by the report, not independent reproductions in this review.

Conclusion

V3.2's architecture and training approach target long-context efficiency and Agent generalization, making it suitable as an open-weight Agent baseline for research. Actual services should nevertheless be validated separately across short/long contexts, thinking/tool use, and provider parsers.

Limitations

  • This is a self-reported DeepSeek technical report. Benchmark selection, prompts, scaffolding, and evaluation details are incomplete; conclusions about rankings should not be based on the abstract alone.

  • Training token counts, top-k, and DSA complexity do not directly determine API costs or user-visible quality.

  • The V3.2, V3.2-Exp, and Speciale checkpoints and capabilities differ; comparisons across variants can be confounded.

  • Artificial Analysis and Fiction.liveBench results should be checked against their original pages.

Reproduction steps

  1. Select the model variant, weights/service, and task split; fix thinking, tools, context length, and output budget.

  2. Run standard reasoning, coding, long-context, tool-calling, and Agent-outcome tasks in separate layers.

  3. For long contexts, record tokens/s, time to first token, GPU/cost, and accuracy; for Agents, save complete tool traces.

  4. Compare GPT and other open models using the same harness, distinguishing figures from the official report, external results, and your own reproduction results.

What this supports

  • supports the training rationale around DSA and Agent-synthetic data

What this does not support

  • does not derive your success rate, latency, or cost from training scale

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

arXiv / DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models · DeepSeek-AI · Original publication date 2025-12-03 · Site edit date 2026-09-20

Open original source

DeepSeek V3.2

Compare DeepSeek V3.2 in Tabbit

Download the Tabbit client to check model access

Related reviews

DeepSeek V3.2 Coding Agent Results on the SWE-bench LeaderboardV3.2 high and Reasoner on mini-SWE-agent; entries dated 2026-02-17/2025-12-01, 70.00%/$0.45 and 60.00%/$0.03; Verified uses 500 instances, full logs undisclosed.DeepSeek-V3.2 Official Release: Reasoning and Agent Positioning of V3.2 and SpecialeV3.2 and Speciale, official release 2025-12-01; V3.2 App/Web/API with thinking and non-thinking tools, Speciale API-only and no tools at launch; full harness undisclosed.Reddit LocalLLaMA: Experience Boundaries for DeepSeek V3.2 Agent CodingV3.2 in a personal Claude Code discussion around Dec 2025; provider, snapshot, tasks, logs and repeats were not pinned; SWE-bench was called approximate.DeepSeek-V3.2's Long-Context and Agent Evidence-Anchoring WorkflowDeepSeek's technical report shows that V3.2 uses large-scale environment and complex-instruction synthesis to train Agent generalization; when using it, organize tool results, task constraints, and verifiable outcomes into a trajectory instead of relying on a single “please think autonomously” instruction.DeepSeek V3.2 Thinking Tool Calls and Multi-turn State ConfigurationV3.2 integrates thinking directly into tool use. Multi-turn tool requests must fully pass back the previous turn's `reasoning_content`; otherwise, the API may return an error or lose the reasoning state. The thinking mode should not be configured with temperature/top_p.