The official model card shows that DeepSeek-V4.1-Flash is a multimodal MoE with a 552B backbone and 8B (prefill) / 16B (decode) activated per token; with the specified maximum reasoning effort and Agent harness, it scores 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, both above DeepSeek-V4-Flash in the table. Those numbers demonstrate the combined performance of the “model + harness + parameters” configuration and cannot be generalized as uniform capability across all clients or ordinary conversational tasks.
Tasks it is suitable for evaluating: Code Agents, terminal operations, repository modifications, tool calls, automation, and comparisons with the reference models listed in the model card under fixed evaluation conditions; multimodal capability can first be screened with MMMU-Pro, CVBench, DocVQA, and RefCOCO-avg.
Tasks it is not suitable for extrapolating to: Generalizing official Agent scores to arbitrary Agent frameworks, chat products, production systems, real-world long-context business tasks, or safety-critical scenarios; results from different harnesses cannot be treated directly as a fixed score for the same model.
Applicable model version: The open weights for deepseek-ai/DeepSeek-V4.1-Flash on Hugging Face; API versions, product endpoints, and later revisions require separate verification.
Test environment or client: The Hugging Face weights and local inference method listed in the model card; Code Agent results use DeepSeek Harness Minimal mode, DeepSWE v1.1 separately uses the mini-SWE harness as required by its official setup, SEC-Bench Pro uses the Claude Code harness, and the vision Agent uses the Claude Code harness.
Reasoning tier and parameters: The entire Instruct table uses the local-model evaluation setting reasoning_effort=100; temperature=1.0 and top_p=0.95. This is not the parameter syntax for hosted APIs, which use string tiers such as low/high/max. The model card recommends top_p=0.95 or 1.0, a 1M-token context window, and max_tokens ≥ 256K.
The official results are divided into Base Model, Instruct Model, and comparisons across different Agent scaffolds.
All Base Models are evaluated in the official internal framework under the same evaluation settings; the model card treats scores within 0.3 of one another as equivalent. Shot counts vary by benchmark in the tables, so scores from different benchmarks cannot be treated as being on the same scale.
The Instruct Model supports continuous reasoning effort from 1 to 100. The model card explicitly states that the entire table below uses the maximum setting, reasoning_effort=100, with sampling at temperature=1.0 and top_p=0.95. It first summarizes the Code Agent benchmarks as using DeepSeek Harness Minimal mode with a 1M-token context, then explicitly lists the exceptions: DeepSWE v1.1 uses mini-SWE to follow the official setup, while SEC-Bench Pro uses Claude Code. Reproduction should select the harness according to each benchmark's exception notes. Chartography, BabyVision, and ZeroBench use the Claude Code harness with a 512K-token context; Agent's Last Exam and AutomationBench use their respective official scaffolds. The model card says that all Agent evaluations use temperature=1.0 and top_p=0.95.
In the scaffold comparison, DeepSWE v1.1 takes N=8 samples per question and Terminal-Bench 2.1 takes N=3 samples per question; the evaluations use Linux containers, max_steps=500, a 1M context limit, and the same temperature=1.0 and top_p=0.95 sampling. This Terminal-Bench 2.1 evaluation group does not allow network access.
In the model card's Agent comparison table, DeepSeek-V4.1-Flash scores higher than DeepSeek-V4-Flash in the same table on Terminal-Bench 2.1 (90.6 vs 82.7), Terminal-Bench 3.0 (30.0 vs 7.6), Terminal-Bench 4.0 (31.2 vs 7.0), DeepSWE v1.1 (74.2 vs 54.4), CyberGym (88.1 vs 76.7), and AutomationBench (54.8 vs 37.7). HLE's 36.8 and 37.8 with † use different definitions and cannot be compared directly; the scores for the same text-only subset are 39.1† and 37.8†.
The model card also reports results for the same model across scaffolds: DeepSWE v1.1 ranges from 69.8 with Claude Code to 74.2 with mini-SWE, while Terminal-Bench 2.1 ranges from 84.1 with Codex to 90.6 with DSH Minimal. This spread is enough to show that the Agent scaffold changes the observable score.
The following numbers are transcribed from the official model card; — means the model card does not provide a result for that column, and † is the model card's footnote marker.
| Benchmark (metric) | Shots | DeepSeek-V4-Flash-Base | DeepSeek-V4-Pro-Base | DeepSeek-V4.1-Flash-Base |
|---|---|---|---|---|
| Architecture | — | MoE | MoE | MoE |
| # Backbone Params | — | 284B | 1.6T | 552B |
| # Activated Params | — | 13B | 49B | 8B / 16B |
| AGIEval (EM) | 3–5-shot | 83.9 | 84.4 | 83.4 |
| MMLU-Pro (EM) | 5-shot | 68.3 | 73.5 | 74.1 |
| C-Eval (EM) | 5-shot | 92.1 | 93.1 | 92.1 |
| MultiLoKo (LLM-Judge) | 5-shot | 42.6 | 50.9 | 45.5 |
| SimpleQA-Verified (EM) | 25-shot | 30.1 | 55.2 | 42.3 |
| SuperGPQA (EM) | 5-shot | 46.5 | 53.9 | 53.1 |
| BBH (EM) | 3-shot | 86.9 | 87.5 | 86.1 |
| BBEH (EM) | 1-shot | 25.4 | 29.8 | 27.2 |
| DROP (F1) | 1-shot | 88.6 | 88.7 | 87.9 |
| HellaSwag (EM) | 0-shot | 85.7 | 88.0 | 87.2 |
| BigCodeBench (Pass@1) | 3-shot | 56.8 | 59.2 | 60.6 |
| HumanEval (Pass@1) | 0-shot | 69.5 | 76.8 | 79.4 |
| GSM8K (EM) | 8-shot | 90.8 | 92.6 | 93.0 |
| MATH (EM) | 4-shot | 57.4 | 64.5 | 61.1 |
| MGSM (EM) | 8-shot | 85.7 | 84.4 | 80.2 |
| LongBench-V2 (EM) | 1-shot | 44.7 | 51.5 | 45.2 |
| MMMU-Pro (EM) | 4-shot | — | — | 56.5 |
| CVBench (EM) | 4-shot | — | — | 77.9 |
| DocVQA (LLM-Judge) | 4-shot | — | — | 95.6 |
| RefCOCO-avg (Acc@0.5) | 0-shot | — | — | 86.0 |
| Benchmark (metric) | Opus-5.0 | GPT-5.6 Sol | K3 | GLM-5.3 | DS-V4-Pro | DS-V4-Flash | DS-V4.1-Flash |
|---|---|---|---|---|---|---|---|
| GPQA Diamond (Pass@1) | 93.4 | 94.1 | 92.9 | 88.1 | 92.4 | 89.9 | 90.9 |
| HLE (Pass@1) | 56.3 | 44.5 | 43.5 | 42.0† | 42.7† | 37.8† | 36.8 (39.1†) |
| Codeforces (Rating) | — | — | — | — | 3348 | 3289 | 3471 |
| MathArena Apex (Pass@1) | — | — | 65.6 | — | 65.3 | 58.6 | 65.6 |
| Terminal-Bench 2.1 (Pass@1) | 89.1 | 88.8 | 88.3 | 88.2 | 87.9 | 82.7 | 90.6 |
| Terminal-Bench 3.0 (Pass@1) | 43.3 | 34.4 | 17.7 | 28.3 | 11.8 | 7.6 | 30.0 |
| Terminal-Bench 4.0 (Pass@1) | 51.8 | 39.9 | 12.6 | 37.9 | 12.4 | 7.0 | 31.2 |
| DeepSWE v1.1 (Resolved) | 74.0 | 73.0 | 67.5 | 66.9 | 62.7 | 54.4 | 74.2 |
| ProgramBench (Almost@1) | 37.0 | 23.0 | 17.5 | 19.0 | 15.5 | — | 20.3 |
| NL2Repo-Bench (Score) | 75.3 | 56.8 | 58.0 | 58.0 | 61.5 | 54.2 | 64.0 |
| CyberGym (Pass@1) | — | 84.5 | 80.0 | 84.5 | 83.3 | 76.7 | 88.1 |
| SEC-Bench Pro (Pass@1) | — | 74.3 | — | — | 56.4 | 30.9 | 62.8 |
| ExploitGym (Pass@1) | 22.1 | 33.7 | — | 15.0 | 5.4 | 1.8 | 15.3 |
| HLE w/ tools (Pass@1) | 63.6 | — | 59.8 | 62.5 | 60.0 | 51.5 | 63.9 |
| AutomationBench (Pass@1) | 50.3 | 45.8 | 46.7 | 48.8 | 43.2 | 37.7 | 54.8 |
| Agent's Last Exam (Pass@1) | 28.6 | 26.7 | 27.6 | 28.5 | 25.7 | 25.2 | 31.8 |
| Chartography w/ tools (Pass@1) | 84.0 | 79.9 | 68.1 | — | — | — | 78.9 |
| BabyVision w/ tools (Pass@1) | 94.1 | 88.9 | 85.7 | — | — | — | 89.6 |
| ZeroBench-main w/ tools (Pass@5) | 52.0 | 53.0 | 41.0 | — | — | — | 49.0 |
† Official footnote: HLE text-only subset.
| Benchmark (metric) | Claude Code | Codex | OpenCode | Pi | mini-SWE | DSH Minimal | DSH Standard | DSH PTC |
|---|---|---|---|---|---|---|---|---|
| DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |
| Terminal-Bench 2.1 (Pass@1) | 88.0 | 84.1 | 85.0 | 86.1 | 90.3 | 90.6 | 85.8 | 85.8 |
The strongest signal in the official results is within the Agent task bucket. Under the fixed settings listed in the model card, V4.1-Flash scores higher than V4-Flash on Terminal-Bench, DeepSWE, CyberGym, and AutomationBench, and has a Codeforces rating of 3471; this supports considering it as a candidate model for Code Agents and terminal tasks.
Benchmarks are not a unified capability score. The Base table uses different shot counts and metrics; the Instruct table mixes Pass@1, Pass@5, Resolved, Rating, Score, and LLM-Judge, so it cannot be sorted across columns to produce a single “overall capability” score.
The harness changes the result. For the same model, scaffold results range from 65.5 to 74.2 on DeepSWE v1.1 and from 84.1 to 90.6 on Terminal-Bench 2.1. Any repeat evaluation must fix the Agent scaffold, tool protocol, step count, context, network access, and sampling parameters.
Maximum reasoning effort limits extrapolation. All official Instruct results use reasoning_effort=100, which does not represent performance at lower effort or the default setting; the model card provides no curves showing how each benchmark changes with reasoning effort.
Long-context support does not equal long-task performance. The model card gives a 1M-token context window and a LongBench-V2 Base result of 45.2, but provides no business data, long-session failure samples, or complete variance, so it cannot guarantee any arbitrary 1M-token workflow.
The official table does not provide all information required for independent verification. The page does not publish the complete inputs, outputs, repeated experiments, confidence intervals, or all harness implementations for each benchmark in the table; some comparison columns are external model names, and the evaluation sources and run details cannot be fully reconstructed from the table alone.
Efficiency and capability require separate validation. The model card states a 552B total backbone, 8B/16B activated parameters, and a 1M context, but does not provide standardized hardware, quantization, concurrency, throughput, or latency data in this results table, so production cost cannot be inferred directly from parameter scale.
Fix the commit of the Hugging Face model repository, weight precision, tokenizer, inference engine, and prompt encoding; the model card says this version has no Jinja chat template, and the repository provides encoding/encoding.py as a reference implementation.
For Instruct reproduction, start with the model card parameters: reasoning_effort=100, temperature=1.0, top_p=0.95, and a 1M context; also record max_tokens, and do not mix results with different context limits.
Run Code Agents separately with the harnesses specified by the model card: DeepSeek Harness Minimal mode, mini-SWE, Claude Code, and each benchmark's official scaffold; for Terminal-Bench 2.1 reproduction, disable network access and record the Linux container, N, max_steps=500, and tool-call logs.
Save the raw input, model output, patch, test results, failure reason, tool-call count, and token usage for every task; calculate metrics such as Pass@1, Pass@5, Resolved, and Rating separately, without merging them into an overall score.
First reproduce the public DeepSWE v1.1 and Terminal-Bench 2.1 results, then repeat the same task sets in the target production Agent; report the difference between the official results and those from your own harness, and mark the collection date and model commit.
DeepSeek V4.1 Flash