Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaDeepSeek V4.1 Flash

DeepSeek-V4.1-Flash Official Model Card Benchmarks: Agent Strengths and Harness Boundaries

Original source

Hugging Face (DeepSeek official model card)

AuthorDeepSeek-AI

Tabbit curation2026-09-16

Read original

One-sentence takeaway

The official model card shows that DeepSeek-V4.1-Flash is a multimodal MoE with a 552B backbone and 8B (prefill) / 16B (decode) activated per token; with the specified maximum reasoning effort and Agent harness, it scores 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, both above DeepSeek-V4-Flash in the table. Those numbers demonstrate the combined performance of the “model + harness + parameters” configuration and cannot be generalized as uniform capability across all clients or ordinary conversational tasks.

Use cases

  • Tasks it is suitable for evaluating: Code Agents, terminal operations, repository modifications, tool calls, automation, and comparisons with the reference models listed in the model card under fixed evaluation conditions; multimodal capability can first be screened with MMMU-Pro, CVBench, DocVQA, and RefCOCO-avg.

  • Tasks it is not suitable for extrapolating to: Generalizing official Agent scores to arbitrary Agent frameworks, chat products, production systems, real-world long-context business tasks, or safety-critical scenarios; results from different harnesses cannot be treated directly as a fixed score for the same model.

  • Applicable model version: The open weights for deepseek-ai/DeepSeek-V4.1-Flash on Hugging Face; API versions, product endpoints, and later revisions require separate verification.

  • Test environment or client: The Hugging Face weights and local inference method listed in the model card; Code Agent results use DeepSeek Harness Minimal mode, DeepSWE v1.1 separately uses the mini-SWE harness as required by its official setup, SEC-Bench Pro uses the Claude Code harness, and the vision Agent uses the Claude Code harness.

  • Reasoning tier and parameters: The entire Instruct table uses the local-model evaluation setting reasoning_effort=100; temperature=1.0 and top_p=0.95. This is not the parameter syntax for hosted APIs, which use string tiers such as low/high/max. The model card recommends top_p=0.95 or 1.0, a 1M-token context window, and max_tokens ≥ 256K.

Evaluation method

The official results are divided into Base Model, Instruct Model, and comparisons across different Agent scaffolds.

All Base Models are evaluated in the official internal framework under the same evaluation settings; the model card treats scores within 0.3 of one another as equivalent. Shot counts vary by benchmark in the tables, so scores from different benchmarks cannot be treated as being on the same scale.

The Instruct Model supports continuous reasoning effort from 1 to 100. The model card explicitly states that the entire table below uses the maximum setting, reasoning_effort=100, with sampling at temperature=1.0 and top_p=0.95. It first summarizes the Code Agent benchmarks as using DeepSeek Harness Minimal mode with a 1M-token context, then explicitly lists the exceptions: DeepSWE v1.1 uses mini-SWE to follow the official setup, while SEC-Bench Pro uses Claude Code. Reproduction should select the harness according to each benchmark's exception notes. Chartography, BabyVision, and ZeroBench use the Claude Code harness with a 512K-token context; Agent's Last Exam and AutomationBench use their respective official scaffolds. The model card says that all Agent evaluations use temperature=1.0 and top_p=0.95.

In the scaffold comparison, DeepSWE v1.1 takes N=8 samples per question and Terminal-Bench 2.1 takes N=3 samples per question; the evaluations use Linux containers, max_steps=500, a 1M context limit, and the same temperature=1.0 and top_p=0.95 sampling. This Terminal-Bench 2.1 evaluation group does not allow network access.

Key results

In the model card's Agent comparison table, DeepSeek-V4.1-Flash scores higher than DeepSeek-V4-Flash in the same table on Terminal-Bench 2.1 (90.6 vs 82.7), Terminal-Bench 3.0 (30.0 vs 7.6), Terminal-Bench 4.0 (31.2 vs 7.0), DeepSWE v1.1 (74.2 vs 54.4), CyberGym (88.1 vs 76.7), and AutomationBench (54.8 vs 37.7). HLE's 36.8 and 37.8 with † use different definitions and cannot be compared directly; the scores for the same text-only subset are 39.1† and 37.8†.

The model card also reports results for the same model across scaffolds: DeepSWE v1.1 ranges from 69.8 with Claude Code to 74.2 with mini-SWE, while Terminal-Bench 2.1 ranges from 84.1 with Codex to 90.6 with DSH Minimal. This spread is enough to show that the Agent scaffold changes the observable score.

Raw data

The following numbers are transcribed from the official model card; — means the model card does not provide a result for that column, and † is the model card's footnote marker.

Base Model

Benchmark (metric)ShotsDeepSeek-V4-Flash-BaseDeepSeek-V4-Pro-BaseDeepSeek-V4.1-Flash-Base
Architecture—MoEMoEMoE
# Backbone Params—284B1.6T552B
# Activated Params—13B49B8B / 16B
AGIEval (EM)3–5-shot83.984.483.4
MMLU-Pro (EM)5-shot68.373.574.1
C-Eval (EM)5-shot92.193.192.1
MultiLoKo (LLM-Judge)5-shot42.650.945.5
SimpleQA-Verified (EM)25-shot30.155.242.3
SuperGPQA (EM)5-shot46.553.953.1
BBH (EM)3-shot86.987.586.1
BBEH (EM)1-shot25.429.827.2
DROP (F1)1-shot88.688.787.9
HellaSwag (EM)0-shot85.788.087.2
BigCodeBench (Pass@1)3-shot56.859.260.6
HumanEval (Pass@1)0-shot69.576.879.4
GSM8K (EM)8-shot90.892.693.0
MATH (EM)4-shot57.464.561.1
MGSM (EM)8-shot85.784.480.2
LongBench-V2 (EM)1-shot44.751.545.2
MMMU-Pro (EM)4-shot——56.5
CVBench (EM)4-shot——77.9
DocVQA (LLM-Judge)4-shot——95.6
RefCOCO-avg (Acc@0.5)0-shot——86.0

Instruct Model: frontier comparison (maximum reasoning effort)

Benchmark (metric)Opus-5.0GPT-5.6 SolK3GLM-5.3DS-V4-ProDS-V4-FlashDS-V4.1-Flash
GPQA Diamond (Pass@1)93.494.192.988.192.489.990.9
HLE (Pass@1)56.344.543.542.0†42.7†37.8†36.8 (39.1†)
Codeforces (Rating)————334832893471
MathArena Apex (Pass@1)——65.6—65.358.665.6
Terminal-Bench 2.1 (Pass@1)89.188.888.388.287.982.790.6
Terminal-Bench 3.0 (Pass@1)43.334.417.728.311.87.630.0
Terminal-Bench 4.0 (Pass@1)51.839.912.637.912.47.031.2
DeepSWE v1.1 (Resolved)74.073.067.566.962.754.474.2
ProgramBench (Almost@1)37.023.017.519.015.5—20.3
NL2Repo-Bench (Score)75.356.858.058.061.554.264.0
CyberGym (Pass@1)—84.580.084.583.376.788.1
SEC-Bench Pro (Pass@1)—74.3——56.430.962.8
ExploitGym (Pass@1)22.133.7—15.05.41.815.3
HLE w/ tools (Pass@1)63.6—59.862.560.051.563.9
AutomationBench (Pass@1)50.345.846.748.843.237.754.8
Agent's Last Exam (Pass@1)28.626.727.628.525.725.231.8
Chartography w/ tools (Pass@1)84.079.968.1———78.9
BabyVision w/ tools (Pass@1)94.188.985.7———89.6
ZeroBench-main w/ tools (Pass@5)52.053.041.0———49.0

† Official footnote: HLE text-only subset.

Across Agent scaffolds

Benchmark (metric)Claude CodeCodexOpenCodePimini-SWEDSH MinimalDSH StandardDSH PTC
DeepSWE v1.1 (Resolved)69.865.665.566.274.272.670.567.6
Terminal-Bench 2.1 (Pass@1)88.084.185.086.190.390.685.885.8

Conclusions and limitations

  1. The strongest signal in the official results is within the Agent task bucket. Under the fixed settings listed in the model card, V4.1-Flash scores higher than V4-Flash on Terminal-Bench, DeepSWE, CyberGym, and AutomationBench, and has a Codeforces rating of 3471; this supports considering it as a candidate model for Code Agents and terminal tasks.

  2. Benchmarks are not a unified capability score. The Base table uses different shot counts and metrics; the Instruct table mixes Pass@1, Pass@5, Resolved, Rating, Score, and LLM-Judge, so it cannot be sorted across columns to produce a single “overall capability” score.

  3. The harness changes the result. For the same model, scaffold results range from 65.5 to 74.2 on DeepSWE v1.1 and from 84.1 to 90.6 on Terminal-Bench 2.1. Any repeat evaluation must fix the Agent scaffold, tool protocol, step count, context, network access, and sampling parameters.

  4. Maximum reasoning effort limits extrapolation. All official Instruct results use reasoning_effort=100, which does not represent performance at lower effort or the default setting; the model card provides no curves showing how each benchmark changes with reasoning effort.

  5. Long-context support does not equal long-task performance. The model card gives a 1M-token context window and a LongBench-V2 Base result of 45.2, but provides no business data, long-session failure samples, or complete variance, so it cannot guarantee any arbitrary 1M-token workflow.

  6. The official table does not provide all information required for independent verification. The page does not publish the complete inputs, outputs, repeated experiments, confidence intervals, or all harness implementations for each benchmark in the table; some comparison columns are external model names, and the evaluation sources and run details cannot be fully reconstructed from the table alone.

  7. Efficiency and capability require separate validation. The model card states a 552B total backbone, 8B/16B activated parameters, and a 1M context, but does not provide standardized hardware, quantization, concurrency, throughput, or latency data in this results table, so production cost cannot be inferred directly from parameter scale.

Reproduction notes

  1. Fix the commit of the Hugging Face model repository, weight precision, tokenizer, inference engine, and prompt encoding; the model card says this version has no Jinja chat template, and the repository provides encoding/encoding.py as a reference implementation.

  2. For Instruct reproduction, start with the model card parameters: reasoning_effort=100, temperature=1.0, top_p=0.95, and a 1M context; also record max_tokens, and do not mix results with different context limits.

  3. Run Code Agents separately with the harnesses specified by the model card: DeepSeek Harness Minimal mode, mini-SWE, Claude Code, and each benchmark's official scaffold; for Terminal-Bench 2.1 reproduction, disable network access and record the Linux container, N, max_steps=500, and tool-call logs.

  4. Save the raw input, model output, patch, test results, failure reason, tool-call count, and token usage for every task; calculate metrics such as Pass@1, Pass@5, Resolved, and Rating separately, without merging them into an overall score.

  5. First reproduce the public DeepSWE v1.1 and Terminal-Bench 2.1 results, then repeat the same task sets in the target production Agent; report the difference between the official results and those from your own harness, and mark the collection date and model commit.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

DeepSeek V4.1 Flash

Use and compare models in Tabbit

DeepSeek V4.1 Flash

Related reviews

CommunityX2026-09-15

DeepSeek-V4.1-Flash (Max): Task Cost and Net Improvement in Agent Arena

CommunityX (Artificial Analysis)2026-09-11

Artificial Analysis: DeepSeek V4.1 Flash's Intelligence, Cost, and Hallucination Boundaries

MediaAI IQ

DeepSeek V4.1 Flash on the AI IQ Leaderboard: Composite Score and Benchmark Coverage

CommunityReddit, r/DeepSeek2026-09-15

DeepSeek V4.1 Flash Fixing Legacy Code in OpenCode: Community Experience and False-Positive Boundaries

DeepSeek V4.1 Flash

Related prompts

OfficialDeepSeek API Docs

DeepSeek-V4.1-Flash: API Model Aliases and First Call

OfficialDeepSeek API Docs

DeepSeek V4.1 Flash Thinking Mode and Reasoning Parameter Configuration

OfficialDeepSeek API Docs

DeepSeek V4.1 Flash: Image Input and Vision Configuration

OfficialDeepSeek API Docs

DeepSeek V4.1 Flash: JSON Question-and-Answer Extraction Prompt