At the Reasoning, Max Effort setting, Artificial Analysis gives DeepSeek V4.1 Flash an Artificial Analysis Intelligence Index score of about 40. It is attractive for agent tasks, long-context work, and cost, but its outputs are extremely verbose, and its AA-Omniscience Non-Hallucination Rate is only 4%, so “cheap” cannot be equated directly with reliable.
Tasks it is suitable for evaluating: Cross-model comparisons under the fixed AA testing methodology; agent workflows, long-context tasks, SaaS operations, and cost-sensitive API selection.
Tasks it is not suitable for extrapolating to: Treating the AA score as a universal measure of ordinary chat, production workloads, or all reasoning tasks; extrapolating the cost of a single task to different input lengths, cache hit rates, output lengths, or off-peak pricing.
Applicable model version: DeepSeek V4.1 Flash; the auxiliary page corresponds to DeepSeek V4.1 Flash (Reasoning, Max Effort), identifies it as an open-weight model, and lists a 1M-token context window.
Test environment or client: An independent Artificial Analysis evaluation; the auxiliary page says the measurements were completed on dedicated hardware. The specific hardware, prompts, sample count, number of repetitions, temperature, and tool configuration were not specified.
Reasoning setting and parameters: Reasoning, Max Effort (page short name: max). The main post records the official API prices: $0.30/1M tokens for input, $1.20/1M tokens for output, and $0.006/1M tokens for cached input (98% discount); an additional 50% discount applies during off-peak periods. The thread does not provide temperature, top_p, maximum output, or tool parameters.
Auxiliary sources: the AA model page, the task cost explanation from the same author, and the AA-Omniscience results. These supplement the same evaluation and are not counted as separate sources.
The main post compares DeepSeek V4.1 Flash with DeepSeek V4 Pro 0813, V4 Flash 0731, and other models using Artificial Analysis's independent metrics. The metrics include Intelligence Index, GDPval-AA v2, AA-LCR v1.1, AutomationBench-AA, Terminal-Bench v4.0, Token Use, Cost per Task, and AA-Omniscience. In a page snapshot from 2026-09-16, the auxiliary model page shows the current Artificial Analysis Intelligence Index as v4.3, comprising 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.
The auxiliary page lists this model as Reasoning, Max Effort and says the evaluation results were measured independently by Artificial Analysis. The page does not disclose the complete questions, prompts, sampling repetitions, confidence intervals, or detailed harness for each evaluation. The results below should therefore be treated as a snapshot under this methodology and are insufficient to independently reconstruct the full experiment.
Overall metric: The AI Index shows a score of 40; the exact value retained by the auxiliary page is 39.5454, ranked #6/113. The main post says it surpasses DeepSeek V4 Pro 0813 (1.6T), but its verbosity places it slightly below the Intelligence vs. Cost Pareto frontier.
Agents and long context: Terminal-Bench v4.0 27%; GDPval-AA v2 rose from 1468 to 1632 Elo; AA-LCR v1.1 was 84%, tied with GPT-5.6 Sol and Gemini 3.8 Flash at 84%. AutomationBench-AA was 69%, tied with GPT-6 Astra and above Grok 4.6 at 67%, V4 Flash 0731 at 54%, V4 Pro 0813 at 57%, and GLM-5.3 at 62%. This metric covers 657 tasks across 39 SaaS applications and scores whether the goal was completed without triggering guardrails.
Cost and verbosity: Each Intelligence Index task used about 89k tokens, with a task cost of $0.27. The main post says this is about one-seventh the cost of GLM-5.3 ($2.01) and Kimi K3 ($2.00), and about 1/2.5 the cost of V4 Pro 0813 ($0.67). The auxiliary page gives an exact task cost of $0.2652, an output speed of 209.8 tokens/s (rounded to 210 on the page), 250M tokens used cumulatively for the evaluation, and a cost of $476.89 to run the full Intelligence Index.
Hallucination boundary: A subsequent post by the same author says that the AA-Omniscience metric “increased by 9 points to -5.3”; accuracy rose from 40% to 46%, while the Non-Hallucination Rate fell from 8% to 4%. The auxiliary page explains that the index ranges from -100 to 100, with 0 meaning that the numbers of correct and incorrect answers are equal, and a negative value meaning that errors outnumber correct answers. Therefore, -5.3 cannot be interpreted as reliability having become positive.
| Item | Original figure or wording |
|---|---|
| Artificial Analysis Intelligence Index | 40 (exact auxiliary-page value: 39.5454) |
| GDPval-AA v2 | 1468 → 1632 Elo, +164 Elo |
| AA-LCR v1.1 | 84% |
| AutomationBench-AA | 69%; GPT-6 Astra 69%, Grok 4.6 67% |
| Terminal-Bench v4.0 | 27%; the main post also writes V4.0 Flash 27% and says “more than double,” so the two statements are inconsistent |
| Intelligence Index Token Use | 89k tokens per task |
| Cost per Intelligence Index task | $0.27 |
| API pricing | $0.30 input, $1.20 output, $0.006 cached input / 1M tokens; 98% cache discount |
| Context window | 1M tokens |
| AA-Omniscience | Index -5.3; accuracy 40% → 46%; Non-Hallucination Rate 8% → 4% |
| Metric | Page value |
|---|---|
| Displayed model setting | Reasoning, Max Effort / max |
| Artificial Analysis Intelligence Index | 40; exact value 39.5454; #6/113 |
| Speed | 209.8 output tokens/s; #4/113 |
| Cost | $0.30 input, $1.20 output, 98% cache discount; $0.27/task; #19/113 |
| Verbosity | 250M output tokens from Intelligence Index; #37/113 |
| Index version | v4.3, 10 evaluations |
| Context window | 1M tokens |
| Total / active parameters | Page lists 552B / 16B |
AA-Omniscience's Non-Hallucination Rate and overall Accuracy are different metrics; 4% cannot be interpreted as meaning that only 4% of all the model's answers are correct, and it cannot replace the 46% accuracy reported on the same page.
The cost advantage comes from the pricing and caching methodology, not low output volume. The main post explicitly reports the $0.27 per-task cost alongside the verbosity of 89k tokens per task. Budgets should therefore be calculated from actual output length, cache hit rate, and whether off-peak pricing applies.
Agent performance cannot replace a reliability evaluation. AutomationBench-AA's 69% indicates strong performance on this set of 657 tasks across 39 applications, but the AA-Omniscience index remains -5.3 and the Non-Hallucination Rate is 4%. Additional verification, refusal handling, and citation checks are required for fact verification or high-risk decisions.
Version and setting determine comparability. This snapshot corresponds to the auxiliary page's Reasoning, Max Effort setting and AI Index v4.3. If the page updates the index composition or uses another reasoning setting, its results cannot be mixed directly with this record's score of 40, cost of $0.27, or index of -5.3.
The main post contains internal contradictions that must be retained. The opening calls the model 552B, while later model details say “763B total parameters”; the auxiliary page currently says 552B total and 16B active. The Terminal-Bench comparison also conflicts by stating “27% and more than double 27%.” This article preserves the original wording and does not correct it without evidence.
A public summary cannot be treated as a fully reproducible experiment. The sources do not specify the complete task sample, prompts, sampling parameters, number of repetitions, confidence intervals, hardware details, or failed samples. The scores should be treated as an independent evaluation snapshot under the AA methodology.
Record the model identifier DeepSeek V4.1 Flash, reasoning setting Reasoning, Max Effort, AI Index version v4.3, collection date, and page version. Do not mix in results from ordinary or lower-reasoning settings.
Follow AA's published pricing methodology to calculate input, output, and cached-input costs separately, and mark the additional 50% off-peak discount separately. Retain input tokens, output tokens, cache-hit status, and actual cost for each task.
For agent retesting, at minimum fix the AutomationBench-AA task set, SaaS tool permissions, guardrail rules, and completion criteria. A Terminal-Bench retest must also specify the version, environment, tools, network, and sample setup.
A reliability retest should report the AA-Omniscience Index, Accuracy, and Non-Hallucination Rate together; do not report only the Intelligence Index. Save error, hallucination, and refusal samples to avoid deriving production reliability from a single mean.
To verify the main post, consult the English original on X. The page's automatic translation renders ~4x less per token incorrectly as “about 4 times”; the correct meaning is “about 4 times cheaper per token.”
DeepSeek V4.1 Flash