At the time of collection, the AI IQ page estimated DeepSeek V4.1 Flash at IQ 116, ranked #38/138. This IQ is a derived estimate across six equally weighted capability dimensions; missing benchmarks enter a conservative imputation process, so it should not be treated as a direct intelligence quotient or as a single complete independent test. The visible benchmark coverage is concentrated in Academic Reasoning (4/6) and Reliability (1/7), while the other four dimensions are 0/3, 0/5, 0/6, and 0/6.
Tasks it is suitable for evaluating: Viewing the model's overall ranking, capability-dimension scores, and included benchmark results under the same AI IQ methodology; using coverage to judge how much evidence supports each score.
Tasks it is not suitable for extrapolating to: Treating 116 as a human IQ, treating #38/138 as a stable rank, or using it to assert the model's performance on untested tasks, arbitrary clients, and production workflows; Effective Cost also cannot be treated as a single API token price.
Applicable model version: The DeepSeek V4.1 Flash identified on the page; the API, open-weight deployments, and other inference configurations require separate verification.
Test environment or client: The public AI IQ model profile and the public benchmark data it includes; the specific runtime harness, sample count, and parameters are not itemized on the model page.
Reasoning setting and parameters: Not specified.
The supporting source for the scoring and cost definitions is AI IQ Methodology. The model score itself follows the snapshot of the primary source collected on the date above.
The model page combines six equally weighted dimensions into Composite IQ: Abstract, Mathematical, Academic, Programmatic, Computer Use, and Reliability. The methodology linked from the page says that each benchmark's raw score is first converted through piecewise linear interpolation along calibration points at IQ 70, 85, 100, 115, 130, 145, and 160, and the benchmarks within each dimension are then averaged; missing benchmarks and missing dimensions are conservatively imputed only within the derived-scoring process. A composite score is produced only when at least 2/6 dimensions have direct or effectively inherited measurement evidence.
Visible coverage is the number of benchmarks covered for a dimension / the total number of benchmarks listed on the page. It measures evidence coverage; it is not the percentage of that dimension's IQ and does not mean that every item was independently run by AI IQ. The AI IQ methodology also explicitly warns that all dimension values are estimates; missing-value imputation, the calibration points, and the benchmark mix all affect the result.
The composite score can be summarized as:
$$ \mathrm{IQ}=\frac{\mathrm{Abstract}+\mathrm{Math}+\mathrm{Academic}+\mathrm{Programmatic}+\mathrm{Computer}+\mathrm{Reliability}}{6} $$
The page also provides “Effective Cost.” The same site's methodology defines it as:
$$ \mathrm{Effective\ Cost}=\mathrm{Sticker\ Price}\times\mathrm{Usage\ Multiplier} $$
Here, Sticker Price uses 1M input + 1M output tokens as the standard workload, while Usage Multiplier reflects the actual token or task cost used by the model to complete similar tasks. It is therefore a comparison metric adjusted for task efficiency, not the API's input or output unit price.
At the time of collection, the page showed IQ 116, IQ Rank #38, and Effective Cost $3.6328. The FAQ describes the rank as 38th among 138 public models and says that the score combines six equally weighted dimensions with conservative estimates for missing coverage; this is a page snapshot from 2026-09-16 and cannot be treated as a permanent ranking.
Scores and coverage across the six dimensions:
| Dimension | IQ | Coverage | How to read it |
|---|---|---|---|
| Abstract Reasoning | 99 | 0/3 | None of the 3 benchmarks listed on the page has direct coverage; the score includes estimated components |
| Mathematical Reasoning | 116 | 0/5 | None of the 5 benchmarks listed on the page has direct coverage; the score includes estimated components |
| Academic Reasoning | 136 | 4/6 | 4 of the 6 benchmarks are covered |
| Programmatic Reasoning | 122 | 0/6 | None of the 6 benchmarks listed on the page has direct coverage; the score includes estimated components |
| Computer Use | 120 | 0/6 | None of the 6 benchmarks listed on the page has direct coverage; the score includes estimated components |
| Reliability | 103 | 1/7 | 1 of the 7 benchmarks is covered |
The model page publicly lists the following benchmark results. The profile does not specify runtime parameters, sample size, number of repetitions, or a complete harness, so these are recorded as values included by the AI IQ page and are not rewritten as results from a standardized independent experiment.
| Benchmark | Score | Dimension |
|---|---|---|
| AA Omniscience | -5.3 | Reliability |
| CritPt | 14.2857 | Academic Reasoning |
| Humanity's Last Exam | 39.2493 | Academic Reasoning |
| MMMU-Pro | 76.9942 | Academic Reasoning |
| SciCode | 51.8519 | Academic Reasoning |
The page's FAQ also gives a pricing and runtime overview: $0.30 / 1M tokens for input and $1.20 / 1M tokens for output; a context window of 1,000,000 tokens; independent measurement of about 190 tokens/s; and a typical response time of about 14 seconds. The model page does not detail the measurement setup for the last two items, so they cannot be used to reproduce performance.
The FAQ gives the model's publication date as 2026-09-10; this is the model release date, not the publication date of the leaderboard page.
Under the same site's 1M input + 1M output convention, the public unit prices correspond to:
$$ \mathrm{Sticker\ Price}=0.30+1.20=$1.50 $$
The page's $3.6328 is the Effective Cost after adjusting for task usage. Back-calculating from the two numbers shown on the page alone gives an adjustment multiplier of approximately $3.6328 / $1.50=2.42$, but the page does not disclose this model's specific multiplier, paired-task records, or measurement details. This back-calculated value therefore cannot be treated as an independent cost measurement.
116 is a derived AI IQ metric. It uses human IQ as an intuitive yardstick and is not equivalent to human intelligence in the psychometric sense. Equal weighting across six dimensions also does not mean that the model has the same amount of direct evidence for each task category.
Coverage is central to reading this page. Academic Reasoning's 4/6 is the page's most substantial direct coverage, while Reliability has only 1/7; Abstract, Math, Programmatic, and Computer Use all have 0 coverage. A value such as 0/3 means that the page lacks direct benchmark coverage; it does not mean the model scored 0.
The composite score includes missing-value handling. The methodology says that missing items are conservatively imputed in the scoring copy and distinguishes source-backed data from derived scoring; therefore, the page can show scores without all six dimensions having been fully measured directly, and the results should not be presented as such.
#38/138 is a dynamic date-specific snapshot. This record only states the rank displayed on the page when it was collected on 2026-09-16; new models, data updates, calibration changes, or changes in leaderboard scope may alter it.
Effective Cost is not the API unit price. The API list prices are $0.30/M for input and $1.20/M for output; $3.6328 applies the task-usage adjustment to a standard I/O workload and is suitable for comparing task efficiency, but it cannot directly replace a supplier billing estimate.
This is not a complete independent retest. The page provides included scores and a derived ranking, but does not list complete inputs, outputs, samples, repeated experiments, confidence intervals, or a harness for this page; it should be cited separately from an official model card or a reproducible experiment.
Record the access date, page URL, model name, IQ, rank, Effective Cost, six dimension scores, and coverage; dynamic leaderboards must retain the snapshot date.
To reproduce an individual benchmark, separately obtain the benchmark's original source, version, inputs, scoring metric, model configuration, harness, sample size, and repetition strategy; the AI IQ page itself is insufficient to reconstruct these conditions.
To recalculate the composite score, obtain AI IQ's current benchmark calibration points and missing-value waterfall, then follow the methodology: convert each benchmark to IQ, handle missing values within each dimension, and finally calculate the equally weighted average across the six dimensions. The five public results in the table cannot be averaged directly to produce 116.
Cost comparisons should report the supplier's API unit prices and Effective Cost together, clearly stating the standard workload of 1M input + 1M output and the source of the Usage Multiplier.
DeepSeek V4.1 Flash