AIIQ's current model page gives Gemini 3.8 Flash an AI IQ of 125, ranked #14/131, and lists scores from 8 source benchmarks. Academic reasoning has 5/6 coverage, programmatic reasoning only 1/6, and reliability 2/7, while abstract reasoning, mathematics, and computer use are respectively 0/3, 0/5, and 0/6. The page therefore supports the conclusion that the model performs strongly on a limited number of public benchmarks, but does not support treating the composite IQ or estimates for missing dimensions as a complete capability evaluation.
Tasks suitable for observation: Long-document extraction/synthesis/reasoning, expert academic question answering, multimodal chart and image understanding, scientific research code, and terminal development and system tasks in isolated containers.
Tasks not suitable for direct inference: Browser or desktop automation, complete software-engineering repository repair, general-purpose visual capabilities, low-latency experience, or production-grade reliability; the model page shows Computer Use at 0/6, with no source benchmark coverage for that dimension.
Applicable model version: The exact model named on the AIIQ page, Gemini 3.8 Flash; do not extend the result to Gemini 3.7 Flash, other unlabeled configurations of Gemini 3.8 Flash, or vendor routing.
Applicable client, Agent, or API: AIIQ's model and benchmark pages; the pages do not disclose Gemini API's specific system prompt, sampling parameters, or complete Agent harness.
Recommended parameters: AIIQ does not provide a reproducible temperature, reasoning tier, stop condition, or retry configuration; any retest should fix and report these variables independently.
AIIQ organizes 6 dimensions with equal weighting: Abstract, Math, Academic, Programmatic, Computer Use, and Reliability. The composite IQ on the model page is the displayed value across these 6 dimensions.
AIIQ's methodology says the Artificial Analysis Intelligence Index is the primary aggregation source and provides scores for HLE, GPQA Diamond, SciCode, Terminal-Bench 2.1, CritPt, MMMU-Pro, AA Omniscience, and AA Long Context Reasoning; the Terminal-Bench page also lists Vals.ai as a source/fallback source.
The cost-efficiency charts on the benchmark pages expand configurations of the same model across multiple reasoning tiers, but the score bars retain one canonical row per model. The model scores on the pages therefore cannot be interpreted as a single public run using one unified harness.
The AIIQ page also lists a context window of 1,048,576 tokens, input/output pricing of $0.75/$3.75 per million tokens, approximately 305 tokens/s, and a typical response time of 15 seconds; these are page-supplied or derived metrics, not measurements reproduced in this note.
Model information: AIIQ labels Gemini 3.8 Flash as Google and Proprietary; the model-page FAQ says it was publicly released on 2026-09-02.
Composite calculation: Each raw score s is first converted by piecewise-linear interpolation f_b(s) against that benchmark's expected IQ-score steps of 70, 85, 100, 115, 130, 145, and 160. Scores are then averaged across benchmarks within each dimension, followed by an equally weighted average across the 6 dimensions:
$$\mathrm{IQ}=\frac{1}{6}\sum_{d=1}^{6}\mathrm{IQ}_d$$
Missing-value rule: The methodology explicitly says that missing benchmarks receive conservative imputation within the scoring pipeline. Thus, coverage such as 0/6 means “no source benchmark coverage,” not a score of zero and not a directly measured result.
Boundary of the page's raw materials: No model-specific original questions, per-question outputs, random seeds, complete call logs, sampling configuration, or independent rerun reports were found. The verifiable scope is the scores, source labels, task semantics, and calculation rules listed on the pages.
| Item | AIIQ page value | Interpretation boundary |
|---|---|---|
| Composite AI IQ | 125 | Derived composite value across 6 dimensions, not a raw benchmark score |
| IQ ranking | #14 / 131 | Dynamic leaderboard position at the time of collection |
| Abstract Reasoning | 103 (0/3) | No ARC source-benchmark coverage listed; the dimension value includes missing-value handling |
| Mathematical Reasoning | 126 (0/5) | No mathematical source-benchmark coverage listed; cannot be treated as a direct math evaluation |
| Academic Reasoning | 141 (5/6) | The most thoroughly covered dimension among the 8 reported results |
| Programmatic Reasoning | 132 (1/6) | Supported primarily by the single Terminal-Bench 2.1 result |
| Computer Use | 124 (0/6) | No source-benchmark coverage for this dimension; the page value is a derived estimate |
| Reliability | 122 (2/7) | Covered by reliability signals including AA Omniscience and AA Long Context Reasoning |
| Benchmark | Score | Task boundary given on the AIIQ page | Dimension |
|---|---|---|---|
| AA Long Context Reasoning | 81 | Extraction, linking, and reasoning in 10K–100K-token long documents; the page lists academic papers, company financial reports, government consultations, legal documents, industry reports, marketing materials, and surveys | Reliability |
| AA Omniscience | 29.55 | Artificial Analysis's factual reliability index; correct answers add points, hallucinations subtract points, and refusals incur no penalty; 0 means correct and incorrect answers offset each other | Reliability |
| CritPt | 18.286 | Original critical-point analysis research-physics problems, scored by original source points; AIIQ labels it as having low gameability | Academic |
| GPQA Diamond | 95.253 | 198 public four-choice, PhD-level science questions; the random baseline is 25%, and the public question set creates saturation and memorization risks at the high end | Academic |
| Humanity's Last Exam | 47.822 | 3,000 expert-contributed questions, screened at creation to exclude questions that existing models could answer; the page marks it as 76% exact-match with low gameability | Academic |
| MMMU-Pro | 85.607 | Multimodal academic reasoning covering text, diagrams, charts, and images | Academic |
| SciCode | 53.588 | Implementing scientific research problems in code; requires identifying scientific concepts, recalling domain facts, deriving numerical methods/simulations, and turning them into computation | Academic |
| Terminal-Bench 2.1 | 87.64 | 89 isolated Docker-container shell tasks covering practical system administration and development; the page reports pass@1 | Programmatic |
Most benchmark-page chart notes are marked Data updated Sep 7, 2026; the Terminal-Bench 2.1 page is marked Sep 6, 2026. These dates differ from the model-page collection date, so dynamic refreshes should not be treated as one experimental batch.
AIIQ's 8 results provide quantifiable readings for Gemini 3.8 Flash on source rows including HLE, GPQA Diamond, MMMU-Pro, SciCode, and Terminal-Bench 2.1; Terminal-Bench's 87.64% covers only 89 Docker/shell tasks and cannot be generalized to all coding work.
Academic IQ 141 has 5/6 coverage and is a relatively data-supported dimension; composite IQ 125 still mixes in conservative imputation for many missing dimensions and cannot be equated with “all six capabilities were tested.”
The reliability signals have clear boundaries: AA Omniscience does not penalize refusals, while AA Long Context Reasoning covers only 10K–100K-token documents. They cannot replace fact checking, citation accuracy, or longer-context regression testing in production scenarios.
AIIQ is an aggregation display: the main scores come from Artificial Analysis, while Terminal-Bench also uses Vals.ai as a fallback/cross-source; the pages do not provide model-level original questions, a complete harness, or independent run logs. It is suitable as a source index and a quick overview of task boundaries, not as an AIIQ self-test or a unified competition score.
AIIQ's Composite IQ, per-dimension IQs, and ranking are derived metrics; the calibration steps between benchmark raw scores and IQ, missing-value imputation, and equal weighting can materially affect the final numbers.
Coverage values such as 0/3, 0/5, and 0/6 are not failing scores, but the corresponding dimension values may still be displayed on the page. When citing them, coverage must be reported as well; do not quote the IQ alone.
Benchmarks differ in question type, scale, scoring, and gameability: GPQA Diamond uses public four-choice questions, HLE uses expert questions, MMMU-Pro uses multimodal questions, and Terminal-Bench uses executable shell tasks. Their scores cannot be averaged horizontally into one “accuracy” figure.
The page presents third-party published/aggregated source-backed data and does not disclose independent reproduction configurations, raw outputs, or confidence intervals for Gemini 3.8 Flash; AIIQ's source labels do not mean that the source platforms or Google endorse all derived values.
Prices, response times, throughput, and leaderboard positions change with vendors, caching, concurrency, and data refreshes; the page's $35.75 effective cost is a derived cost and should not be conflated with the listed $0.75/$3.75 prices.
Fix the exact model version, vendor/API route, reasoning tier, temperature, context limit, stop conditions, retry strategy, and tool schema, and record HLE, GPQA, MMMU-Pro, SciCode, and Terminal-Bench separately.
For HLE, GPQA, MMMU-Pro, and SciCode, save the dataset version, modality inputs, per-question outputs, and scoring scripts; for Terminal-Bench, save the 89 container tasks, command traces, final states, and pass@1 determinations.
Test long-context tasks separately at 10K, 50K, 100K, and the target business length, recording extraction recall, citation accuracy, hallucination rate, refusal rate, and end-to-end latency; do not treat AA Long Context Reasoning's 81 as a guarantee for a 1M context window.
Recalculate the raw scores using the same-version Artificial Analysis/source-platform readings, then apply the piecewise-linear steps in AIIQ's methodology to calculate dimension and composite IQs; report source-backed coverage and imputed items separately.
Build additional task sets for Computer Use, browsing, desktop use, and repository-level software engineering; this AIIQ page has no source coverage for these dimensions, so its Computer Use IQ of 124 cannot replace direct measurement.
The AIIQ Gemini 3.8 Flash model page directly lists IQ 125, #14, 1M context, $0.75/$3.75 pricing, 8 benchmark scores, and coverage for the 6 dimensions.
The AIIQ benchmark pages provide task semantics and chart notes respectively: long-document extraction/synthesis/reasoning for AA Long Context Reasoning, a factual reliability index for AA Omniscience, multimodal reasoning over text/charts/images for MMMU-Pro, and 89 Docker shell tasks for Terminal-Bench 2.1.
The AIIQ methodology page provides the equal-weight formula across 6 dimensions, piecewise-linear interpolation on benchmark IQ steps, and the conservative missing-value imputation rule, and identifies the Artificial Analysis Intelligence Index as the primary aggregation source.
The source labels on AIIQ benchmark pages show Artificial Analysis; Terminal-Bench 2.1 also shows Vals.ai Terminal-Bench 2.1, so this note treats it as an aggregation/fallback source rather than an independent AIIQ rerun.
AIIQ describes the model page as “source-backed benchmark results and derived ranking context.” This accurately defines the evidence boundary: the page contains source scores as well as ranking context produced by interpolation and imputation, and the two should not be conflated with a complete model evaluation run.
Gemini 3.8 Flash