Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaGemini 3.8 Flash

Gemini 3.8 Flash: AI IQ Capability Benchmarks and Task Boundaries

Original source

AI IQ (AIIQ, Liberated Software LLC)

AuthorAI IQ

Source date2026-09-02

Tabbit curation2026-09-08

Read original

One-sentence takeaway

AIIQ's current model page gives Gemini 3.8 Flash an AI IQ of 125, ranked #14/131, and lists scores from 8 source benchmarks. Academic reasoning has 5/6 coverage, programmatic reasoning only 1/6, and reliability 2/7, while abstract reasoning, mathematics, and computer use are respectively 0/3, 0/5, and 0/6. The page therefore supports the conclusion that the model performs strongly on a limited number of public benchmarks, but does not support treating the composite IQ or estimates for missing dimensions as a complete capability evaluation.

Use cases

  • Tasks suitable for observation: Long-document extraction/synthesis/reasoning, expert academic question answering, multimodal chart and image understanding, scientific research code, and terminal development and system tasks in isolated containers.

  • Tasks not suitable for direct inference: Browser or desktop automation, complete software-engineering repository repair, general-purpose visual capabilities, low-latency experience, or production-grade reliability; the model page shows Computer Use at 0/6, with no source benchmark coverage for that dimension.

  • Applicable model version: The exact model named on the AIIQ page, Gemini 3.8 Flash; do not extend the result to Gemini 3.7 Flash, other unlabeled configurations of Gemini 3.8 Flash, or vendor routing.

  • Applicable client, Agent, or API: AIIQ's model and benchmark pages; the pages do not disclose Gemini API's specific system prompt, sampling parameters, or complete Agent harness.

  • Recommended parameters: AIIQ does not provide a reproducible temperature, reasoning tier, stop condition, or retry configuration; any retest should fix and report these variables independently.

Test environment

  • AIIQ organizes 6 dimensions with equal weighting: Abstract, Math, Academic, Programmatic, Computer Use, and Reliability. The composite IQ on the model page is the displayed value across these 6 dimensions.

  • AIIQ's methodology says the Artificial Analysis Intelligence Index is the primary aggregation source and provides scores for HLE, GPQA Diamond, SciCode, Terminal-Bench 2.1, CritPt, MMMU-Pro, AA Omniscience, and AA Long Context Reasoning; the Terminal-Bench page also lists Vals.ai as a source/fallback source.

  • The cost-efficiency charts on the benchmark pages expand configurations of the same model across multiple reasoning tiers, but the score bars retain one canonical row per model. The model scores on the pages therefore cannot be interpreted as a single public run using one unified harness.

  • The AIIQ page also lists a context window of 1,048,576 tokens, input/output pricing of $0.75/$3.75 per million tokens, approximately 305 tokens/s, and a typical response time of 15 seconds; these are page-supplied or derived metrics, not measurements reproduced in this note.

Input/configuration

  • Model information: AIIQ labels Gemini 3.8 Flash as Google and Proprietary; the model-page FAQ says it was publicly released on 2026-09-02.

  • Composite calculation: Each raw score s is first converted by piecewise-linear interpolation f_b(s) against that benchmark's expected IQ-score steps of 70, 85, 100, 115, 130, 145, and 160. Scores are then averaged across benchmarks within each dimension, followed by an equally weighted average across the 6 dimensions:

    $$\mathrm{IQ}=\frac{1}{6}\sum_{d=1}^{6}\mathrm{IQ}_d$$

  • Missing-value rule: The methodology explicitly says that missing benchmarks receive conservative imputation within the scoring pipeline. Thus, coverage such as 0/6 means “no source benchmark coverage,” not a score of zero and not a directly measured result.

  • Boundary of the page's raw materials: No model-specific original questions, per-question outputs, random seeds, complete call logs, sampling configuration, or independent rerun reports were found. The verifiable scope is the scores, source labels, task semantics, and calculation rules listed on the pages.

Results data

Model-page summary

ItemAIIQ page valueInterpretation boundary
Composite AI IQ125Derived composite value across 6 dimensions, not a raw benchmark score
IQ ranking#14 / 131Dynamic leaderboard position at the time of collection
Abstract Reasoning103 (0/3)No ARC source-benchmark coverage listed; the dimension value includes missing-value handling
Mathematical Reasoning126 (0/5)No mathematical source-benchmark coverage listed; cannot be treated as a direct math evaluation
Academic Reasoning141 (5/6)The most thoroughly covered dimension among the 8 reported results
Programmatic Reasoning132 (1/6)Supported primarily by the single Terminal-Bench 2.1 result
Computer Use124 (0/6)No source-benchmark coverage for this dimension; the page value is a derived estimate
Reliability122 (2/7)Covered by reliability signals including AA Omniscience and AA Long Context Reasoning

Source benchmark readings for Gemini 3.8 Flash

BenchmarkScoreTask boundary given on the AIIQ pageDimension
AA Long Context Reasoning81Extraction, linking, and reasoning in 10K–100K-token long documents; the page lists academic papers, company financial reports, government consultations, legal documents, industry reports, marketing materials, and surveysReliability
AA Omniscience29.55Artificial Analysis's factual reliability index; correct answers add points, hallucinations subtract points, and refusals incur no penalty; 0 means correct and incorrect answers offset each otherReliability
CritPt18.286Original critical-point analysis research-physics problems, scored by original source points; AIIQ labels it as having low gameabilityAcademic
GPQA Diamond95.253198 public four-choice, PhD-level science questions; the random baseline is 25%, and the public question set creates saturation and memorization risks at the high endAcademic
Humanity's Last Exam47.8223,000 expert-contributed questions, screened at creation to exclude questions that existing models could answer; the page marks it as 76% exact-match with low gameabilityAcademic
MMMU-Pro85.607Multimodal academic reasoning covering text, diagrams, charts, and imagesAcademic
SciCode53.588Implementing scientific research problems in code; requires identifying scientific concepts, recalling domain facts, deriving numerical methods/simulations, and turning them into computationAcademic
Terminal-Bench 2.187.6489 isolated Docker-container shell tasks covering practical system administration and development; the page reports pass@1Programmatic

Most benchmark-page chart notes are marked Data updated Sep 7, 2026; the Terminal-Bench 2.1 page is marked Sep 6, 2026. These dates differ from the model-page collection date, so dynamic refreshes should not be treated as one experimental batch.

Conclusions

  1. AIIQ's 8 results provide quantifiable readings for Gemini 3.8 Flash on source rows including HLE, GPQA Diamond, MMMU-Pro, SciCode, and Terminal-Bench 2.1; Terminal-Bench's 87.64% covers only 89 Docker/shell tasks and cannot be generalized to all coding work.

  2. Academic IQ 141 has 5/6 coverage and is a relatively data-supported dimension; composite IQ 125 still mixes in conservative imputation for many missing dimensions and cannot be equated with “all six capabilities were tested.”

  3. The reliability signals have clear boundaries: AA Omniscience does not penalize refusals, while AA Long Context Reasoning covers only 10K–100K-token documents. They cannot replace fact checking, citation accuracy, or longer-context regression testing in production scenarios.

  4. AIIQ is an aggregation display: the main scores come from Artificial Analysis, while Terminal-Bench also uses Vals.ai as a fallback/cross-source; the pages do not provide model-level original questions, a complete harness, or independent run logs. It is suitable as a source index and a quick overview of task boundaries, not as an AIIQ self-test or a unified competition score.

Limitations

  • AIIQ's Composite IQ, per-dimension IQs, and ranking are derived metrics; the calibration steps between benchmark raw scores and IQ, missing-value imputation, and equal weighting can materially affect the final numbers.

  • Coverage values such as 0/3, 0/5, and 0/6 are not failing scores, but the corresponding dimension values may still be displayed on the page. When citing them, coverage must be reported as well; do not quote the IQ alone.

  • Benchmarks differ in question type, scale, scoring, and gameability: GPQA Diamond uses public four-choice questions, HLE uses expert questions, MMMU-Pro uses multimodal questions, and Terminal-Bench uses executable shell tasks. Their scores cannot be averaged horizontally into one “accuracy” figure.

  • The page presents third-party published/aggregated source-backed data and does not disclose independent reproduction configurations, raw outputs, or confidence intervals for Gemini 3.8 Flash; AIIQ's source labels do not mean that the source platforms or Google endorse all derived values.

  • Prices, response times, throughput, and leaderboard positions change with vendors, caching, concurrency, and data refreshes; the page's $35.75 effective cost is a derived cost and should not be conflated with the listed $0.75/$3.75 prices.

Reproduction steps

  1. Fix the exact model version, vendor/API route, reasoning tier, temperature, context limit, stop conditions, retry strategy, and tool schema, and record HLE, GPQA, MMMU-Pro, SciCode, and Terminal-Bench separately.

  2. For HLE, GPQA, MMMU-Pro, and SciCode, save the dataset version, modality inputs, per-question outputs, and scoring scripts; for Terminal-Bench, save the 89 container tasks, command traces, final states, and pass@1 determinations.

  3. Test long-context tasks separately at 10K, 50K, 100K, and the target business length, recording extraction recall, citation accuracy, hallucination rate, refusal rate, and end-to-end latency; do not treat AA Long Context Reasoning's 81 as a guarantee for a 1M context window.

  4. Recalculate the raw scores using the same-version Artificial Analysis/source-platform readings, then apply the piecewise-linear steps in AIIQ's methodology to calculate dimension and composite IQs; report source-backed coverage and imputed items separately.

  5. Build additional task sets for Computer Use, browsing, desktop use, and repository-level software engineering; this AIIQ page has no source coverage for these dimensions, so its Computer Use IQ of 124 cannot replace direct measurement.

Original evidence and data

  • The AIIQ Gemini 3.8 Flash model page directly lists IQ 125, #14, 1M context, $0.75/$3.75 pricing, 8 benchmark scores, and coverage for the 6 dimensions.

  • The AIIQ benchmark pages provide task semantics and chart notes respectively: long-document extraction/synthesis/reasoning for AA Long Context Reasoning, a factual reliability index for AA Omniscience, multimodal reasoning over text/charts/images for MMMU-Pro, and 89 Docker shell tasks for Terminal-Bench 2.1.

  • The AIIQ methodology page provides the equal-weight formula across 6 dimensions, piecewise-linear interpolation on benchmark IQ steps, and the conservative missing-value imputation rule, and identifies the Artificial Analysis Intelligence Index as the primary aggregation source.

  • The source labels on AIIQ benchmark pages show Artificial Analysis; Terminal-Bench 2.1 also shows Vals.ai Terminal-Bench 2.1, so this note treats it as an aggregation/fallback source rather than an independent AIIQ rerun.

Source excerpt or observation (short quote for compliance only)

AIIQ describes the model page as “source-backed benchmark results and derived ranking context.” This accurately defines the evidence boundary: the page contains source scores as well as ranking context produced by interpolation and imputation, and the two should not be conflated with a complete model evaluation run.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Gemini 3.8 Flash

Use and compare models in Tabbit

Gemini 3.8 Flash

Related reviews

OfficialGoogle Blog (The Keyword)2026-09-02

Gemini 3.8 Flash: Google’s Official Benchmarks and Reproduction Boundaries

MediaArtificial Analysis (official model pages, methodology, and release article; the official X account was used to discover and cross-check the release post)2026-09-02

Gemini 3.8 Flash: Artificial Analysis Intelligence, Speed, Pricing, and Latency

MediaVals AI2026-09-05

Vals AI Finance Agent v2: Professional Finance Agent Benchmark for Gemini 3.8 Flash

MediaSimpleBench official leaderboard and project

SimpleBench: Gemini 3.8 Flash on Everyday Reasoning and Language-Trap Questions

Gemini 3.8 Flash

Related prompts

OfficialGoogle AI for Developers / Google DeepMind2026-09-02

Gemini 3.8 Flash: Google’s Official Model Parameters and API Configuration

OfficialGoogle AI for Developers2026-06-10

Gemini 3.8 Flash: Google's Official Structured Prompting and Agent Workflow

OfficialGoogle AI for Developers

Gemini 3.8 Flash: Google's Official Function-Calling Configuration and Tool Workflow

OfficialGoogle AI for Developers2026-09-02

Gemini 3.8 Flash: Google's Official Structured Output Configuration