Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialGemini 3.8 Flash

Gemini 3.8 Flash: Google’s Official Benchmarks and Reproduction Boundaries

Original source

Google Blog (The Keyword)

AuthorTulsee Doshi, Raluca Ada Popa / Google DeepMind

Source date2026-09-02

Tabbit curation2026-09-08

Read original

One-sentence takeaway

In the comparison table embedded in its launch announcement, Google compares Gemini 3.8 Flash with Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra. Gemini 3.8 Flash scores higher than 3.7 Flash on every corresponding score entry in the table, but does not outperform the other vendors’ models on every benchmark; these figures are vendor-reported evaluation results published by Google, not independent re-tests.

Use cases

  • Suitable tasks: Use the official table as a version-migration reference for long-horizon coding, agentic terminals, professional knowledge work, legal/financial workflows, document and chart understanding, and bioinformatics tasks.

  • Unsuitable tasks: Use it to establish absolute cross-vendor rankings, promise online success rates, or infer “no tools” from an item whose tool status is “not specified.”

  • Applicable model version: Gemini 3.8 Flash in Google’s launch announcement (released 2026-09-02).

  • Applicable clients, agents, or APIs: The announcement mentions Google Antigravity, Gemini API / Google AI Studio, Android Studio, Stitch, Gemini Enterprise, and the Gemini app; the table does not guarantee specific availability.

Test environment

  • Model columns in the official table: Gemini 3.8 Flash, Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra.

  • Pricing basis: Input pricing is labeled “$/1M tokens, no caching”; the $0.75 input and $3.75 output prices for 3.8/3.7 are introductory prices, with a note below the table stating that they revert to $1.50 / $7.50 after 2026-12-31.

  • Tool annotations: The original image explicitly labels CharXiv Reasoning as No tools and OSWorld-2.0 as Partial score with batch tool enabled; tool status is not specified for the other items in the table, so no inference can be made.

  • Page notes: The body says that 3.8 Flash performs more reasoning steps on complex tasks and iteratively calls tools, while higher effort may use more tokens; the table does not provide the effort, tool harness, or sampling configuration for each item.

Raw data table

The following is a line-by-line transcription of the original image embedded in Google’s page; percentages retain %, GDPval-AA v2 retains Elo, and the original image’s color or bold highlighting is not treated as statistical significance.

Pricing (original image)

ItemGemini 3.8 FlashGemini 3.7 FlashClaude Opus 5Claude Sonnet 5GPT-5.6 SolGPT-5.6 Terra
Input price ($/1M tokens, no caching)$0.75 ($1.50 regular)$0.75 ($1.50 regular)$5.00$2.00$4.00$2.00
Output price ($/1M tokens)$3.75 ($7.50 regular)$3.75 ($7.50 regular)$25.00$10.00$20.00$12.00

Benchmark scores (original image)

BenchmarkTask/metricTool annotationGemini 3.8 FlashGemini 3.7 FlashClaude Opus 5Claude Sonnet 5GPT-5.6 SolGPT-5.6 Terra
DeepSWE v1.1Long-horizon software engineeringNot specified73.7%65.3%74.0%53.8%72.7%69.6%
GDPval-AA v2Knowledge work; EloNot specified154514821824158417101528
Vals Finance Agent v2Financial analyst tasksNot specified61.4%59.0%58.6%53.9%53.8%54.4%
Harvey's Legal Agent BenchmarkComplex legal workflows; All pass rateNot specified10.0%8.8%6.7%5.0%2.5%0.8%
Terminal-bench 2.1Agentic terminal codingNot specified89.4%85.8%89.1%80.4%88.8%87.4%
Terminal-bench 4.0General agent capabilitiesNot specified19.1%11.2%51.8%12.4%37.3%23.6%
GDP.PDFExpert PDF document comprehension; All pass rateNot specified35.0%34.0%37.0%28.0%40.0%29.0%
CharXiv ReasoningInformation synthesis from complex chartsNo tools86.2%84.5%83.7%70.1%85.8%85.9%
LVBenchLong video understandingNot specified; 3.8 has agentic/static tracks87.8% (agentic); 87.1% (static)85.4%75.4%68.5%82.1%78.9%
HLE-VerifiedMultidisciplinary expert reasoningNot specified54.9%53.6%54.4%31.0%54.5%51.1%
OSWorld-2.0Agentic computer use; Partial scorebatch tool enabled59.0%50.6%75.4%42.6%62.6%50.2%
BioMysteryBench (Human Solvable)Bioinformatics research workflowsNot specified88.8%87.1%90.1%87.5%79.5%83.8%
BioMysteryBench (Human Difficult)Bioinformatics research workflowsNot specified56.5%43.5%49.4%34.1%44.7%49.4%
LABBench2Biology real-world research tasksNot specified86.2%82.1%84.2%80.1%82.1%81.2%

Interpreting the results

  • Version difference between 3.8 and 3.7: DeepSWE v1.1 is 73.7% vs. 65.3%, GDPval-AA v2 is 1545 vs. 1482, Vals Finance Agent v2 is 61.4% vs. 59.0%, Terminal-bench 2.1 is 89.4% vs. 85.8%, and OSWorld-2.0 is 59.0% vs. 50.6%; the largest gap is on the BioMysteryBench Human Difficult subtask, at 56.5% vs. 43.5%.

  • No across-the-board lead: Claude Opus 5 scores higher on DeepSWE v1.1 (74.0%), GDPval-AA v2 (1824), Terminal-bench 4.0 (51.8%), GDP.PDF (37.0%), and BioMysteryBench Human Solvable (90.1%); GPT-5.6 Sol scores higher on GDP.PDF (40.0%); and Claude Opus 5 scores higher on OSWorld-2.0 (75.4%).

  • Items where 3.8 is higher in the table: Vals Finance Agent v2, Harvey's Legal Agent Benchmark, Terminal-bench 2.1, CharXiv Reasoning, LVBench (both 3.8 values are higher than the comparison values), HLE-Verified, BioMysteryBench Human Difficult, and LABBench2. Here, “higher” refers only to the point estimates in this table; it does not mean the results can be combined across tasks into a single overall score.

Vendor-reported limits

  • This is a comparison table published by Google in its launch announcement; even where the benchmark names belong to external projects, the page does not disclose each project’s complete prompts, data split/version, sample size, number of repetitions, random seeds, confidence intervals, model snapshots, thinking/effort settings, or scoring scripts.

  • Tool conditions can be verified only in the two places explicitly marked in the original image: CharXiv is No tools, and OSWorld-2.0 is batch tool enabled. The tools, agent scaffold, human intervention, and failure retries for the other rows are unknown, so the scores cannot be attributed to the base models alone.

  • The body explicitly says that 3.8 Flash performs additional reasoning steps and iteratively calls tools on complex tasks, while higher effort may consume more tokens; this means that “quality improvements” and “increased compute/tool budgets” may coexist.

  • Pricing also reflects a launch-time promotion: the table’s $0.75 / $3.75 is not permanent pricing; cost comparisons should record the price effective date, caching status, input/output tokens, and tool-call costs.

  • Google’s positioning of 3.8 Flash as the “most intelligent workhorse model” is a vendor positioning statement; it cannot replace acceptance testing on specific business tasks, nor can the limited set of items in the table be extrapolated to all tasks.

Reproduction notes

  1. Record the exact Gemini 3.8 Flash model string, snapshot/date, client, effort, context, output limit, and pricing version used in practice; do not record only “Gemini Flash.”

  2. Fix the same data version, inputs, scoring rules, and number of repetitions for each table item; when reproducing CharXiv, keep tools disabled, and record separately whether the batch tool is enabled for OSWorld-2.0.

  3. For DeepSWE, terminal, legal/financial agents, document, video, computer-use, and bioinformatics tasks, record success rates, scores, input/output tokens, tool calls, latency, cost, and failure types separately.

  4. Report the difference between 3.8 and 3.7, the differences from the other models, and confidence intervals separately; do not average different benchmarks or tool conditions into a single overall score.

  5. If Google’s undisclosed harness, split, or parameters cannot be obtained, mark the result as “the official conditions were not reproduced,” and treat this article’s table only as a vendor-reported baseline from the time of release.

Source excerpts or observations (compliance short quote only)

Google calls Gemini 3.8 Flash the “most intelligent workhorse model”; this is product positioning, not an independent evaluation conclusion.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Gemini 3.8 Flash

Use and compare models in Tabbit

Gemini 3.8 Flash

Related reviews

MediaArtificial Analysis (official model pages, methodology, and release article; the official X account was used to discover and cross-check the release post)2026-09-02

Gemini 3.8 Flash: Artificial Analysis Intelligence, Speed, Pricing, and Latency

MediaAI IQ (AIIQ, Liberated Software LLC)2026-09-02

Gemini 3.8 Flash: AI IQ Capability Benchmarks and Task Boundaries

MediaVals AI2026-09-05

Vals AI Finance Agent v2: Professional Finance Agent Benchmark for Gemini 3.8 Flash

MediaSimpleBench official leaderboard and project

SimpleBench: Gemini 3.8 Flash on Everyday Reasoning and Language-Trap Questions

Gemini 3.8 Flash

Related prompts

OfficialGoogle AI for Developers / Google DeepMind2026-09-02

Gemini 3.8 Flash: Google’s Official Model Parameters and API Configuration

OfficialGoogle AI for Developers2026-06-10

Gemini 3.8 Flash: Google's Official Structured Prompting and Agent Workflow

OfficialGoogle AI for Developers

Gemini 3.8 Flash: Google's Official Function-Calling Configuration and Tool Workflow

OfficialGoogle AI for Developers2026-09-02

Gemini 3.8 Flash: Google's Official Structured Output Configuration