In the comparison table embedded in its launch announcement, Google compares Gemini 3.8 Flash with Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra. Gemini 3.8 Flash scores higher than 3.7 Flash on every corresponding score entry in the table, but does not outperform the other vendors’ models on every benchmark; these figures are vendor-reported evaluation results published by Google, not independent re-tests.
Suitable tasks: Use the official table as a version-migration reference for long-horizon coding, agentic terminals, professional knowledge work, legal/financial workflows, document and chart understanding, and bioinformatics tasks.
Unsuitable tasks: Use it to establish absolute cross-vendor rankings, promise online success rates, or infer “no tools” from an item whose tool status is “not specified.”
Applicable model version: Gemini 3.8 Flash in Google’s launch announcement (released 2026-09-02).
Applicable clients, agents, or APIs: The announcement mentions Google Antigravity, Gemini API / Google AI Studio, Android Studio, Stitch, Gemini Enterprise, and the Gemini app; the table does not guarantee specific availability.
Model columns in the official table: Gemini 3.8 Flash, Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra.
Pricing basis: Input pricing is labeled “$/1M tokens, no caching”; the $0.75 input and $3.75 output prices for 3.8/3.7 are introductory prices, with a note below the table stating that they revert to $1.50 / $7.50 after 2026-12-31.
Tool annotations: The original image explicitly labels CharXiv Reasoning as No tools and OSWorld-2.0 as Partial score with batch tool enabled; tool status is not specified for the other items in the table, so no inference can be made.
Page notes: The body says that 3.8 Flash performs more reasoning steps on complex tasks and iteratively calls tools, while higher effort may use more tokens; the table does not provide the effort, tool harness, or sampling configuration for each item.
The following is a line-by-line transcription of the original image embedded in Google’s page; percentages retain %, GDPval-AA v2 retains Elo, and the original image’s color or bold highlighting is not treated as statistical significance.
| Item | Gemini 3.8 Flash | Gemini 3.7 Flash | Claude Opus 5 | Claude Sonnet 5 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Input price ($/1M tokens, no caching) | $0.75 ($1.50 regular) | $0.75 ($1.50 regular) | $5.00 | $2.00 | $4.00 | $2.00 |
| Output price ($/1M tokens) | $3.75 ($7.50 regular) | $3.75 ($7.50 regular) | $25.00 | $10.00 | $20.00 | $12.00 |
| Benchmark | Task/metric | Tool annotation | Gemini 3.8 Flash | Gemini 3.7 Flash | Claude Opus 5 | Claude Sonnet 5 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|---|---|
| DeepSWE v1.1 | Long-horizon software engineering | Not specified | 73.7% | 65.3% | 74.0% | 53.8% | 72.7% | 69.6% |
| GDPval-AA v2 | Knowledge work; Elo | Not specified | 1545 | 1482 | 1824 | 1584 | 1710 | 1528 |
| Vals Finance Agent v2 | Financial analyst tasks | Not specified | 61.4% | 59.0% | 58.6% | 53.9% | 53.8% | 54.4% |
| Harvey's Legal Agent Benchmark | Complex legal workflows; All pass rate | Not specified | 10.0% | 8.8% | 6.7% | 5.0% | 2.5% | 0.8% |
| Terminal-bench 2.1 | Agentic terminal coding | Not specified | 89.4% | 85.8% | 89.1% | 80.4% | 88.8% | 87.4% |
| Terminal-bench 4.0 | General agent capabilities | Not specified | 19.1% | 11.2% | 51.8% | 12.4% | 37.3% | 23.6% |
| GDP.PDF | Expert PDF document comprehension; All pass rate | Not specified | 35.0% | 34.0% | 37.0% | 28.0% | 40.0% | 29.0% |
| CharXiv Reasoning | Information synthesis from complex charts | No tools | 86.2% | 84.5% | 83.7% | 70.1% | 85.8% | 85.9% |
| LVBench | Long video understanding | Not specified; 3.8 has agentic/static tracks | 87.8% (agentic); 87.1% (static) | 85.4% | 75.4% | 68.5% | 82.1% | 78.9% |
| HLE-Verified | Multidisciplinary expert reasoning | Not specified | 54.9% | 53.6% | 54.4% | 31.0% | 54.5% | 51.1% |
| OSWorld-2.0 | Agentic computer use; Partial score | batch tool enabled | 59.0% | 50.6% | 75.4% | 42.6% | 62.6% | 50.2% |
| BioMysteryBench (Human Solvable) | Bioinformatics research workflows | Not specified | 88.8% | 87.1% | 90.1% | 87.5% | 79.5% | 83.8% |
| BioMysteryBench (Human Difficult) | Bioinformatics research workflows | Not specified | 56.5% | 43.5% | 49.4% | 34.1% | 44.7% | 49.4% |
| LABBench2 | Biology real-world research tasks | Not specified | 86.2% | 82.1% | 84.2% | 80.1% | 82.1% | 81.2% |
Version difference between 3.8 and 3.7: DeepSWE v1.1 is 73.7% vs. 65.3%, GDPval-AA v2 is 1545 vs. 1482, Vals Finance Agent v2 is 61.4% vs. 59.0%, Terminal-bench 2.1 is 89.4% vs. 85.8%, and OSWorld-2.0 is 59.0% vs. 50.6%; the largest gap is on the BioMysteryBench Human Difficult subtask, at 56.5% vs. 43.5%.
No across-the-board lead: Claude Opus 5 scores higher on DeepSWE v1.1 (74.0%), GDPval-AA v2 (1824), Terminal-bench 4.0 (51.8%), GDP.PDF (37.0%), and BioMysteryBench Human Solvable (90.1%); GPT-5.6 Sol scores higher on GDP.PDF (40.0%); and Claude Opus 5 scores higher on OSWorld-2.0 (75.4%).
Items where 3.8 is higher in the table: Vals Finance Agent v2, Harvey's Legal Agent Benchmark, Terminal-bench 2.1, CharXiv Reasoning, LVBench (both 3.8 values are higher than the comparison values), HLE-Verified, BioMysteryBench Human Difficult, and LABBench2. Here, “higher” refers only to the point estimates in this table; it does not mean the results can be combined across tasks into a single overall score.
This is a comparison table published by Google in its launch announcement; even where the benchmark names belong to external projects, the page does not disclose each project’s complete prompts, data split/version, sample size, number of repetitions, random seeds, confidence intervals, model snapshots, thinking/effort settings, or scoring scripts.
Tool conditions can be verified only in the two places explicitly marked in the original image: CharXiv is No tools, and OSWorld-2.0 is batch tool enabled. The tools, agent scaffold, human intervention, and failure retries for the other rows are unknown, so the scores cannot be attributed to the base models alone.
The body explicitly says that 3.8 Flash performs additional reasoning steps and iteratively calls tools on complex tasks, while higher effort may consume more tokens; this means that “quality improvements” and “increased compute/tool budgets” may coexist.
Pricing also reflects a launch-time promotion: the table’s $0.75 / $3.75 is not permanent pricing; cost comparisons should record the price effective date, caching status, input/output tokens, and tool-call costs.
Google’s positioning of 3.8 Flash as the “most intelligent workhorse model” is a vendor positioning statement; it cannot replace acceptance testing on specific business tasks, nor can the limited set of items in the table be extrapolated to all tasks.
Record the exact Gemini 3.8 Flash model string, snapshot/date, client, effort, context, output limit, and pricing version used in practice; do not record only “Gemini Flash.”
Fix the same data version, inputs, scoring rules, and number of repetitions for each table item; when reproducing CharXiv, keep tools disabled, and record separately whether the batch tool is enabled for OSWorld-2.0.
For DeepSWE, terminal, legal/financial agents, document, video, computer-use, and bioinformatics tasks, record success rates, scores, input/output tokens, tool calls, latency, cost, and failure types separately.
Report the difference between 3.8 and 3.7, the differences from the other models, and confidence intervals separately; do not average different benchmarks or tool conditions into a single overall score.
If Google’s undisclosed harness, split, or parameters cannot be obtained, mark the result as “the official conditions were not reproduced,” and treat this article’s table only as a vendor-reported baseline from the time of release.
Google calls Gemini 3.8 Flash the “most intelligent workhorse model”; this is product positioning, not an independent evaluation conclusion.
Gemini 3.8 Flash