Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaGemini 3.8 Flash

Gemini 3.8 Flash: Artificial Analysis Intelligence, Speed, Pricing, and Latency

Original source

Artificial Analysis (official model pages, methodology, and release article; the official X account was used to discover and cross-check the release post)

AuthorArtificial Analysis

Source date2026-09-02

Tabbit curation2026-09-08

Read original

One-sentence takeaway

Gemini 3.8 Flash's speed and latency vary substantially by reasoning tier: Artificial Analysis's current v4.3 pages record 285.9 tok/s and 17.36 seconds to the first answer token for high, 246.3 tok/s and 8.29 seconds for medium, and 265.3 tok/s and 0.78 seconds for low; however, the Intelligence Index also changed from the v4.2 figures of 59/57/52 at release on 2026-09-02 to 41/40/34 in the current v4.3 snapshot. These figures come from different test versions and cannot be directly compared as evidence that the model has regressed.

Use cases

  • Suitable tasks: Coding, Agent, document, and enterprise workflows that require long context, multimodal input, tool calling, and relatively fast answer generation.

  • Unsuitable tasks: Using the high tier as the default for low-latency chat, or treating one-off tok/s as total task duration; high has materially higher TTFT and reasoning-token overhead.

  • Applicable model version: gemini-3.8-flash; this article records Artificial Analysis's high, medium, and low reasoning pages without treating the three tiers as three API models.

  • Applicable client, Agent, or API: The Google API used by Artificial Analysis's tests; Google's official page lists Gemini API, Google AI Studio, Gemini App, Google Antigravity, and other available entry points.

  • Recommended reasoning tier and parameters: Start with low for latency-sensitive tasks; test medium for a quality/latency tradeoff; use high for complex, long-horizon Agents. Google's official documentation lists low, medium, and high, and does not support minimal.

Test environment, versions, and measurement timing

  • v4.2 release snapshot: The Gemini 3.8 Flash article and X post published by Artificial Analysis on 2026-09-02 used Intelligence Index v4.2; high/medium/low were 59/57/52, respectively. The article also reported about 300 tok/s and 2.5 minutes per task for high, and about 0.8 minutes per task for low.

  • v4.3 current snapshot: When the Artificial Analysis model page was opened on 2026-09-08, it was labeled Intelligence Index v4.3; this version includes 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.

  • Performance measurement timing: Artificial Analysis's performance methodology specifies that standard 1k, 10k, and vision workloads are tested eight times per day, with web values shown as the median (P50) over the past 72 hours; the table below is therefore a page snapshot from 2026-09-08, not a fixed permanent value.

  • Service and parameters: The primary performance test server is located in Google Cloud us-central1-a; the default temperature for reasoning models is 0.6, and performance benchmarks count tokens with o200k_base. Intelligence Index costs instead use the token counts reported by the provider API.

  • Independent/self-reported boundary: Intelligence Index, output speed, TTFT, cost, and task-token figures are independently measured or calculated by Artificial Analysis; Google's model status, modalities, context, and pricing are vendor-reported specifications. Artificial Analysis's statement that the tests are "based on the Google API" does not mean Google endorses its evaluation scores.

Results

Current Artificial Analysis v4.3 model pages

Reasoning tierIntelligence IndexOutput speedTTFT (first answer token)Cost per Intelligence Index taskAPI price (input/output per 1M tokens)Context window
high41285.9 tok/s17.36 s$1.24$0.75 / $3.751M
medium40246.3 tok/s8.29 s$0.93$0.75 / $3.751M
low34265.3 tok/s0.78 sUnknown (shown as N/A on the page)$0.75 / $3.751M

All speed, TTFT, and pricing figures above come from the corresponding Artificial Analysis model pages; the pages note that speed and TTFT are based on the Google API and that prices may vary by provider. The cache-hit discount is 90%. The high page also shows approximately 170M output tokens generated cumulatively for the Intelligence Index evaluation, versus approximately 95M for medium; these are total evaluation volumes, not per-request limits.

Artificial Analysis v4.2 release figures

Reasoning tierIntelligence IndexCost per taskAdditional speed/time details
high59$0.58Approximately 300 tok/s; approximately 2.5 minutes per task; approximately 48k average output tokens
medium57$0.41The release article gives the cost but no independent TTFT for this tier
low52$0.24Approximately 0.8 minutes per task; approximately 14k average output tokens

The v4.2 costs use the release article's definition of “per Intelligence Index task.” The article says high costs about 40% more than Gemini 3.7 Flash's $0.40 because average output tokens increased by about 30% and the number of Agent evaluation turns increased. The v4.2 $0.58 and v4.3 page's $1.24 for high cannot be used to infer a pricing trend because the evaluation set changed from v4.2 to v4.3.

Cross-check against official specifications

ItemGoogle's official pageBoundary
API model IDgemini-3.8-flash; stable versionAPI documentation string; does not mean third-party platforms update in sync
Input/outputText, images, video, audio, and PDF input; text outputModality specification, not a guarantee for every client UI
Token limits1,048,576 input; 65,536 outputQuota limits, not recommended values for every request
Reasoning tierslow, medium, high; minimal is not supportedSpecific quality, latency, and token consumption require measurement
StatusGeneral availability / stable versionAvailability by region, account, and third-party provider still requires separate confirmation

Conclusions

  1. First distinguish evaluation versions: The v4.2 scores in the September 2 release article were high 59, medium 57, and low 52; the v4.3 scores on the current model page on September 8 were 41, 40, and 34. The versions, evaluation sets, and page timestamps differ, so the two groups of figures should not be combined into a single ranking.

  2. low is the latency-first tier: The current page shows just 0.78 seconds TTFT and 265.3 tok/s output speed; however, its Intelligence Index is 34, and the v4.3 page does not yet provide its cost per task.

  3. high prioritizes quality and task capability: It has a v4.3 score of 41 and a speed of 285.9 tok/s, but TTFT is 17.36 seconds; interactive products should include reasoning wait time in their experience budget.

  4. Price cannot be judged only per million tokens: Current input/output prices are both $0.75/$3.75, with a 90% cache-hit discount, but per-task cost is also determined by input, cached, reasoning, and answer tokens plus evaluation turns.

  5. The 1M context window is a capability boundary, not a quality guarantee: Both Google and Artificial Analysis list 1M tokens; real long-document workflows still need dedicated testing by input position, retrieval recall, citation accuracy, and end-to-end latency.

Limitations

  • The current Artificial Analysis page uses v4.3, while the release article uses v4.2; the evaluation sets differ, so a “score drop” or cross-version ranking change cannot be calculated directly.

  • Performance figures are P50 values over the past 72 hours; TTFT is affected by server location, network conditions, and provider queuing, so results from us-central1-a may not represent mainland China or other regions.

  • Artificial Analysis's output speed for reasoning models may be calculated from the latter 80% of the answer chunk, while the performance benchmark and Intelligence Index use different token-counting conventions; tokenizer efficiency limits price comparisons.

  • The v4.3 low page shows N/A for per-task cost; its current value cannot be inferred from v4.2's $0.24 or from the high/medium costs.

  • Evaluations such as AA-Briefcase and GDPval-AA v2 include private test sets or internal scoring processes; the platform's public methodology improves reproducibility, but does not constitute a third-party audit.

  • Results reported on the Google DeepMind page for DeepSWE, Vals Finance Agent, Harvey Legal Agent, HLE-Verified, and others are materials published by Google, and this article does not conflate them with independent Artificial Analysis measurements.

Reproduction steps

  1. Fix the API model ID at gemini-3.8-flash, run low, medium, and high separately, and record the first reasoning token, first answer token, full response time, input/output/reasoning tokens, and failed retries.

  2. For the same set of coding, tool-calling, PDF/long-document, and multimodal tasks, calculate TTFT P50/P95, output tok/s, task completion rate, cost per completed task, and human takeover rate.

  3. Repeat the performance measurements across at least two regions or providers, reporting network location, account type, request concurrency, and API version; do not generalize Artificial Analysis's Google API results to other providers.

  4. Save the evaluation-set version (v4.2 or v4.3), page collection time, and pricing snapshot; rerun the same task set whenever the model or Artificial Analysis methodology changes.

  5. Test 100k-, 500k-, and near-1M-token contexts separately for retrieval, citation accuracy, answer tokens, TTFT, and end-to-end latency, and avoid describing “supports 1M” as “quality remains constant within 1M.”

Original evidence and data

  • The current Artificial Analysis high model page records: v4.3 Intelligence Index 41, 285.9 tok/s, 17.36s TTFT, $1.24/task, $0.75/$3.75 input/output pricing, a 90% cache discount, and a 1M context window.

  • The current Artificial Analysis medium/low model pages record: medium at 40, 246.3 tok/s, 8.29s TTFT, and $0.93/task; low at 34, 265.3 tok/s, 0.78s TTFT, and N/A per-task cost; both have a 1M context window.

  • The Artificial Analysis release article dated 2026-09-02 and its corresponding X post record v4.2 high/medium/low at 59/57/52, costs of $0.58/$0.41/$0.24, approximately 300 tok/s and 2.5 minutes per task for high, and approximately 0.8 minutes per task for low.

  • Artificial Analysis's performance methodology states that standard metrics use P50 over the past 72 hours; TTFT runs from request submission to the first token, with the first answer-token latency including reasoning time; output speed is tokens/s after the first chunk is received.

  • The Google AI Developers model page confirms gemini-3.8-flash, 1,048,576 input tokens, 65,536 output tokens, input modalities, and low/medium/high reasoning tiers; the Google DeepMind model page confirms GA status, 1M input, 64k output, and other product specifications.

Source excerpts or observations (short quotation for compliance only)

Artificial Analysis describes high as “notably fast,” while its current page also reports 17.36 seconds TTFT; this shows that output speed and the user-perceived delay before the first answer are two different metrics. Google positions 3.8 Flash as a Flash model for long-horizon software engineering, autonomous Agents, and complex enterprise workflows. Neither source can replace a same-harness rerun for the target business.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Gemini 3.8 Flash

Use and compare models in Tabbit

Gemini 3.8 Flash

Related reviews

OfficialGoogle Blog (The Keyword)2026-09-02

Gemini 3.8 Flash: Google’s Official Benchmarks and Reproduction Boundaries

MediaAI IQ (AIIQ, Liberated Software LLC)2026-09-02

Gemini 3.8 Flash: AI IQ Capability Benchmarks and Task Boundaries

MediaVals AI2026-09-05

Vals AI Finance Agent v2: Professional Finance Agent Benchmark for Gemini 3.8 Flash

MediaSimpleBench official leaderboard and project

SimpleBench: Gemini 3.8 Flash on Everyday Reasoning and Language-Trap Questions

Gemini 3.8 Flash

Related prompts

OfficialGoogle AI for Developers / Google DeepMind2026-09-02

Gemini 3.8 Flash: Google’s Official Model Parameters and API Configuration

OfficialGoogle AI for Developers2026-06-10

Gemini 3.8 Flash: Google's Official Structured Prompting and Agent Workflow

OfficialGoogle AI for Developers

Gemini 3.8 Flash: Google's Official Function-Calling Configuration and Tool Workflow

OfficialGoogle AI for Developers2026-09-02

Gemini 3.8 Flash: Google's Official Structured Output Configuration