Gemini 3.8 Flash's speed and latency vary substantially by reasoning tier: Artificial Analysis's current v4.3 pages record 285.9 tok/s and 17.36 seconds to the first answer token for high, 246.3 tok/s and 8.29 seconds for medium, and 265.3 tok/s and 0.78 seconds for low; however, the Intelligence Index also changed from the v4.2 figures of 59/57/52 at release on 2026-09-02 to 41/40/34 in the current v4.3 snapshot. These figures come from different test versions and cannot be directly compared as evidence that the model has regressed.
Suitable tasks: Coding, Agent, document, and enterprise workflows that require long context, multimodal input, tool calling, and relatively fast answer generation.
Unsuitable tasks: Using the high tier as the default for low-latency chat, or treating one-off tok/s as total task duration; high has materially higher TTFT and reasoning-token overhead.
Applicable model version: gemini-3.8-flash; this article records Artificial Analysis's high, medium, and low reasoning pages without treating the three tiers as three API models.
Applicable client, Agent, or API: The Google API used by Artificial Analysis's tests; Google's official page lists Gemini API, Google AI Studio, Gemini App, Google Antigravity, and other available entry points.
Recommended reasoning tier and parameters: Start with low for latency-sensitive tasks; test medium for a quality/latency tradeoff; use high for complex, long-horizon Agents. Google's official documentation lists low, medium, and high, and does not support minimal.
v4.2 release snapshot: The Gemini 3.8 Flash article and X post published by Artificial Analysis on 2026-09-02 used Intelligence Index v4.2; high/medium/low were 59/57/52, respectively. The article also reported about 300 tok/s and 2.5 minutes per task for high, and about 0.8 minutes per task for low.
v4.3 current snapshot: When the Artificial Analysis model page was opened on 2026-09-08, it was labeled Intelligence Index v4.3; this version includes 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.
Performance measurement timing: Artificial Analysis's performance methodology specifies that standard 1k, 10k, and vision workloads are tested eight times per day, with web values shown as the median (P50) over the past 72 hours; the table below is therefore a page snapshot from 2026-09-08, not a fixed permanent value.
Service and parameters: The primary performance test server is located in Google Cloud us-central1-a; the default temperature for reasoning models is 0.6, and performance benchmarks count tokens with o200k_base. Intelligence Index costs instead use the token counts reported by the provider API.
Independent/self-reported boundary: Intelligence Index, output speed, TTFT, cost, and task-token figures are independently measured or calculated by Artificial Analysis; Google's model status, modalities, context, and pricing are vendor-reported specifications. Artificial Analysis's statement that the tests are "based on the Google API" does not mean Google endorses its evaluation scores.
| Reasoning tier | Intelligence Index | Output speed | TTFT (first answer token) | Cost per Intelligence Index task | API price (input/output per 1M tokens) | Context window |
|---|---|---|---|---|---|---|
| high | 41 | 285.9 tok/s | 17.36 s | $1.24 | $0.75 / $3.75 | 1M |
| medium | 40 | 246.3 tok/s | 8.29 s | $0.93 | $0.75 / $3.75 | 1M |
| low | 34 | 265.3 tok/s | 0.78 s | Unknown (shown as N/A on the page) | $0.75 / $3.75 | 1M |
All speed, TTFT, and pricing figures above come from the corresponding Artificial Analysis model pages; the pages note that speed and TTFT are based on the Google API and that prices may vary by provider. The cache-hit discount is 90%. The high page also shows approximately 170M output tokens generated cumulatively for the Intelligence Index evaluation, versus approximately 95M for medium; these are total evaluation volumes, not per-request limits.
| Reasoning tier | Intelligence Index | Cost per task | Additional speed/time details |
|---|---|---|---|
| high | 59 | $0.58 | Approximately 300 tok/s; approximately 2.5 minutes per task; approximately 48k average output tokens |
| medium | 57 | $0.41 | The release article gives the cost but no independent TTFT for this tier |
| low | 52 | $0.24 | Approximately 0.8 minutes per task; approximately 14k average output tokens |
The v4.2 costs use the release article's definition of “per Intelligence Index task.” The article says high costs about 40% more than Gemini 3.7 Flash's $0.40 because average output tokens increased by about 30% and the number of Agent evaluation turns increased. The v4.2 $0.58 and v4.3 page's $1.24 for high cannot be used to infer a pricing trend because the evaluation set changed from v4.2 to v4.3.
| Item | Google's official page | Boundary |
|---|---|---|
| API model ID | gemini-3.8-flash; stable version | API documentation string; does not mean third-party platforms update in sync |
| Input/output | Text, images, video, audio, and PDF input; text output | Modality specification, not a guarantee for every client UI |
| Token limits | 1,048,576 input; 65,536 output | Quota limits, not recommended values for every request |
| Reasoning tiers | low, medium, high; minimal is not supported | Specific quality, latency, and token consumption require measurement |
| Status | General availability / stable version | Availability by region, account, and third-party provider still requires separate confirmation |
First distinguish evaluation versions: The v4.2 scores in the September 2 release article were high 59, medium 57, and low 52; the v4.3 scores on the current model page on September 8 were 41, 40, and 34. The versions, evaluation sets, and page timestamps differ, so the two groups of figures should not be combined into a single ranking.
low is the latency-first tier: The current page shows just 0.78 seconds TTFT and 265.3 tok/s output speed; however, its Intelligence Index is 34, and the v4.3 page does not yet provide its cost per task.
high prioritizes quality and task capability: It has a v4.3 score of 41 and a speed of 285.9 tok/s, but TTFT is 17.36 seconds; interactive products should include reasoning wait time in their experience budget.
Price cannot be judged only per million tokens: Current input/output prices are both $0.75/$3.75, with a 90% cache-hit discount, but per-task cost is also determined by input, cached, reasoning, and answer tokens plus evaluation turns.
The 1M context window is a capability boundary, not a quality guarantee: Both Google and Artificial Analysis list 1M tokens; real long-document workflows still need dedicated testing by input position, retrieval recall, citation accuracy, and end-to-end latency.
The current Artificial Analysis page uses v4.3, while the release article uses v4.2; the evaluation sets differ, so a “score drop” or cross-version ranking change cannot be calculated directly.
Performance figures are P50 values over the past 72 hours; TTFT is affected by server location, network conditions, and provider queuing, so results from us-central1-a may not represent mainland China or other regions.
Artificial Analysis's output speed for reasoning models may be calculated from the latter 80% of the answer chunk, while the performance benchmark and Intelligence Index use different token-counting conventions; tokenizer efficiency limits price comparisons.
The v4.3 low page shows N/A for per-task cost; its current value cannot be inferred from v4.2's $0.24 or from the high/medium costs.
Evaluations such as AA-Briefcase and GDPval-AA v2 include private test sets or internal scoring processes; the platform's public methodology improves reproducibility, but does not constitute a third-party audit.
Results reported on the Google DeepMind page for DeepSWE, Vals Finance Agent, Harvey Legal Agent, HLE-Verified, and others are materials published by Google, and this article does not conflate them with independent Artificial Analysis measurements.
Fix the API model ID at gemini-3.8-flash, run low, medium, and high separately, and record the first reasoning token, first answer token, full response time, input/output/reasoning tokens, and failed retries.
For the same set of coding, tool-calling, PDF/long-document, and multimodal tasks, calculate TTFT P50/P95, output tok/s, task completion rate, cost per completed task, and human takeover rate.
Repeat the performance measurements across at least two regions or providers, reporting network location, account type, request concurrency, and API version; do not generalize Artificial Analysis's Google API results to other providers.
Save the evaluation-set version (v4.2 or v4.3), page collection time, and pricing snapshot; rerun the same task set whenever the model or Artificial Analysis methodology changes.
Test 100k-, 500k-, and near-1M-token contexts separately for retrieval, citation accuracy, answer tokens, TTFT, and end-to-end latency, and avoid describing “supports 1M” as “quality remains constant within 1M.”
The current Artificial Analysis high model page records: v4.3 Intelligence Index 41, 285.9 tok/s, 17.36s TTFT, $1.24/task, $0.75/$3.75 input/output pricing, a 90% cache discount, and a 1M context window.
The current Artificial Analysis medium/low model pages record: medium at 40, 246.3 tok/s, 8.29s TTFT, and $0.93/task; low at 34, 265.3 tok/s, 0.78s TTFT, and N/A per-task cost; both have a 1M context window.
The Artificial Analysis release article dated 2026-09-02 and its corresponding X post record v4.2 high/medium/low at 59/57/52, costs of $0.58/$0.41/$0.24, approximately 300 tok/s and 2.5 minutes per task for high, and approximately 0.8 minutes per task for low.
Artificial Analysis's performance methodology states that standard metrics use P50 over the past 72 hours; TTFT runs from request submission to the first token, with the first answer-token latency including reasoning time; output speed is tokens/s after the first chunk is received.
The Google AI Developers model page confirms gemini-3.8-flash, 1,048,576 input tokens, 65,536 output tokens, input modalities, and low/medium/high reasoning tiers; the Google DeepMind model page confirms GA status, 1M input, 64k output, and other product specifications.
Artificial Analysis describes high as “notably fast,” while its current page also reports 17.36 seconds TTFT; this shows that output speed and the user-perceived delay before the first answer are two different metrics. Google positions 3.8 Flash as a Flash model for long-horizon software engineering, autonomous Agents, and complex enterprise workflows. Neither source can replace a same-harness rerun for the target business.
Gemini 3.8 Flash