Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Gemini 3.1 Pro · Media / benchmark · Independent measurement

Artificial Analysis's Comprehensive 182-Model Benchmark and End-to-End Latency Evaluation of Gemini 3.1 Pro Preview

Artificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
The Artificial Analysis tracker, followed from 2026-02-19, compares 182 models priced above $1 per million tokens using Intelligence Index v4.1.1's nine benchmarks.
Published conditions
It records first-party API defaults, total tokens, throughput, TTFT, and cost; prices and preview snapshots can change.

Key data and applicable tasks

One-sentence takeaway

Independent testing by Artificial Analysis shows that Gemini 3.1 Pro Preview achieves an Intelligence Index score of 48 (vs. a price-tier median of 35) with an output generation speed of 121.4 t/s (vs. a price-tier median of 76.2 t/s); however, driven by its extended reasoning chain, its time to first token (TTFT) reaches 32.45s, making it a classic reasoning model characterized by high intelligence, high throughput, and high upfront reasoning latency.

Test environment

  • Evaluation organization: Artificial Analysis independent benchmark platform.

  • Evaluation target: gemini-3.1-pro-preview (Google official first-party API endpoint).

  • Comparison baseline: 182 mainstream large language models within the same price tier (>$1.00/1M tokens).

  • Evaluation framework: Artificial Analysis Intelligence Index v4.1.1 (aggregating 9 benchmarks: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR).

  • Test configuration: Standardized on first-party API default parameters and reasoning settings, measuring end-to-end token consumption, output generation throughput, time to first token (TTFT), and API invocation costs.

Input/configuration

  • Input modalities: Text, image, audio, and video.

  • Output modality: Text.

  • Context window: 1,000,000 tokens (1M).

  • API pricing: $2.00 / 1M tokens for input, $12.00 / 1M tokens for output, with a 90% discount on context cache hits (Cache Hit: $0.20 / 1M tokens).

Results

Evaluation MetricGemini 3.1 Pro MeasuredPrice-Tier Median / RankKey Observations & Characteristics
Intelligence Index (Overall Intelligence)4835 (Rank #40 / 182)Significantly above the peer median, ranking in the top tier
Output Speed (Generation Throughput)121.4 t/s76.2 t/s (Rank #35 / 182)Extremely fast decoding generation, well-suited for batch throughput
Time to First Token (TTFT)32.45s2.79sIncludes deep "thinking" phase, resulting in longer upfront waiting time
Total Generated Tokens (Verbosity)56M72M (Rank #30 / 182)More concise with lower verbosity compared to peer reasoning models
Weighted Cost per Task$0.33Mid-to-high tierDriven by the $12/1M output pricing combined with multi-step reasoning
Total Evaluation Cost (Full Benchmark Suite)$811.90-Standard full evaluation cost covering all 9 sub-benchmarks

Conclusion

In independent controlled benchmarks, Gemini 3.1 Pro Preview demonstrates a standout profile of "high-speed decoding + concise output" (121.4 t/s, with significantly fewer generated tokens than the peer median), making it well-suited for long-document reasoning and high-throughput analytical workloads. However, its 32.45s TTFT makes it unsuitable for low-latency real-time conversational applications or interactive UIs that require sub-second initial responses.

Limitations

  • Testing is based on preview snapshots from the official Google API; performance in production environments may experience fluctuations due to server load variance.

  • The Intelligence Index does not isolate or break down the marginal performance-latency tradeoffs of different thinking_level configurations (low/medium/high) across individual academic sub-benchmarks.

  • Evaluations were conducted exclusively via the first-party API, without accounting for routing overhead introduced by third-party distribution platforms (e.g., OpenRouter, Vertex AI proxies).

Reproduction steps

  1. Access the Google GenAI API and specify gemini-3.1-pro-preview as the model.

  2. Execute the evaluation scripts for the 9 open-source or controlled sub-datasets corresponding to the Artificial Analysis Intelligence Index v4.1.1.

  3. Enable streaming to record the duration from request dispatch to receipt of the first answer token (TTFT) and the subsequent token generation throughput (tokens per second).

  4. Aggregate input and output token counts, and calculate actual per-task costs based on official API pricing.

What this supports

  • It supports like-for-like intelligence, generation-rate, TTFT, and cost comparison

What this does not support

  • It supports like-for-like intelligence, generation-rate, TTFT, and cost comparison, not a repository success rate or fixed price.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Artificial Analysis Research Team · Original publication date 2026-02-19 · Site edit date 2026-09-20

Open original source

Gemini 3.1 Pro

Compare Gemini 3.1 Pro in Tabbit

Download the Tabbit client to check model access

Related reviews

LayerLens Stratix's Six-Benchmark Evaluation of Gemini 3.1 Pro PreviewLayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.MindStudio's Full-Task Evaluation of Three Flagships: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 ProMindStudio compares three flagships with HumanEval, SWE-bench, MATH, GPQA, MMLU Pro, and custom long-document tasks; Gemini's context advantage does not generalize to every code repair.Google Officially Releases Gemini 3.1 Pro: ARC-AGI-2 and Product Positioning BaselineGoogle's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.Google AI Developers Forum: Empirical Instruction-Following Evaluation of Gemini 3.1 Pro Under a Complex 4,000-Word System PromptAn Antigravity Ultra user reports that Gemini 3.1 Pro High can skip planning, compress output, drift in long sessions, or refactor out of scope under a roughly 4,000-word engineering instruction set.Gemini 3.1 Pro: Concise Prompting and Long-Context Question PlacementGoogle's Gemini 3 guide recommends direct, concise prompts and placing the specific question after long context with a short anchoring phrase.Gemini 3.1 Pro Thinking Levels, Structured Outputs, and Tool ConfigurationGoogle's official documentation combines thinking_level, default temperature, tool calls, and JSON schema checks, while separating the customtools endpoint.Spec-Driven Coding Workflow: Claude-Led Planning and Gemini-Isolated ExecutionThe developer-forum case uses Claude for specification and audit, Gemini 3.1 Pro for isolated execution in fresh sessions, and a final audit for changes.Gemini 3.1 Pro: Open-Source Architecture Alignment and Multi-Model Pair Programming WorkflowThis Antigravity community case injects a mature open-source project's architecture into Gemini 3.1 Pro and uses a second model for cross-review and alignment.