Gemini 3.1 Pro

Gemini 3.1 Pro · Reviews and evidence

Which Gemini 3.1 Pro conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

LayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.

LayerLens / Stratix · Read evidence

Artificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.

Artificial Analysis · Read evidence

Selected evidence

OfficialVendor report

Google Officially Releases Gemini 3.1 Pro: ARC-AGI-2 and Product Positioning Baseline

Google's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.

SourceGoogle Blog / Gemini models
Published2026-02-19
Collected2026-08-20
Source-specific observation
The 2026-02-19 Google release covers Gemini 3.1 Pro preview and reports a verified ARC-AGI-2 score of 77.1%.
Published conditions
The release does not publish the ARC prompt, sample split, sampling parameters, or tool trace; the API model name is gemini-3.1-pro-preview.
Capability
Media / benchmarkIndependent measurement

LayerLens Stratix's Six-Benchmark Evaluation of Gemini 3.1 Pro Preview

LayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.

SourceLayerLens / Stratix
Published2026-02-19
Collected2026-08-20
Source-specific observation
The February 19, 2026 LayerLens Stratix test has 14,549 cases across ARC AGI 2, LiveCodeBench, SWE Bench Lite, IF-Evals, BIRD-CRITIC, and BFCL v3.
Published conditions
The article says it used standardized configurations and consistent parameters but does not publish every prompt, temperature, repeat count, or confidence interval.
Capability
Media / benchmarkIndependent measurement

Artificial Analysis's Comprehensive 182-Model Benchmark and End-to-End Latency Evaluation of Gemini 3.1 Pro Preview

Artificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.

SourceArtificial Analysis
Published2026-02-19
Collected2026-08-20
Source-specific observation
The Artificial Analysis tracker, followed from 2026-02-19, compares 182 models priced above $1 per million tokens using Intelligence Index v4.1.1's nine benchmarks.
Published conditions
It records first-party API defaults, total tokens, throughput, TTFT, and cost; prices and preview snapshots can change.
Capability
Media / benchmarkIndependent measurement

MindStudio's Full-Task Evaluation of Three Flagships: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro

MindStudio compares three flagships with HumanEval, SWE-bench, MATH, GPQA, MMLU Pro, and custom long-document tasks; Gemini's context advantage does not generalize to every code repair.

SourceMindStudio Blog
Published2026-03-15
Collected2026-08-20

Unverified: the original source could not be rechecked.

Source-specific observation
The March 15, 2026 MindStudio test includes 164 Python HumanEval pass@1 items, SWE-bench Verified, MATH, GPQA Diamond, MMLU Pro, 120k-token document synthesis, and 5,000-word writing.
Published conditions
Objective tasks were automatically graded and subjective tasks were blind-rated by three independent reviewers across GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro.
Capability

All sources

All sources

7 / 7
OfficialVendor report

Google Officially Releases Gemini 3.1 Pro: ARC-AGI-2 and Product Positioning Baseline

Google's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.

SourceGoogle Blog / Gemini models
Published2026-02-19
Collected2026-08-20
Source-specific observation
The 2026-02-19 Google release covers Gemini 3.1 Pro preview and reports a verified ARC-AGI-2 score of 77.1%.
Published conditions
The release does not publish the ARC prompt, sample split, sampling parameters, or tool trace; the API model name is gemini-3.1-pro-preview.
Capability
Media / benchmarkIndependent measurement

LayerLens Stratix's Six-Benchmark Evaluation of Gemini 3.1 Pro Preview

LayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.

SourceLayerLens / Stratix
Published2026-02-19
Collected2026-08-20
Source-specific observation
The February 19, 2026 LayerLens Stratix test has 14,549 cases across ARC AGI 2, LiveCodeBench, SWE Bench Lite, IF-Evals, BIRD-CRITIC, and BFCL v3.
Published conditions
The article says it used standardized configurations and consistent parameters but does not publish every prompt, temperature, repeat count, or confidence interval.
Capability
Media / benchmarkIndependent measurement

Artificial Analysis's Comprehensive 182-Model Benchmark and End-to-End Latency Evaluation of Gemini 3.1 Pro Preview

Artificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.

SourceArtificial Analysis
Published2026-02-19
Collected2026-08-20
Source-specific observation
The Artificial Analysis tracker, followed from 2026-02-19, compares 182 models priced above $1 per million tokens using Intelligence Index v4.1.1's nine benchmarks.
Published conditions
It records first-party API defaults, total tokens, throughput, TTFT, and cost; prices and preview snapshots can change.
Capability
Media / benchmarkIndependent measurement

MindStudio's Full-Task Evaluation of Three Flagships: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro

MindStudio compares three flagships with HumanEval, SWE-bench, MATH, GPQA, MMLU Pro, and custom long-document tasks; Gemini's context advantage does not generalize to every code repair.

SourceMindStudio Blog
Published2026-03-15
Collected2026-08-20

Unverified: the original source could not be rechecked.

Source-specific observation
The March 15, 2026 MindStudio test includes 164 Python HumanEval pass@1 items, SWE-bench Verified, MATH, GPQA Diamond, MMLU Pro, 120k-token document synthesis, and 5,000-word writing.
Published conditions
Objective tasks were automatically graded and subjective tasks were blind-rated by three independent reviewers across GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro.
Capability
OfficialPersonal experience

Google AI Developers Forum: Empirical Instruction-Following Evaluation of Gemini 3.1 Pro Under a Complex 4,000-Word System Prompt

An Antigravity Ultra user reports that Gemini 3.1 Pro High can skip planning, compress output, drift in long sessions, or refactor out of scope under a roughly 4,000-word engineering instruction set.

SourceGoogle AI Developers Forum
Published2026-04-07
Collected2026-08-20

Unverified: the original source could not be rechecked.

Model/version
Gemini 3.1 Pro; source date 2026-04-07; do not merge snapshots or reasoning tiers.
Platform/harness
Google AI Developers Forum; the source-specific platform and harness remain the unit of observation.
Sample/date boundary
Collected 2026-08-20; Google AI Developers Forum: Empirical Instruction-Following Evaluation of Gemini 3.1 Pro Under a Complex 4,000-Word System Prompt does not establish a universal rate beyond its published sample.
Capability
CommunityPersonal experience

Reddit Discussion of Gemini 3.1 Pro's Static Benchmarks and Arena Deployment Choices

A Reddit discussion contrasts ARC-AGI-2 and HLE release scores with Arena preference rankings, arguing that correctness, tool success, cost, and blind preference should be tested separately before deployment.

SourceReddit / r/LocalLLM
Published2026-02-19
Collected2026-08-20

Unverified: the original source could not be rechecked.

Model/version
Gemini 3.1 Pro; source date 2026-02-19; do not merge snapshots or reasoning tiers.
Platform/harness
Reddit / r/LocalLLM; the source-specific platform and harness remain the unit of observation.
Sample/date boundary
Collected 2026-08-20; Reddit Discussion of Gemini 3.1 Pro's Static Benchmarks and Arena Deployment Choices does not establish a universal rate beyond its published sample.
Capability
CommunityPersonal experience

Reddit Community Hands-on: Gemini 3.1 Pro Extended Thinking, 1M Context Synthesis, and API vs. Web Differences

Reddit power users report stronger million-token synthesis and long-session state retention with high thinking in AI Studio or the API than in the consumer web client, while noting frequent serving changes.

SourceReddit / r/GeminiAI
Published2026-08-08
Collected2026-08-20

Unverified: the original source could not be rechecked.

Model/version
Gemini 3.1 Pro; source date 2026-08-08; do not merge snapshots or reasoning tiers.
Platform/harness
Reddit / r/GeminiAI; the source-specific platform and harness remain the unit of observation.
Sample/date boundary
Collected 2026-08-20; Reddit Community Hands-on: Gemini 3.1 Pro Extended Thinking, 1M Context Synthesis, and API vs. Web Differences does not establish a universal rate beyond its published sample.
Capability

Gemini 3.1 Pro

Compare Gemini 3.1 Pro in Tabbit

Model access, features, and permissions depend on your current client account.