Google's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.
Google Blog / Gemini models · Read evidenceGemini 3.1 Pro · Reviews and evidence
Which Gemini 3.1 Pro conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
LayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.
LayerLens / Stratix · Read evidenceArtificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.
Artificial Analysis · Read evidenceSelected evidence
Google Officially Releases Gemini 3.1 Pro: ARC-AGI-2 and Product Positioning Baseline
Google's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.
- Source-specific observation
- The 2026-02-19 Google release covers Gemini 3.1 Pro preview and reports a verified ARC-AGI-2 score of 77.1%.
- Published conditions
- The release does not publish the ARC prompt, sample split, sampling parameters, or tool trace; the API model name is gemini-3.1-pro-preview.
LayerLens Stratix's Six-Benchmark Evaluation of Gemini 3.1 Pro Preview
LayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.
- Source-specific observation
- The February 19, 2026 LayerLens Stratix test has 14,549 cases across ARC AGI 2, LiveCodeBench, SWE Bench Lite, IF-Evals, BIRD-CRITIC, and BFCL v3.
- Published conditions
- The article says it used standardized configurations and consistent parameters but does not publish every prompt, temperature, repeat count, or confidence interval.
Artificial Analysis's Comprehensive 182-Model Benchmark and End-to-End Latency Evaluation of Gemini 3.1 Pro Preview
Artificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.
- Source-specific observation
- The Artificial Analysis tracker, followed from 2026-02-19, compares 182 models priced above $1 per million tokens using Intelligence Index v4.1.1's nine benchmarks.
- Published conditions
- It records first-party API defaults, total tokens, throughput, TTFT, and cost; prices and preview snapshots can change.
MindStudio's Full-Task Evaluation of Three Flagships: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro
MindStudio compares three flagships with HumanEval, SWE-bench, MATH, GPQA, MMLU Pro, and custom long-document tasks; Gemini's context advantage does not generalize to every code repair.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The March 15, 2026 MindStudio test includes 164 Python HumanEval pass@1 items, SWE-bench Verified, MATH, GPQA Diamond, MMLU Pro, 120k-token document synthesis, and 5,000-word writing.
- Published conditions
- Objective tasks were automatically graded and subjective tasks were blind-rated by three independent reviewers across GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro.
All sources
All sources
Google Officially Releases Gemini 3.1 Pro: ARC-AGI-2 and Product Positioning Baseline
Google's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.
- Source-specific observation
- The 2026-02-19 Google release covers Gemini 3.1 Pro preview and reports a verified ARC-AGI-2 score of 77.1%.
- Published conditions
- The release does not publish the ARC prompt, sample split, sampling parameters, or tool trace; the API model name is gemini-3.1-pro-preview.
LayerLens Stratix's Six-Benchmark Evaluation of Gemini 3.1 Pro Preview
LayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.
- Source-specific observation
- The February 19, 2026 LayerLens Stratix test has 14,549 cases across ARC AGI 2, LiveCodeBench, SWE Bench Lite, IF-Evals, BIRD-CRITIC, and BFCL v3.
- Published conditions
- The article says it used standardized configurations and consistent parameters but does not publish every prompt, temperature, repeat count, or confidence interval.
Artificial Analysis's Comprehensive 182-Model Benchmark and End-to-End Latency Evaluation of Gemini 3.1 Pro Preview
Artificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.
- Source-specific observation
- The Artificial Analysis tracker, followed from 2026-02-19, compares 182 models priced above $1 per million tokens using Intelligence Index v4.1.1's nine benchmarks.
- Published conditions
- It records first-party API defaults, total tokens, throughput, TTFT, and cost; prices and preview snapshots can change.
MindStudio's Full-Task Evaluation of Three Flagships: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro
MindStudio compares three flagships with HumanEval, SWE-bench, MATH, GPQA, MMLU Pro, and custom long-document tasks; Gemini's context advantage does not generalize to every code repair.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The March 15, 2026 MindStudio test includes 164 Python HumanEval pass@1 items, SWE-bench Verified, MATH, GPQA Diamond, MMLU Pro, 120k-token document synthesis, and 5,000-word writing.
- Published conditions
- Objective tasks were automatically graded and subjective tasks were blind-rated by three independent reviewers across GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro.
Google AI Developers Forum: Empirical Instruction-Following Evaluation of Gemini 3.1 Pro Under a Complex 4,000-Word System Prompt
An Antigravity Ultra user reports that Gemini 3.1 Pro High can skip planning, compress output, drift in long sessions, or refactor out of scope under a roughly 4,000-word engineering instruction set.
Unverified: the original source could not be rechecked.
- Model/version
- Gemini 3.1 Pro; source date 2026-04-07; do not merge snapshots or reasoning tiers.
- Platform/harness
- Google AI Developers Forum; the source-specific platform and harness remain the unit of observation.
- Sample/date boundary
- Collected 2026-08-20; Google AI Developers Forum: Empirical Instruction-Following Evaluation of Gemini 3.1 Pro Under a Complex 4,000-Word System Prompt does not establish a universal rate beyond its published sample.
Reddit Discussion of Gemini 3.1 Pro's Static Benchmarks and Arena Deployment Choices
A Reddit discussion contrasts ARC-AGI-2 and HLE release scores with Arena preference rankings, arguing that correctness, tool success, cost, and blind preference should be tested separately before deployment.
Unverified: the original source could not be rechecked.
- Model/version
- Gemini 3.1 Pro; source date 2026-02-19; do not merge snapshots or reasoning tiers.
- Platform/harness
- Reddit / r/LocalLLM; the source-specific platform and harness remain the unit of observation.
- Sample/date boundary
- Collected 2026-08-20; Reddit Discussion of Gemini 3.1 Pro's Static Benchmarks and Arena Deployment Choices does not establish a universal rate beyond its published sample.
Reddit Community Hands-on: Gemini 3.1 Pro Extended Thinking, 1M Context Synthesis, and API vs. Web Differences
Reddit power users report stronger million-token synthesis and long-session state retention with high thinking in AI Studio or the API than in the consumer web client, while noting frequent serving changes.
Unverified: the original source could not be rechecked.
- Model/version
- Gemini 3.1 Pro; source date 2026-08-08; do not merge snapshots or reasoning tiers.
- Platform/harness
- Reddit / r/GeminiAI; the source-specific platform and harness remain the unit of observation.
- Sample/date boundary
- Collected 2026-08-20; Reddit Community Hands-on: Gemini 3.1 Pro Extended Thinking, 1M Context Synthesis, and API vs. Web Differences does not establish a universal rate beyond its published sample.
Gemini 3.1 Pro
Compare Gemini 3.1 Pro in Tabbit
Model access, features, and permissions depend on your current client account.