Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Gemini 3.1 Pro · Official source · Vendor report

Google Officially Releases Gemini 3.1 Pro: ARC-AGI-2 and Product Positioning Baseline

Google's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Source-specific observation
The 2026-02-19 Google release covers Gemini 3.1 Pro preview and reports a verified ARC-AGI-2 score of 77.1%.
Published conditions
The release does not publish the ARC prompt, sample split, sampling parameters, or tool trace; the API model name is gemini-3.1-pro-preview.

Key data and applicable tasks

One-sentence takeaway

Google positions Gemini 3.1 Pro as a core model for complex problems, multimodal reasoning, and agentic workflows, and reports an ARC-AGI-2 verified score of 77.1%; however, the release page provides only limited benchmark context.

Test environment

  • Version: Gemini 3.1 Pro, released as a preview on 2026-02-19.

  • Access points: Gemini API/AI Studio, Gemini CLI, Google Antigravity, Android Studio, Vertex AI, Gemini Enterprise, Gemini app, and NotebookLM.

  • Benchmark: ARC-AGI-2, which the official description says tests a model's ability to solve entirely new logical patterns.

  • Configuration: The release page calls this a verified score but does not disclose the complete prompts, sample split, sampling parameters, or tool traces in the article.

Inputs/configuration

The official release page does not provide the complete ARC-AGI-2 inputs or execution harness. At the product-example level, Google presents the model for complex problems, multimodal interpretation, data synthesis, and creative projects; the developer preview path is gemini-3.1-pro-preview.

Results

  • ARC-AGI-2 verified: 77.1%.

  • Google says this score is more than twice that of Gemini 3 Pro; the release page does not provide Gemini 3 Pro's complete configuration information in the same article.

  • Product positioning: advanced reasoning across complex problems and modalities; at launch, Google explicitly said it would continue validating ambitious agentic workflows before gradually moving toward general availability.

Conclusion

The ARC-AGI-2 result supports treating Gemini 3.1 Pro as a candidate with strong abstract reasoning; if the user's task involves coding agents, SQL, or multi-tool orchestration, it must be evaluated separately at the task level, and cannot be inferred from a single ARC score.

Limitations

  • This is an official self-report; the release article does not provide complete raw samples, prompts, random seeds, costs, failure types, or confidence intervals.

  • “Verified” indicates that a verification process took place, but it is still not equivalent to an independent third-party reproduction.

  • The preview model's API, pricing, rate limits, and behavior may change; the launch-day state should not be treated as a permanent specification.

  • The article explicitly presents agentic workflows as an area for continued validation, so it cannot be claimed that the model's Agent capabilities are already comprehensively stable.

Reproduction steps

  1. Use a fixed gemini-3.1-pro-preview snapshot and the official permitted ARC-AGI-2 evaluation protocol.

  2. Record the thinking level, temperature, tools, input version, outputs, and time spent per question; do not mix Gemini app results with API results.

  3. Build separate coding, SQL, and multi-tool task sets, and report success rate, tool-call accuracy, latency, and cost as separate metrics.

  4. Use the official 77.1% as the release baseline, clearly indicating whether your own harness is comparable.

What this supports

  • It supports ARC-AGI-2 as an official positioning signal

What this does not support

  • It supports ARC-AGI-2 as an official positioning signal, not score reproduction, universal task win rates, or a stable preview SLA.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Google Blog / Gemini models · The Gemini Team · Original publication date 2026-02-19 · Site edit date 2026-09-20

Open original source

Gemini 3.1 Pro

Compare Gemini 3.1 Pro in Tabbit

Download the Tabbit client to check model access

Related reviews

LayerLens Stratix's Six-Benchmark Evaluation of Gemini 3.1 Pro PreviewLayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.Artificial Analysis's Comprehensive 182-Model Benchmark and End-to-End Latency Evaluation of Gemini 3.1 Pro PreviewArtificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.MindStudio's Full-Task Evaluation of Three Flagships: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 ProMindStudio compares three flagships with HumanEval, SWE-bench, MATH, GPQA, MMLU Pro, and custom long-document tasks; Gemini's context advantage does not generalize to every code repair.Google AI Developers Forum: Empirical Instruction-Following Evaluation of Gemini 3.1 Pro Under a Complex 4,000-Word System PromptAn Antigravity Ultra user reports that Gemini 3.1 Pro High can skip planning, compress output, drift in long sessions, or refactor out of scope under a roughly 4,000-word engineering instruction set.Gemini 3.1 Pro: Concise Prompting and Long-Context Question PlacementGoogle's Gemini 3 guide recommends direct, concise prompts and placing the specific question after long context with a short anchoring phrase.Gemini 3.1 Pro Thinking Levels, Structured Outputs, and Tool ConfigurationGoogle's official documentation combines thinking_level, default temperature, tool calls, and JSON schema checks, while separating the customtools endpoint.Spec-Driven Coding Workflow: Claude-Led Planning and Gemini-Isolated ExecutionThe developer-forum case uses Claude for specification and audit, Gemini 3.1 Pro for isolated execution in fresh sessions, and a final audit for changes.Gemini 3.1 Pro: Open-Source Architecture Alignment and Multi-Model Pair Programming WorkflowThis Antigravity community case injects a mature open-source project's architecture into Gemini 3.1 Pro and uses a second model for cross-review and alignment.