Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Gemini 3.1 Pro · Media / benchmark · Independent measurement

LayerLens Stratix's Six-Benchmark Evaluation of Gemini 3.1 Pro Preview

LayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
The February 19, 2026 LayerLens Stratix test has 14,549 cases across ARC AGI 2, LiveCodeBench, SWE Bench Lite, IF-Evals, BIRD-CRITIC, and BFCL v3.
Published conditions
The article says it used standardized configurations and consistent parameters but does not publish every prompt, temperature, repeat count, or confidence interval.

Key data and applicable tasks

One-sentence takeaway

Across 14,549 test cases on Stratix, LayerLens measured a wide task spread for Gemini 3.1 Pro Preview, from 92.3% on ARC-AGI-2 to 32.5% on BIRD-CRITIC. The results suggest that it is better suited to abstract reasoning, while SQL and repository-level engineering require routing or supporting tools.

Test environment

  • Platform: LayerLens Stratix.

  • Sample: 14,549 test cases.

  • Categories/benchmarks: ARC AGI 2, LiveCodeBench, SWE Bench Lite, IF-Evals, BIRD-CRITIC, and BFCL v3, covering abstract reasoning, coding, software engineering, instruction following, SQL, and function calling.

  • Configuration: The article says it used standardized benchmark configurations and consistent parameters; it does not disclose each benchmark's complete prompt, temperature, number of repetitions, or confidence interval.

Input/configuration

The article reports an evaluation of Gemini 3.1 Pro Preview dated 2026-02-19, but does not state whether it exactly matches the specific snapshot of the public API gemini-3.1-pro-preview. Testing was distributed across six standard benchmarks, and the results came from automated evaluations rather than selected demos.

Results

BenchmarkScoreObservation
ARC AGI 292.3%Strongest result: abstract patterns and multi-step reasoning
IF-Evals78.4%Generally good with multiple instructions, but fails on conflicting constraints
BFCL v371.2%Function calling is above average, but complex multi-tool chains are unstable
LiveCodeBench61.0%Standard coding tasks are acceptable; performance declines on long or multi-file tasks
SWE Bench Lite48.7%Moderate performance on repository-level engineering; cross-file dependencies are difficult
BIRD-CRITIC32.5%Weak at SQL and data reasoning; prone to errors with multi-table joins, subqueries, and schema inference

Conclusion

This evaluation supports routing by task: try Gemini 3.1 Pro first for abstract reasoning and structured instructions; for SQL, repository-level code, and complex multi-tool chains, pair it with schema, test, and tool validation, or evaluate other models.

Limitations

  • The 92.3% on ARC-AGI-2 does not match Google's published verified score of 77.1%. The difference may come from the dataset, harness, decoding, or model version; the rankings cannot be compared directly.

  • LayerLens did not publish the complete original inputs, run logs, seeds, costs, or per-question result downloads in the article. Reproduction requires Stratix or obtaining the data separately.

  • The article's author is LayerLens's marketing lead, and the platform has a product relationship; the evidence is stronger than an opinion but still requires third-party verification.

  • “2M context” does not match the 1,048,576 input tokens collected from the Google API model page; use the current specifications for the target endpoint.

Reproduction steps

  1. Fix the model snapshot, API endpoint, thinking level, temperature, and max output.

  2. Use the same versions of ARC AGI 2, IF-Evals, BFCL v3, LiveCodeBench, SWE Bench Lite, and BIRD-CRITIC; record each sample's input, output, tool calls, and failure reason.

  3. Report all six scores and confidence intervals; do not let an average conceal the task gap between 92.3% and 32.5%.

  4. Show the Google 77.1% baseline alongside these results and explain dataset/harness differences in the report.

What this supports

  • It supports routing by abstract reasoning, software engineering, SQL, and function calling

What this does not support

  • It supports routing by abstract reasoning, software engineering, SQL, and function calling; it does not establish a current API snapshot or universal ranking.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

LayerLens / Stratix · Jake Meany (LayerLens) · Original publication date 2026-02-19 · Site edit date 2026-09-20

Open original source

Gemini 3.1 Pro

Compare Gemini 3.1 Pro in Tabbit

Download the Tabbit client to check model access

Related reviews

Artificial Analysis's Comprehensive 182-Model Benchmark and End-to-End Latency Evaluation of Gemini 3.1 Pro PreviewArtificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.MindStudio's Full-Task Evaluation of Three Flagships: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 ProMindStudio compares three flagships with HumanEval, SWE-bench, MATH, GPQA, MMLU Pro, and custom long-document tasks; Gemini's context advantage does not generalize to every code repair.Google Officially Releases Gemini 3.1 Pro: ARC-AGI-2 and Product Positioning BaselineGoogle's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.Google AI Developers Forum: Empirical Instruction-Following Evaluation of Gemini 3.1 Pro Under a Complex 4,000-Word System PromptAn Antigravity Ultra user reports that Gemini 3.1 Pro High can skip planning, compress output, drift in long sessions, or refactor out of scope under a roughly 4,000-word engineering instruction set.Gemini 3.1 Pro: Concise Prompting and Long-Context Question PlacementGoogle's Gemini 3 guide recommends direct, concise prompts and placing the specific question after long context with a short anchoring phrase.Gemini 3.1 Pro Thinking Levels, Structured Outputs, and Tool ConfigurationGoogle's official documentation combines thinking_level, default temperature, tool calls, and JSON schema checks, while separating the customtools endpoint.Spec-Driven Coding Workflow: Claude-Led Planning and Gemini-Isolated ExecutionThe developer-forum case uses Claude for specification and audit, Gemini 3.1 Pro for isolated execution in fresh sessions, and a final audit for changes.Gemini 3.1 Pro: Open-Source Architecture Alignment and Multi-Model Pair Programming WorkflowThis Antigravity community case injects a mature open-source project's architecture into Gemini 3.1 Pro and uses a second model for cross-review and alignment.