Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Gemini 3.1 Pro · Media / benchmark · Independent measurement

MindStudio's Full-Task Evaluation of Three Flagships: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro

MindStudio compares three flagships with HumanEval, SWE-bench, MATH, GPQA, MMLU Pro, and custom long-document tasks; Gemini's context advantage does not generalize to every code repair.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
The March 15, 2026 MindStudio test includes 164 Python HumanEval pass@1 items, SWE-bench Verified, MATH, GPQA Diamond, MMLU Pro, 120k-token document synthesis, and 5,000-word writing.
Published conditions
Objective tasks were automatically graded and subjective tasks were blind-rated by three independent reviewers across GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro.

Key data and applicable tasks

One-sentence takeaway

In a unified multi-model benchmark spanning code generation, long-form creative writing, graduate-level reasoning, mathematics, and long-document synthesis, Gemini 3.1 Pro holds an overwhelming advantage in its 2M-token ultra-long context and output cost (60% cheaper than Claude). However, it trails GPT-5.4 and Claude Opus 4.6 in complex multi-file bug fixing and literary-grade subjective writing.

Test environment

  • Platform: MindStudio model evaluation and multi-agent orchestration pipeline.

  • Evaluated models: OpenAI GPT-5.4, Anthropic Claude Opus 4.6, Google Gemini 3.1 Pro.

  • Evaluated benchmarks and tasks:

    • Standard benchmarks: HumanEval (164 Python problems, pass@1), SWE-bench Verified (real-world GitHub issue resolution rate), MATH (competition-grade mathematics), GPQA Diamond (high-difficulty scientific reasoning), and MMLU Pro (comprehensive evaluation across 57 academic disciplines).

    • Custom empirical tasks: 5,000-word long-form literary writing, brand marketing copy (strict constraint following), 120k-token multi-document research synthesis report, and 6 categories of SVG spatial and visual code generation.

  • Evaluation method: Automated grading for objective coding and math tasks; blind review by 3 independent human evaluators for subjective tasks, scored across prose quality, instruction following, and narrative coherence.

Inputs/configuration

  • All models used identical prompts and default temperature settings across standard benchmarks.

  • The long-document synthesis task used a uniform input of multiple research reports and academic materials totaling 80,000–150,000 tokens.

  • Pricing and parameter specifications:

    • GPT-5.4: Input $15.00 / 1M, Output $60.00 / 1M, Context 128k tokens, Generation speed ~80 TPS

    • Claude Opus 4.6: Input $20.00 / 1M, Output $100.00 / 1M, Context 200k tokens, Generation speed ~55 TPS

    • Gemini 3.1 Pro: Input $12.50 / 1M, Output $37.50 / 1M, Context 2M tokens, Generation speed ~75 TPS

Results

1. Standard Benchmark Comparison

BenchmarkGPT-5.4Claude Opus 4.6Gemini 3.1 ProAnalysis & Observations
HumanEval (pass@1)93.1%90.4%89.2%Gemini tends to lock into incorrect assumptions prematurely on ambiguous prompts; performance is close on explicit algorithmic problems
SWE-bench Verified52.7%50.3%48.1%On real-world cross-file repository repairs, Gemini partially closes the gap leveraging its large context window
GPQA Diamond83.9%87.4%82.1%Claude leads in multi-step deep scientific derivations
MMLU Pro92.3%91.7%90.8%All three achieve exceptionally high knowledge breadth (minimal gap)
MATH94.8%94.1%94.6%All three are essentially on par in competition math (differences within error margins)

2. Custom Long-Form Writing and SVG Generation (Scores 1–10, 3-Judge Blind Review Average)

Test TaskGPT-5.4Claude Opus 4.6Gemini 3.1 ProKey Findings
5,000-Word Literary Writing7.88.67.3Gemini satisfies plot requirements but produces relatively mechanical prose; Claude is best at pacing and subtext
Strictly Constrained Marketing Copy8.28.07.5GPT-5.4 adheres most strictly to negative constraints; Gemini's tone is more generic
120k-Token Research Report SynthesisGood (misses some deep subtle connections)Excellent (best cross-document integration)Solid (complete information retrieval, but summaries lean generic)Gemini handles million-scale single-pass throughput effortlessly; Claude excels in nuance within 200k tokens
Complex SVG & Spatial LayoutExcellent (best layering and z-index)Good (best at animations and flowcharts)Fair (prone to adding redundant viewBox elements requiring manual cleanup)Gemini exhibits slight shortcomings in complex spatial layout SVG code generation

Conclusion

MindStudio provides clear guidelines for model selection:

  1. Default first choice for coding and engineering execution: GPT-5.4 (highest accuracy, fastest generation speed).

  2. High-quality long-form writing and extreme scientific reasoning: Claude Opus 4.6 (leads in prose quality and GPQA).

  3. Ultra-long document retrieval, full codebase reading, and production-grade high-throughput cost-sensitive tasks: Gemini 3.1 Pro (2M-token context, output cost is only 37.5% of Claude's, offering the highest engineering cost-performance).

Limitations

  • The empirical review was published in mid-March 2026; subsequent fine-tuning updates to each model may slightly alter scores.

  • Subjective scoring is constrained by the preference distribution of 3 evaluators; although blind review was employed, literary evaluation is inherently subjective.

Reproduction steps

  1. Prepare the standard test environments for HumanEval and SWE-bench Verified.

  2. Construct a 5,000-word creative writing brief and a 120k-token multi-document dataset, pinning the API endpoints for all three models.

  3. Record pass@1 rates, operational costs, and actual throughput TPS, and organize three evaluator groups to conduct normalized blind scoring.

What this supports

  • It supports relative observations about long documents, cost, coding, and writing

What this does not support

  • It supports relative observations about long documents, cost, coding, and writing, not a Gemini API SLA or quality for every repository.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

MindStudio Blog · Luis Chavez-Mattos (Director of Product, MindStudio) · Original publication date 2026-03-15 · Site edit date 2026-09-20

Open original source

Gemini 3.1 Pro

Compare Gemini 3.1 Pro in Tabbit

Download the Tabbit client to check model access

Related reviews

LayerLens Stratix's Six-Benchmark Evaluation of Gemini 3.1 Pro PreviewLayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.Artificial Analysis's Comprehensive 182-Model Benchmark and End-to-End Latency Evaluation of Gemini 3.1 Pro PreviewArtificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.Google Officially Releases Gemini 3.1 Pro: ARC-AGI-2 and Product Positioning BaselineGoogle's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.Google AI Developers Forum: Empirical Instruction-Following Evaluation of Gemini 3.1 Pro Under a Complex 4,000-Word System PromptAn Antigravity Ultra user reports that Gemini 3.1 Pro High can skip planning, compress output, drift in long sessions, or refactor out of scope under a roughly 4,000-word engineering instruction set.Gemini 3.1 Pro: Concise Prompting and Long-Context Question PlacementGoogle's Gemini 3 guide recommends direct, concise prompts and placing the specific question after long context with a short anchoring phrase.Gemini 3.1 Pro Thinking Levels, Structured Outputs, and Tool ConfigurationGoogle's official documentation combines thinking_level, default temperature, tool calls, and JSON schema checks, while separating the customtools endpoint.Spec-Driven Coding Workflow: Claude-Led Planning and Gemini-Isolated ExecutionThe developer-forum case uses Claude for specification and audit, Gemini 3.1 Pro for isolated execution in fresh sessions, and a final audit for changes.Gemini 3.1 Pro: Open-Source Architecture Alignment and Multi-Model Pair Programming WorkflowThis Antigravity community case injects a mature open-source project's architecture into Gemini 3.1 Pro and uses a second model for cross-review and alignment.