Gemini 3.1 Pro Thinking Levels, Structured Outputs, and Tool Configuration
Google's official documentation combines thinking_level, default temperature, tool calls, and JSON schema checks, while separating the customtools endpoint.
Prepare
task goal, source or reference material, runtime constraints, acceptance criteria
Runtime
Gemini 3.1 Pro client or API; confirm the live model ID, tools, permissions, and version before execution.
Spec-Driven Coding Workflow: Claude-Led Planning and Gemini-Isolated Execution
The developer-forum case uses Claude for specification and audit, Gemini 3.1 Pro for isolated execution in fresh sessions, and a final audit for changes.
Prepare
task goal, source or reference material, runtime constraints, acceptance criteria
Runtime
Gemini 3.1 Pro client or API; confirm the live model ID, tools, permissions, and version before execution.
This Antigravity community case injects a mature open-source project's architecture into Gemini 3.1 Pro and uses a second model for cross-review and alignment.
Prepare
task goal, source or reference material, runtime constraints, acceptance criteria
Runtime
Gemini 3.1 Pro client or API; confirm the live model ID, tools, permissions, and version before execution.
Google Officially Releases Gemini 3.1 Pro: ARC-AGI-2 and Product Positioning Baseline
Google's 2026-02-19 release uses the Gemini 3.1 Pro preview and a verified ARC-AGI-2 score of 77.1% as a product baseline, without publishing the full ARC harness.
Evidence
Vendor report
Boundary
It supports ARC-AGI-2 as an official positioning signal, not score reproduction, universal task win rates, or a stable preview SLA.
LayerLens Stratix's Six-Benchmark Evaluation of Gemini 3.1 Pro Preview
LayerLens Stratix covers 14,549 cases across six benchmarks and shows large task differences for Gemini 3.1 Pro between ARC and BIRD-CRITIC, among others.
Evidence
Independent measurement
Boundary
It supports routing by abstract reasoning, software engineering, SQL, and function calling; it does not establish a current API snapshot or universal ranking.
Artificial Analysis's Comprehensive 182-Model Benchmark and End-to-End Latency Evaluation of Gemini 3.1 Pro Preview
Artificial Analysis compares 182 similarly priced models on first-party APIs and reports Gemini 3.1 Pro Preview at Intelligence Index 48, 121.4 t/s, and 32.45 seconds TTFT, combining high throughput with high startup latency.
Evidence
Independent measurement
Boundary
It supports like-for-like intelligence, generation-rate, TTFT, and cost comparison, not a repository success rate or fixed price.