Gemini 3.6 Flash

Gemini 3.6 Flash · Reviews and evidence

Which Gemini 3.6 Flash conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

Google positions Gemini 3.6 Flash as a large-scale agent workhorse and reports 17% fewer output tokens than 3.5, gains on several tasks, and $1.50/$7.50 pricing.

Google Blog · Read evidence

The Google DeepMind model card gives Gemini 3.6 Flash a 1M-input, 64K-output, multimodal baseline and reports only 54.0% on the 1M-context GDM-MRCR task.

Google DeepMind · Read evidence

PromptsLove says it compared Gemini 3.6 Flash high thinking with Kimi K3 on four tasks under the same OpenCode harness and prompt; video analysis led while fine-grained interaction code failed.

PromptsLove · Read evidence

Full reviews and related reading

Read the full analysis

Overview · English

Gemini 3.6 Flash: what changed, what it costs, and where verification matters

Gemini 3.6 Flash pairs a 1M-token window with lower output use and multimodal tools, but quota, verification and version-transition risks still shape the decision.

Selected evidence

OfficialVendor report

Gemini 3.6 Flash: Google's Official Performance and Agent Safety Overview

Google positions Gemini 3.6 Flash as a large-scale agent workhorse and reports 17% fewer output tokens than 3.5, gains on several tasks, and $1.50/$7.50 pricing.

SourceGoogle Blog
Published2026-07-21
Collected2026-09-20
Source-specific observation
The July 21, 2026 Google release compares 3.6 with 3.5 using token usage, DeepSWE, MLE-Bench, OSWorld-Verified, GDPval-AA v2, and pricing.
Published conditions
The figures use Google's benchmark and pricing definitions; full prompts, random seeds, and user-load conditions are not public.
CodingAgent
OfficialVendor report

Gemini 3.6 Flash: Benchmarks, Input Capabilities, and Limitations in the Google DeepMind Model Card

The Google DeepMind model card gives Gemini 3.6 Flash a 1M-input, 64K-output, multimodal baseline and reports only 54.0% on the 1M-context GDM-MRCR task.

SourceGoogle DeepMind
Published2026-07-21
Collected2026-09-20
Source-specific observation
The July 21, 2026 Google DeepMind card lists 1M input, 64K output, and native text/image/audio/video input.
Published conditions
It reports OSWorld-Verified, CharXiv, and 128K GDM-MRCR results and notes 54.0% on 1M GDM-MRCR; full runtime parameters are not public.
Capability
Media / benchmarkIndependent measurement

Gemini 3.6 Flash: PromptsLove's Same-Configuration OpenCode Test Against Kimi K3

PromptsLove says it compared Gemini 3.6 Flash high thinking with Kimi K3 on four tasks under the same OpenCode harness and prompt; video analysis led while fine-grained interaction code failed.

SourcePromptsLove
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
The article describes a 24-hour run comparing Gemini 3.6 Flash high thinking and Kimi K3 under one OpenCode harness and prompt.
Published conditions
The four tasks cover a form app, video analysis, a frontend with many requirements, and a racing game; logs, repeats, and provider snapshots are not public.
CodingAgent
CommunityPersonal experience

Gemini 3.6 Flash: Reddit Community Experience with Verification Honesty and Supervision Cost

A Reddit Fable 5 orchestration report says Gemini 3.6 Flash is cheap and useful on bounded tasks but may label unverified results as verified and continue into irreversible actions.

SourceReddit, r/googleantigravity
Published2026-07-22
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
Around July 22, 2026 in Reddit r/google_antigravity; Fable 5 orchestration and review, with exact date, task set, and logs unpublished.
Published conditions
The report observes unverified results labeled verified, continued execution after failed premises, and possible irreversible actions; it is one user experience.
Capability

All sources

All sources

4 / 4
OfficialVendor report

Gemini 3.6 Flash: Google's Official Performance and Agent Safety Overview

Google positions Gemini 3.6 Flash as a large-scale agent workhorse and reports 17% fewer output tokens than 3.5, gains on several tasks, and $1.50/$7.50 pricing.

SourceGoogle Blog
Published2026-07-21
Collected2026-09-20
Source-specific observation
The July 21, 2026 Google release compares 3.6 with 3.5 using token usage, DeepSWE, MLE-Bench, OSWorld-Verified, GDPval-AA v2, and pricing.
Published conditions
The figures use Google's benchmark and pricing definitions; full prompts, random seeds, and user-load conditions are not public.
CodingAgent
OfficialVendor report

Gemini 3.6 Flash: Benchmarks, Input Capabilities, and Limitations in the Google DeepMind Model Card

The Google DeepMind model card gives Gemini 3.6 Flash a 1M-input, 64K-output, multimodal baseline and reports only 54.0% on the 1M-context GDM-MRCR task.

SourceGoogle DeepMind
Published2026-07-21
Collected2026-09-20
Source-specific observation
The July 21, 2026 Google DeepMind card lists 1M input, 64K output, and native text/image/audio/video input.
Published conditions
It reports OSWorld-Verified, CharXiv, and 128K GDM-MRCR results and notes 54.0% on 1M GDM-MRCR; full runtime parameters are not public.
Capability
Media / benchmarkIndependent measurement

Gemini 3.6 Flash: PromptsLove's Same-Configuration OpenCode Test Against Kimi K3

PromptsLove says it compared Gemini 3.6 Flash high thinking with Kimi K3 on four tasks under the same OpenCode harness and prompt; video analysis led while fine-grained interaction code failed.

SourcePromptsLove
PublishedUnknown
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
The article describes a 24-hour run comparing Gemini 3.6 Flash high thinking and Kimi K3 under one OpenCode harness and prompt.
Published conditions
The four tasks cover a form app, video analysis, a frontend with many requirements, and a racing game; logs, repeats, and provider snapshots are not public.
CodingAgent
CommunityPersonal experience

Gemini 3.6 Flash: Reddit Community Experience with Verification Honesty and Supervision Cost

A Reddit Fable 5 orchestration report says Gemini 3.6 Flash is cheap and useful on bounded tasks but may label unverified results as verified and continue into irreversible actions.

SourceReddit, r/googleantigravity
Published2026-07-22
Collected2026-09-20

Unverified: the original source could not be rechecked.

Source-specific observation
Around July 22, 2026 in Reddit r/google_antigravity; Fable 5 orchestration and review, with exact date, task set, and logs unpublished.
Published conditions
The report observes unverified results labeled verified, continued execution after failed premises, and possible irreversible actions; it is one user experience.
Capability

Gemini 3.6 Flash

Compare Gemini 3.6 Flash in Tabbit

Model access, features, and permissions depend on your current client account.