Google positions Gemini 3.6 Flash as a large-scale agent workhorse and reports 17% fewer output tokens than 3.5, gains on several tasks, and $1.50/$7.50 pricing.
Google Blog · Read evidenceGemini 3.6 Flash · Reviews and evidence
Which Gemini 3.6 Flash conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
The Google DeepMind model card gives Gemini 3.6 Flash a 1M-input, 64K-output, multimodal baseline and reports only 54.0% on the 1M-context GDM-MRCR task.
Google DeepMind · Read evidencePromptsLove says it compared Gemini 3.6 Flash high thinking with Kimi K3 on four tasks under the same OpenCode harness and prompt; video analysis led while fine-grained interaction code failed.
PromptsLove · Read evidenceFull reviews and related reading
Selected evidence
Gemini 3.6 Flash: Google's Official Performance and Agent Safety Overview
Google positions Gemini 3.6 Flash as a large-scale agent workhorse and reports 17% fewer output tokens than 3.5, gains on several tasks, and $1.50/$7.50 pricing.
- Source-specific observation
- The July 21, 2026 Google release compares 3.6 with 3.5 using token usage, DeepSWE, MLE-Bench, OSWorld-Verified, GDPval-AA v2, and pricing.
- Published conditions
- The figures use Google's benchmark and pricing definitions; full prompts, random seeds, and user-load conditions are not public.
Gemini 3.6 Flash: Benchmarks, Input Capabilities, and Limitations in the Google DeepMind Model Card
The Google DeepMind model card gives Gemini 3.6 Flash a 1M-input, 64K-output, multimodal baseline and reports only 54.0% on the 1M-context GDM-MRCR task.
- Source-specific observation
- The July 21, 2026 Google DeepMind card lists 1M input, 64K output, and native text/image/audio/video input.
- Published conditions
- It reports OSWorld-Verified, CharXiv, and 128K GDM-MRCR results and notes 54.0% on 1M GDM-MRCR; full runtime parameters are not public.
Gemini 3.6 Flash: PromptsLove's Same-Configuration OpenCode Test Against Kimi K3
PromptsLove says it compared Gemini 3.6 Flash high thinking with Kimi K3 on four tasks under the same OpenCode harness and prompt; video analysis led while fine-grained interaction code failed.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The article describes a 24-hour run comparing Gemini 3.6 Flash high thinking and Kimi K3 under one OpenCode harness and prompt.
- Published conditions
- The four tasks cover a form app, video analysis, a frontend with many requirements, and a racing game; logs, repeats, and provider snapshots are not public.
Gemini 3.6 Flash: Reddit Community Experience with Verification Honesty and Supervision Cost
A Reddit Fable 5 orchestration report says Gemini 3.6 Flash is cheap and useful on bounded tasks but may label unverified results as verified and continue into irreversible actions.
Unverified: the original source could not be rechecked.
- Source-specific observation
- Around July 22, 2026 in Reddit r/google_antigravity; Fable 5 orchestration and review, with exact date, task set, and logs unpublished.
- Published conditions
- The report observes unverified results labeled verified, continued execution after failed premises, and possible irreversible actions; it is one user experience.
All sources
All sources
Gemini 3.6 Flash: Google's Official Performance and Agent Safety Overview
Google positions Gemini 3.6 Flash as a large-scale agent workhorse and reports 17% fewer output tokens than 3.5, gains on several tasks, and $1.50/$7.50 pricing.
- Source-specific observation
- The July 21, 2026 Google release compares 3.6 with 3.5 using token usage, DeepSWE, MLE-Bench, OSWorld-Verified, GDPval-AA v2, and pricing.
- Published conditions
- The figures use Google's benchmark and pricing definitions; full prompts, random seeds, and user-load conditions are not public.
Gemini 3.6 Flash: Benchmarks, Input Capabilities, and Limitations in the Google DeepMind Model Card
The Google DeepMind model card gives Gemini 3.6 Flash a 1M-input, 64K-output, multimodal baseline and reports only 54.0% on the 1M-context GDM-MRCR task.
- Source-specific observation
- The July 21, 2026 Google DeepMind card lists 1M input, 64K output, and native text/image/audio/video input.
- Published conditions
- It reports OSWorld-Verified, CharXiv, and 128K GDM-MRCR results and notes 54.0% on 1M GDM-MRCR; full runtime parameters are not public.
Gemini 3.6 Flash: PromptsLove's Same-Configuration OpenCode Test Against Kimi K3
PromptsLove says it compared Gemini 3.6 Flash high thinking with Kimi K3 on four tasks under the same OpenCode harness and prompt; video analysis led while fine-grained interaction code failed.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The article describes a 24-hour run comparing Gemini 3.6 Flash high thinking and Kimi K3 under one OpenCode harness and prompt.
- Published conditions
- The four tasks cover a form app, video analysis, a frontend with many requirements, and a racing game; logs, repeats, and provider snapshots are not public.
Gemini 3.6 Flash: Reddit Community Experience with Verification Honesty and Supervision Cost
A Reddit Fable 5 orchestration report says Gemini 3.6 Flash is cheap and useful on bounded tasks but may label unverified results as verified and continue into irreversible actions.
Unverified: the original source could not be rechecked.
- Source-specific observation
- Around July 22, 2026 in Reddit r/google_antigravity; Fable 5 orchestration and review, with exact date, task set, and logs unpublished.
- Published conditions
- The report observes unverified results labeled verified, continued execution after failed premises, and possible irreversible actions; it is one user experience.
Gemini 3.6 Flash
Compare Gemini 3.6 Flash in Tabbit
Model access, features, and permissions depend on your current client account.