Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Gemini 3.6 Flash · Official source · Vendor report

Gemini 3.6 Flash: Benchmarks, Input Capabilities, and Limitations in the Google DeepMind Model Card

The Google DeepMind model card gives Gemini 3.6 Flash a 1M-input, 64K-output, multimodal baseline and reports only 54.0% on the 1M-context GDM-MRCR task.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Source-specific observation
The July 21, 2026 Google DeepMind card lists 1M input, 64K output, and native text/image/audio/video input.
Published conditions
It reports OSWorld-Verified, CharXiv, and 128K GDM-MRCR results and notes 54.0% on 1M GDM-MRCR; full runtime parameters are not public.

Key data and applicable tasks

One-sentence takeaway

The model card provides a verifiable version baseline for Gemini 3.6 Flash: 1M-token input, 64K-token output, and native text/image/audio/video input. It has an advantage on OSWorld-Verified, CharXiv, and 128K GDM-MRCR, but SWE-Bench Pro, DeepSWE, and GDPval remain materially dependent on task type and competitors, while GDM-MRCR at the 1M context length is only 54.0%.

Use cases

  • Suitable tasks: Multimodal file understanding, computer use, long-document retrieval, chart synthesis, and code/Agent loops of moderate complexity.

  • Unsuitable tasks: Assuming that a 1M context means stable high recall at the 1M length, running high-risk tools without supervision, or treating the official benchmark comparison table as a fair competition under the same harness.

  • Applicable model version: Gemini 3.6 Flash, with the model card published in July 2026.

  • Applicable client, Agent, or API: Gemini API, Google AI Studio, and the Agent tool paths described in the model card.

  • Recommended reasoning level and parameters: The model card does not disclose the complete temperature/prompt configuration; thinking configuration should be checked against the official Gemini 3 developer guide.

Test environment

The model card says the evaluation covers reasoning, coding, agentic, multimodal, and long-context capabilities, and presents the results alongside Gemini 3.5 Flash, Gemini 3.1 Pro, GPT-5.6 Luna, Grok 4.5, and Claude Sonnet 5. The link for the detailed methodology points to DeepMind's evals methodology page.

Input/configuration

  • Input: Text, images, audio, and video; context window up to 1M.

  • Output: Text, up to 64K tokens.

  • Pricing (model card table): $1.50/1M for input and $7.50/1M for output, with no cached-input price.

  • Model dependency: Based on Gemini 3.5 Flash; knowledge cutoff is March 2026.

  • Tools/tasks: The table includes different types of setups, including the Terminus-2 harness for Terminal-bench, OSWorld-Verified computer use, and the long-context GDM-MRCR.

Result data

BenchmarkGemini 3.6 FlashGemini 3.5 FlashObservation
SWE-Bench Pro58.7%55.1%An improvement, but below GPT-5.6 Luna at 62.7%, Grok 4.5 at 64.7%, and Claude Sonnet 5 at 63.2% in the table
DeepSWE v1.149%37%A clear improvement in long-horizon software engineering
Terminal-Bench 2.178.0%76.2%Uses the Terminus-2 harness
MLE-Bench63.9%49.7%Improvement in machine learning engineering
GDPval-AA v214211349Knowledge-work Elo
OSWorld-Verified83.0%78.4%Gemini 3.6 Flash is the highest in the table
CharXiv (no tools)85.2%84.2%Chart reasoning
CharXiv (with tools)89.4%84.9%Higher with tools
GDM-MRCR v2 (128K)91.8%77.3%Long-context improvement
GDM-MRCR v2 (1M pointwise)54.0%26.6%Significantly below 128K at the 1M length

Conclusion

The model card breaks down “strong at Agent and multimodal tasks” into concrete boundaries: performance is strong on OSWorld, CharXiv, and 128K retrieval, while SWE/DeepSWE is competitive but not an across-the-board leader. The advertised 1M capacity cannot replace recall regression at real context lengths. In production routing, it can be prioritized for document, chart, video, and computer-use subtasks, with an escalation path retained for more difficult code.

Limitations

  • The harnesses, tools, number of repetitions, and scoring methods for the benchmarks in the table are not fully identical, so they cannot be simply averaged across benchmarks.

  • The model card results are evaluations published by Google; independent third-party reproduction has not yet been reported.

  • The model card may be updated as the model improves, and the knowledge cutoff date creates differences in timeliness.

  • Known limitations include hallucination, occasional slowness/timeout, the knowledge cutoff, and erroneous refusals caused by safety policies; high-risk tasks require human review.

Reproduction steps

  1. Record the model card version, publication date, knowledge cutoff, and complete table; keep the 128K and 1M GDM-MRCR results separate.

  2. Prepare retrieval sets at 128K, 300K, 600K, and close to 1M using the same model ID and API version, and calculate needle recall, citation accuracy, latency, and cost.

  3. Run SWE/Agent, OSWorld-style, and multimodal document tasks with the same tool schema, recording complete traces and failure reasons.

  4. Label all official figures as “official results” and do not raise the evidence level until they have been reproduced with an independent harness.

Source excerpt or observation (for compliant short quotation only)

The model card explicitly lists two long-context readings—“1M (pointwise) 54.0%” and 128K 91.8%—so both boundaries must be shown when promoting the context window.

What this supports

  • It supports context-size and task-boundary analysis

What this does not support

  • It supports context-size and task-boundary analysis, not success rates for every agent or repository.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Google DeepMind · Google DeepMind · Original publication date 2026-07-21 · Site edit date 2026-09-20

Open original source

Gemini 3.6 Flash

Compare Gemini 3.6 Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Gemini 3.6 Flash: what changed, what it costs, and where verification matters

Gemini 3.6 Flash pairs a 1M-token window with lower output use and multimodal tools, but quota, verification and version-transition risks still shape the decision.

Related reviews

Gemini 3.6 Flash: Reddit Community Experience with Verification Honesty and Supervision CostA Reddit Fable 5 orchestration report says Gemini 3.6 Flash is cheap and useful on bounded tasks but may label unverified results as verified and continue into irreversible actions.Gemini 3.6 Flash: Google's Official Performance and Agent Safety OverviewGoogle positions Gemini 3.6 Flash as a large-scale agent workhorse and reports 17% fewer output tokens than 3.5, gains on several tasks, and $1.50/$7.50 pricing.Gemini 3.6 Flash: PromptsLove's Same-Configuration OpenCode Test Against Kimi K3PromptsLove says it compared Gemini 3.6 Flash high thinking with Kimi K3 on four tasks under the same OpenCode harness and prompt; video analysis led while fine-grained interaction code failed.Gemini 3.6 Flash: Google Official API Capabilities and Thinking Configuration ChecklistGoogle's model page lists Gemini 3.6 Flash's stable ID, modalities, context, caching, tools, and thinking settings as an integration baseline.Gemini 3.6 Flash: PromptsRush's Long-Context, Multimodal, and Agent PromptsPromptsRush turns Gemini 3.6 Flash long-context and multimodal agent work into task contracts with citations, schemas, and verification steps.