Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialGemini 3.6 Flash

Gemini 3.6 Flash: Benchmarks, Input Capabilities, and Limitations in the Google DeepMind Model Card

Original source

Google DeepMind

AuthorGoogle DeepMind

Source date2026-07-21

Tabbit curation2026-08-19

Read original

One-sentence takeaway

The model card provides a verifiable version baseline for Gemini 3.6 Flash: 1M-token input, 64K-token output, and native text/image/audio/video input. It has an advantage on OSWorld-Verified, CharXiv, and 128K GDM-MRCR, but SWE-Bench Pro, DeepSWE, and GDPval remain materially dependent on task type and competitors, while GDM-MRCR at the 1M context length is only 54.0%.

Use cases

  • Suitable tasks: Multimodal file understanding, computer use, long-document retrieval, chart synthesis, and code/Agent loops of moderate complexity.

  • Unsuitable tasks: Assuming that a 1M context means stable high recall at the 1M length, running high-risk tools without supervision, or treating the official benchmark comparison table as a fair competition under the same harness.

  • Applicable model version: Gemini 3.6 Flash, with the model card published in July 2026.

  • Applicable client, Agent, or API: Gemini API, Google AI Studio, and the Agent tool paths described in the model card.

  • Recommended reasoning level and parameters: The model card does not disclose the complete temperature/prompt configuration; thinking configuration should be checked against the official Gemini 3 developer guide.

Test environment

The model card says the evaluation covers reasoning, coding, agentic, multimodal, and long-context capabilities, and presents the results alongside Gemini 3.5 Flash, Gemini 3.1 Pro, GPT-5.6 Luna, Grok 4.5, and Claude Sonnet 5. The link for the detailed methodology points to DeepMind's evals methodology page.

Input/configuration

  • Input: Text, images, audio, and video; context window up to 1M.

  • Output: Text, up to 64K tokens.

  • Pricing (model card table): $1.50/1M for input and $7.50/1M for output, with no cached-input price.

  • Model dependency: Based on Gemini 3.5 Flash; knowledge cutoff is March 2026.

  • Tools/tasks: The table includes different types of setups, including the Terminus-2 harness for Terminal-bench, OSWorld-Verified computer use, and the long-context GDM-MRCR.

Result data

BenchmarkGemini 3.6 FlashGemini 3.5 FlashObservation
SWE-Bench Pro58.7%55.1%An improvement, but below GPT-5.6 Luna at 62.7%, Grok 4.5 at 64.7%, and Claude Sonnet 5 at 63.2% in the table
DeepSWE v1.149%37%A clear improvement in long-horizon software engineering
Terminal-Bench 2.178.0%76.2%Uses the Terminus-2 harness
MLE-Bench63.9%49.7%Improvement in machine learning engineering
GDPval-AA v214211349Knowledge-work Elo
OSWorld-Verified83.0%78.4%Gemini 3.6 Flash is the highest in the table
CharXiv (no tools)85.2%84.2%Chart reasoning
CharXiv (with tools)89.4%84.9%Higher with tools
GDM-MRCR v2 (128K)91.8%77.3%Long-context improvement
GDM-MRCR v2 (1M pointwise)54.0%26.6%Significantly below 128K at the 1M length

Conclusion

The model card breaks down “strong at Agent and multimodal tasks” into concrete boundaries: performance is strong on OSWorld, CharXiv, and 128K retrieval, while SWE/DeepSWE is competitive but not an across-the-board leader. The advertised 1M capacity cannot replace recall regression at real context lengths. In production routing, it can be prioritized for document, chart, video, and computer-use subtasks, with an escalation path retained for more difficult code.

Limitations

  • The harnesses, tools, number of repetitions, and scoring methods for the benchmarks in the table are not fully identical, so they cannot be simply averaged across benchmarks.

  • The model card results are evaluations published by Google; independent third-party reproduction has not yet been reported.

  • The model card may be updated as the model improves, and the knowledge cutoff date creates differences in timeliness.

  • Known limitations include hallucination, occasional slowness/timeout, the knowledge cutoff, and erroneous refusals caused by safety policies; high-risk tasks require human review.

Reproduction steps

  1. Record the model card version, publication date, knowledge cutoff, and complete table; keep the 128K and 1M GDM-MRCR results separate.

  2. Prepare retrieval sets at 128K, 300K, 600K, and close to 1M using the same model ID and API version, and calculate needle recall, citation accuracy, latency, and cost.

  3. Run SWE/Agent, OSWorld-style, and multimodal document tasks with the same tool schema, recording complete traces and failure reasons.

  4. Label all official figures as “official results” and do not raise the evidence level until they have been reproduced with an independent harness.

Source excerpt or observation (for compliant short quotation only)

The model card explicitly lists two long-context readings—“1M (pointwise) 54.0%” and 128K 91.8%—so both boundaries must be shown when promoting the context window.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Gemini 3.6 Flash

Use and compare models in Tabbit

Gemini 3.6 Flash

Related reviews

OfficialGoogle Blog2026-07-21

Gemini 3.6 Flash: Google's Official Performance and Agent Safety Overview

MediaPromptsLove

Gemini 3.6 Flash: PromptsLove's Same-Configuration OpenCode Test Against Kimi K3

CommunityReddit, r/googleantigravity2026-07-22

Gemini 3.6 Flash: Reddit Community Experience with Verification Honesty and Supervision Cost

Gemini 3.6 Flash

Related prompts

OfficialGoogle AI for Developers2026-07-30

Gemini 3.6 Flash: Google Official API Capabilities and Thinking Configuration Checklist

MediaPromptsRush2026-07-25

Gemini 3.6 Flash: PromptsRush's Long-Context, Multimodal, and Agent Prompts