TabbitBlog

Gemini 3.5 Flash: What It Is, What It Costs, and Whether It Still Fits

A sourced Gemini 3.5 Flash overview covering its 1M context, multimodal tools, $1.50/$9 API pricing, legacy status, evidence limits and migration choices.

In this article
  1. Key takeaways
  2. Gemini 3.5 Flash at a glance
  3. Where it sits in the Flash family
  4. Benchmarks with the harness attached
  5. What users report
  6. Price, data use and access are separate decisions
  7. Run a scenario self-check
  8. A practical browser boundary: Tabbit Browser
  9. Verdict
  10. Sources and questions
  11. Is Gemini 3.5 Flash still available?
  12. Does it accept one million tokens?
  13. Is $9 per million the total cost?
  14. Is it better than Gemini 3.6, 3.7 or 3.8?
  15. Is it safe for autonomous agents?
  16. Can I use it in Tabbit Browser?

Gemini 3.5 Flash is Google’s May 19, 2026 Flash release for multimodal input, coding and agentic workflows. It still makes sense when you need a stable 1M-token API route and broad tools, but Google now calls it a legacy Flash model. That label means “older stable route,” not “already shut down.”

The practical anchor is the gap between list price and accepted-work cost. Google’s paid Standard table lists $1.50 per million input tokens and $9 per million output tokens, including thinking. An independent Artificial Analysis run was estimated at $1,552, while one Reddit user’s saved-evals project found 3.5 weak on its own vision task. Those are different denominators. Start with the Gemini 3.5 model page, pin the exact model ID and client, and treat Tabbit as a separate availability question.

Key takeaways

  • The stable model ID is gemini-3.5-flash; Google’s current catalogue calls it a legacy Flash model and lists no shutdown date.

  • Google lists text, image, video, audio and PDF input, 1,048,576 input tokens, 65,536 output tokens and thinking.

  • Supported API capabilities include caching, code execution, computer use preview, file search, function calling, grounding, structured outputs and URL context; client/account exposure still matters.

  • Google’s model card reports 76.2% Terminal-Bench 2.1 and 55.1% SWE-Bench Pro with named evaluation harnesses.

  • Gemini API billing, free-tier data use, AI Studio, Antigravity, Enterprise, subscriptions and Tabbit are separate routes.

Gemini 3.5 Flash at a glance

The Gemini 3.5 prompt collection and review collection preserve model-specific source material. They do not prove that your Google account, provider or Tabbit selector exposes every capability.

QuestionGoogle’s checked answerBoundary to keep
Release and stateMay 19, 2026; stable/GA; now listed as legacyLegacy is not a shutdown date.
Model IDgemini-3.5-flashConfirm the ID in the request or selector.
Input/outputText, image, video, audio and PDF input; text outputA client can expose fewer modalities.
Token limits1,048,576 input; 65,536 outputA product or quota can lower the effective window.
ToolsCode execution, file search, function calling, Search/Maps grounding, structured output, URL context and computer use previewAPI capability does not prove account or browser access.
Paid Standard$1.50 input / $9 output / $0.15 cached input per millionBatch, free tier, grounding and subscriptions are separate terms.
ShutdownNo date announced on Google’s deprecation pageRecheck before committing a long migration.

Where it sits in the Flash family

Gemini 3.5 Flash is the baseline this cluster needs to preserve. The Gemini 3.6 overview owns the July 21 efficiency and verification story. The Gemini 3.7 overview owns the later workhorse and 3.8 comparison. The Gemini 3.8 overview owns the September long-horizon route. None of those pages retroactively changes what gemini-3.5-flash meant at launch.

Google’s model catalogue now describes 3.5 as a legacy Flash model, 3.6 and 3.7 as previous-generation stable routes, and 3.8 as its newer long-horizon Flash route. That vocabulary is useful for lifecycle awareness, but not sufficient for a migration decision. Pin date, model ID, thinking level and tools, then repeat the same fixture. The agentic reasoning guide explains how to keep that comparison honest.

Migration question3.5 FlashNeighboring routeWhat to measure
Need a stable baseline?May 2026 GA route3.6/3.7/3.8 are later stable routesSame prompt, effort and tool policy.
Need cheaper list pricing?$1.50/$9 StandardLater routes have their own dated tables and promotionsAccepted output tokens, not only unit rate.
Need newer long-horizon behavior?Strong tools, but legacy status3.7/3.8 claim newer agentic focusCompletion, retries, review time and failure recovery.
Need lifecycle certainty?No shutdown date announcedNewer does not mean permanent eitherStore model IDs and a fallback.

Benchmarks with the harness attached

Google’s model card reports 76.2% on Terminal-Bench 2.1 and 55.1% on SWE-Bench Pro. The methodology document says the Terminal-Bench result uses the Terminus-2 agent harness, while Gemini SWE-Bench results are self-computed, averaged over five runs and use an internal Antigravity harness. That makes the numbers useful directional evidence, not a promise about your repository.

The safety card says Gemini 3.5 Flash is natively multimodal and reasoning-capable, and Google reports no material new Frontier Safety capability compared with its Gemini 3.1 Pro reference. Google also lists known limitations such as hallucinations, occasional slowness or timeouts, ongoing jailbreak-resistance work and a March 2026 knowledge cutoff. A benchmark score does not remove those operating conditions.

Better Stack’s independent analysis reports an Artificial Analysis full-suite estimate of $1,552 for Gemini 3.5 Flash, compared with $282 for Gemini 3 Flash and approximately $870 for Gemini 3.1 Pro. Keep this in an independent cost section: it is the cost of one benchmark workload, not Google’s $9 output rate or a monthly subscription.

What users report

Community evidence is split by task. A Reddit developer ran about 10 saved production-selection evals and says 3.5 underperformed older Gemini variants on most tasks; a five-run vision emotion test placed it 13th in that personal setup. Another reviewer calls it fast but expensive and reports more hallucinations on complex problems, while preferring Gemini 3.1 Pro for deep coding. A different Reddit user says 3.5 Flash is underrated for spreadsheets and document work, with the caveat that the cited comparison used each model’s highest reasoning option.

Two YouTube pages add scenario context rather than proof. One creator built a benchmarking app around an OpenAI-compatible local API; another framed a coding test around speed, cost and capability. These videos are useful for designing a pilot, not for converting a local/provider result into a Google API claim.

Price, data use and access are separate decisions

RouteWhat it may provideWhat it does not prove
Gemini APIModel ID, paid/free tier and token billingA Gemini app subscription or Tabbit picker.
Google AI StudioDeveloper UI and API credentialsEnterprise controls or a fixed quota for every account.
Antigravity / Android StudioProduct-specific agent or development surfaceThe same API billing, tools or data terms.
Gemini EnterpriseOrganization contract, governance and supportConsumer or developer availability.
Tabbit BrowserA browser workspace that may expose a live choiceGoogle API credit, context limit or guaranteed 3.5 access.

Google’s checked paid Standard row is $1.50 input, $9 output including thinking, and $0.15 cached input per million tokens. Batch is listed separately at $0.75 input and $4.50 output. The free tier has a different data-use statement; Google’s pricing page says submitted content may be used to improve products on free access, while paid access says it is not. Confirm the current regional terms before sending sensitive material. Search grounding, Maps grounding, retries, storage and subscriptions are separate lines.

Run a scenario self-check

ScenarioFirst testAcceptance condition
Small coding patchFixed repository, read-only branch, testsDiff is in scope and tests distinguish a real fix from no change.
Document or spreadsheet workDated files with a hand-checked answer keyFields and calculations match; unknowns stay explicit.
Multimodal extractionFive images or PDFs with known fieldsEvery field points to its source; no invented value is accepted.
Browser/computer useDisposable page and no write permissionURLs, tool calls, retries and stop behavior are logged.
Long agent loopFixed budget, thinking level and checkpointHuman review occurs before external side effects.

Record the exact model ID, product, thinking level, input/output/cache tokens, tools, retries, wall time, corrections and acceptance result. If a later 3.6, 3.7 or 3.8 run wins, keep the 3.5 failure reason instead of rewriting it as a family-wide ranking.

A practical browser boundary: Tabbit Browser

When work begins with live pages, grouped tabs or local documents, Tabbit Browser is a separate browser layer. It does not provide Gemini API credits, change Google’s free/paid data policy or guarantee gemini-3.5-flash in the model picker. This draft did not run an authenticated Tabbit task and captured no qualified screenshots, so it makes no claim about availability, latency, quota or context. If the model appears in your selector, start with a public, reversible fixture and record what the live account actually exposes.

For browser workflow context, read the AI browser guide and browser automation guide. Keep the Gemini 3.5 reviews beside the task log.

Tabbit Browser

Verdict

Gemini 3.5 Flash remains a plausible stable baseline for multimodal extraction, fast coding loops and bounded tools when your account still exposes it. Its trade-off is clear: a high output rate and thinking-token bill, legacy lifecycle language, and task-sensitive quality. Move to 3.6, 3.7 or 3.8 only after the same fixture shows that fewer retries, better verification or newer capability repays the migration cost. Keep a pinned fallback and do not treat Tabbit availability as an API entitlement.

Sources and questions

Primary sources are Google’s I/O developer announcement, model page, pricing table, deprecation table and DeepMind model card. Independent context comes from Better Stack.

Is Gemini 3.5 Flash still available?

Google’s catalogue currently lists it as a legacy Flash model, while the deprecation table shows a May 19, 2026 release and no shutdown date announced. Recheck the table and your account before a long migration.

Does it accept one million tokens?

Google lists 1,048,576 input tokens and 65,536 output tokens. A client, account, tool call or product plan may expose less, so test the actual route.

Is $9 per million the total cost?

It is the paid Standard output rate, including thinking tokens. Input, cached input, batch, grounding, retries and subscriptions are separate; one benchmark suite can cost much more because it consumes many tokens.

Is it better than Gemini 3.6, 3.7 or 3.8?

No universal ranking follows from the version number. Compare the same task, client, thinking setting, tools and acceptance test on pinned IDs.

Is it safe for autonomous agents?

Use least-privilege tools, a disposable workspace, tests and a human checkpoint. Google lists hallucinations, timeouts and ongoing jailbreak-resistance work, and community reports show verification failures.

Can I use it in Tabbit Browser?

This draft did not verify a signed-in Tabbit account. Check the live selector and treat a successful task as an account-level observation, not a platform guarantee.

FAQ

What is Gemini 3.5 Flash?

Gemini 3.5 Flash is Google's stable, generally available Flash model released on May 19, 2026 for multimodal, coding and agentic work. Google now labels it a legacy Flash model, but its deprecation table has no shutdown date announced.

What are Gemini 3.5 Flash's limits?

Google lists text, image, video, audio and PDF input, a 1,048,576-token input limit and a 65,536-token output limit. Tools such as code execution, file search, function calling and computer use preview still depend on the client and account.

How much does Gemini 3.5 Flash cost?

Google's paid Standard table showed $1.50 per million input tokens, $9 per million output tokens including thinking, and $0.15 per million cached input tokens when checked on September 20, 2026. Batch pricing and free-tier data terms are separate.

Is Gemini 3.5 Flash good for coding agents?

Google reports 76.2% on Terminal-Bench 2.1 and 55.1% on SWE-Bench Pro under named harnesses. Community evals disagree by task, so use a fixed repository, tests, a budget and a human checkpoint.

Should I move to Gemini 3.6, 3.7 or 3.8 Flash?

Not automatically. Those are separate stable routes with their own release dates, prices and task evidence. Pin the model ID and compare the same fixture before migrating.

Can I use Gemini 3.5 Flash in Tabbit Browser?

This draft did not run an authenticated Tabbit Gemini 3.5 Flash task. Check the live selector and verify one small reversible task; Google API access does not prove Tabbit access.

Take the next step

Let Tabbit work alongside you.

Research across tabs, automate repetitive browser work, and keep every piece of context within reach.