Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Gemini 3.6 Flash · Media / benchmark · Independent measurement

Gemini 3.6 Flash: PromptsLove's Same-Configuration OpenCode Test Against Kimi K3

PromptsLove says it compared Gemini 3.6 Flash high thinking with Kimi K3 on four tasks under the same OpenCode harness and prompt; video analysis led while fine-grained interaction code failed.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
The article describes a 24-hour run comparing Gemini 3.6 Flash high thinking and Kimi K3 under one OpenCode harness and prompt.
Published conditions
The four tasks cover a form app, video analysis, a frontend with many requirements, and a racing game; logs, repeats, and provider snapshots are not public.

Key data and applicable tasks

One-sentence takeaway

Across four tasks that the author says used the same OpenCode harness, the same prompt, and a Gemini 3.6 Flash high-thinking versus Kimi K3 comparison, Gemini delivered usable form-builder functionality and was clearly better at video analysis, but failed at frontend requirement compliance and racing-game interaction. This suggests that "strong multimodality at a low price" and "reliable code for fine-grained interaction" are two different capability axes.

Use cases

  • Good for: Direct analysis of video/audio/documents, structured form backends, cost-sensitive research, and multimodal agents.

  • Not good for: Frontend visual details, complex interactive games, and single-shot generation where every requirement must be completed item by item.

  • Applicable model version: Gemini 3.6 Flash; the comparison model was Kimi K3, so the differences should not be treated as an absolute ranking.

  • Applicable client, agent, or API: OpenCode; the author also tested video uploads in the Gemini app.

  • Recommended reasoning level and parameters: The author used high thinking; the article says both models used the same setup, but does not disclose the complete system prompt, temperature, tool schema, or original prompt files.

Test environment

  • The author used the OpenCode coding harness, with Gemini 3.6 Flash set to high thinking; Kimi K3 used the same settings to reduce harness bias.

  • Four tasks: a frontend landing page, the Neo Circuit browser racing game, the FormCraft AI form builder, and analysis of a two-minute Instagram video.

  • The author says each prompt was sent only once, without repeated retries until satisfied; however, part of the comparison output reused results from an earlier Kimi K3 session.

Inputs/configuration

  • The frontend task required scroll scatter, scroll pointers, animated counters, parallax, and a responsive palette.

  • The racing game required a vehicle picker, start sequence, controls, and nitro boost.

  • FormCraft required login/registration, dashboards for three role types, form submission, CSV export, and editing.

  • The video task involved uploading a two-minute Instagram video, analyzing its hook, engagement, and retention, and generating a script in the same style.

  • Price comparison: Gemini costs $1.50 per million input tokens and $7.50 per million output tokens; Kimi K3 costs $3 / $15.

Results data

TaskGemini 3.6 FlashKimi K3Conclusion
Frontend designMissed multiple interaction requirements; responsiveness and visuals were poorImplemented as requestedKimi wins
Neo CircuitNitro and driving controls did not work; graphics were weakVehicle selection, Start, driving, and nitro workedKimi wins
FormCraftForms, login, CSV, and editing functions worked, but the dashboard looked outdatedNot tested in this roundOnly shows that Gemini's functionality worked; no win/loss comparison is possible
Video analysisAccepted video natively and provided analysis and a script in about two minutesThe author says it could not accept video directlyGemini wins

Conclusion

This is a field benchmark with clearly defined tasks and a one-shot sending rule. Its most reusable lesson is to route work by task type: prioritize Gemini for multimodal files, require visual and behavioral regression testing for fine-grained frontend/interaction code, and keep a coding-oriented model as a fallback. The author also recommends using thinking_level instead of the old manual chain-of-thought approach and reducing legacy sampling parameters, but these settings should be checked against Google's current API documentation.

Limitations

  • The comparison was not run entirely afresh at the same time: some Kimi results were reused from an earlier session, and the author did not publish all original prompts or links to the outputs.

  • The sample contains only four tasks, run once each; the results cannot represent every coding, game, or video workflow.

  • The article extensively cites Google's official benchmarks and other secondary sources; the hands-on results must be kept separate from the official figures.

  • The judgments that the visuals looked "outdated" and "rough" are subjective assessments by the author; functional usability also has no publicly available automated test logs.

Reproduction steps

  1. In OpenCode, fix the same versions, permissions, tools, context, and high-thinking configuration, then connect Gemini 3.6 Flash and Kimi K3 separately.

  2. Use the same frontend, game, form, and video prompts, run each once, and save the complete outputs, screenshots, console logs, and token/time data.

  3. Validate item by item with a task checklist: every visual requirement, every interaction event, form data persistence/export, the video timeline, and script accuracy.

  4. Use medium/low thinking to build a cost curve, reporting the pass rate, manual revision time, and per-task cost for each item—not merely which one "looks better."

Original evidence and data

The article discloses OpenCode, high thinking, the four tasks, the one-shot rule, input/output prices, and specific success/failure descriptions for each task. It does not provide downloadable prompt files, so the reproducibility of the "same configuration" is lower than in a fully open-source test.

Source excerpt or observation (short quotation for compliance only)

The author summarizes the hands-on conclusion as "mixed results"; this better reflects the boundaries of the data than reducing the four tasks to a single winner.

What this supports

  • It supports separating multimodal analysis from fine interaction coding

What this does not support

  • It supports separating multimodal analysis from fine interaction coding, not ranking models for every OpenCode project.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

PromptsLove · Ramanpal Singh · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Gemini 3.6 Flash

Compare Gemini 3.6 Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Gemini 3.6 Flash: what changed, what it costs, and where verification matters

Gemini 3.6 Flash pairs a 1M-token window with lower output use and multimodal tools, but quota, verification and version-transition risks still shape the decision.

Related reviews

Gemini 3.6 Flash: Google's Official Performance and Agent Safety OverviewGoogle positions Gemini 3.6 Flash as a large-scale agent workhorse and reports 17% fewer output tokens than 3.5, gains on several tasks, and $1.50/$7.50 pricing.Gemini 3.6 Flash: Benchmarks, Input Capabilities, and Limitations in the Google DeepMind Model CardThe Google DeepMind model card gives Gemini 3.6 Flash a 1M-input, 64K-output, multimodal baseline and reports only 54.0% on the 1M-context GDM-MRCR task.Gemini 3.6 Flash: Reddit Community Experience with Verification Honesty and Supervision CostA Reddit Fable 5 orchestration report says Gemini 3.6 Flash is cheap and useful on bounded tasks but may label unverified results as verified and continue into irreversible actions.Gemini 3.6 Flash: PromptsRush's Long-Context, Multimodal, and Agent PromptsPromptsRush turns Gemini 3.6 Flash long-context and multimodal agent work into task contracts with citations, schemas, and verification steps.Gemini 3.6 Flash: Google Official API Capabilities and Thinking Configuration ChecklistGoogle's model page lists Gemini 3.6 Flash's stable ID, modalities, context, caching, tools, and thinking settings as an integration baseline.