Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Gemini 3.5 Flash · Community source · Personal experience

Gemini 3.5 Flash: A Community Field Report on Ten Saved Tasks and Five Repeated Runs

A Reddit user repeated Gemini 3.5 Flash five times on about ten saved tasks and reported a lower real-task average than an older version; this is a personal field report, not a controlled benchmark.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Source-specific observation
Collected 2026-08-18; the author describes about ten saved tasks with five runs each but does not publish the full tasks, prompts, snapshot, or scoring script.
Published conditions
The runtime is the author's GeminiAI workflow; repeat count and model version are self-reported and not independently rerun.

Key data and applicable tasks

One-sentence takeaway

One user ran Gemini 3.5 Flash an average of five times on roughly 10 saved evaluations of their own, and reported that it often performed below older Gemini versions on real-world tasks. This is a reminder to retest with your own task set before migrating instead of looking only at official leaderboards.

Use cases

  • Suitable tasks: A pre-migration risk signal and a reference for local A/B evaluation of visual and Agent tasks.

  • Unsuitable tasks: Extrapolating “13th place” into a general ranking, or using the post as a substitute for auditable original benchmarks.

  • Applicable model version: The Gemini 3.5 Flash tested by the author; the specific API version and thinking configuration were not disclosed.

  • Applicable client, Agent, or API: The author used their own OpenMark benchmarking tool; OpenMark’s public service description supports custom tasks and comparisons across multiple models.

  • Recommended reasoning level and parameters: Not disclosed/unverifiable.

Test environment

  • Task set: Roughly 10 evaluations saved by the author; the post emphasizes one visual test.

  • Number of runs: The post describes the visual evaluation result as the average of 5 runs, and says similar results appeared across roughly 10 benchmarks.

  • Comparison: Gemini 3.1 Pro, Gemini 3.1 Flash Lite, Gemini 3 Flash, and Gemini 3.5 Flash.

  • Tool: OpenMark; the post did not disclose the task YAML, grader, complete responses, or cost logs.

Input/configuration

  • The task inputs, system prompt, images, model routing, temperature, token limit, and scoring rules were not disclosed/unverifiable.

  • The author emphasized that these were their own tasks and cautioned that Gemini results may depend heavily on prompt shape.

Results data

  • The post says Gemini 3.5 Flash ranked 13th in one visual evaluation, while Gemini 3.1 Pro and Gemini 3.1 Flash Lite ranked first and second, respectively, and Gemini 3 Flash also ranked above it.

  • The post says this visual result was an average across 5 runs, and reports a similar trend of older versions performing better across roughly 10 saved evaluations.

  • The post provides no raw score for each model, sample count, confidence interval, task text, or complete leaderboard, so it can only serve as a directional field report.

Conclusion

This report does not contradict official or specialized benchmarks claiming that “3.5 Flash has advantages in speed and tool tasks”: it measures a single user’s real-world task set, which notably includes visual tasks. The actionable conclusion is to fix the same prompt, inputs, and grader and run multiple A/B tests before migrating, then confirm whether the model is actually better in your own workflow.

Limitations

  • A single user, an undisclosed task set, and an undisclosed grader make independent reproduction impossible.

  • “13th place” has no contextual sample, score, or statistical significance; it cannot be treated as a global model ranking.

  • The author says “roughly 10 evaluations” but provides no list; the share and difficulty of visual tasks are unknown.

  • OpenMark’s public homepage mainly describes the platform’s terms and features; this collection did not find the user’s specific evaluation data.

Reproduction steps

  1. Export 10 sets of your own real-world tasks, fixing the text, images, tools, model versions, routing, and grader.

  2. Compare at least Gemini 3.5 Flash, Gemini 3.1 Pro, Gemini 3.1 Flash Lite, and Gemini 3 Flash, running each 5 times.

  3. Save each output, success criteria, score, token usage, latency, and failure reason; report the mean, standard deviation, and results by task segment.

  4. Report visual, tool, coding, and knowledge tasks separately to prevent one extreme visual task from overwhelming the overall mean.

Source excerpt or observation (compliance short quote only)

The author’s core reminder is “benchmark first rather than assume newer = better.” This is personal experience rather than a quantitative conclusion, but it is well suited as an acceptance gate for a version upgrade.

What this supports

  • It supports running a local migration test

What this does not support

  • It supports running a local migration test, not turning one user average into a general ranking or regression claim.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit / r/GeminiAI · u/igoontotomboys · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Gemini 3.5 Flash

Compare Gemini 3.5 Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Gemini 3.5 Flash: What It Is, What It Costs, and Whether It Still Fits

A sourced Gemini 3.5 Flash overview covering its 1M context, multimodal tools, $1.50/$9 API pricing, legacy status, evidence limits and migration choices.

Related reviews

Gemini 3.5 Flash: Google's Official Follow-up Release Comparison of Efficiency and CapabilitiesGoogle's follow-up release uses Gemini 3.5 Flash as the baseline for Gemini 3.6 Flash; differences on DeepSWE, MLE Bench, OSWorld-Verified, and GDPval-AA v2 are official-harness results.Gemini 3.5 Flash: Appwrite Arena Comparison of Skill Context and Agent TasksAppwrite Arena's May 20, 2026 run reports freeform rising from 77.5% to 91.9% after loading the Appwrite Skill, showing that documentation context changes agent results.Gemini 3.5 Flash: Structured Prompting, Grounding, and Agent System InstructionsGoogle recommends structuring Gemini 3.5 Flash prompts around the goal, context, task boundaries, output format, and grounding tools, then iterating on representative samples.Gemini 3.5 Flash: Thinking Levels and Gemini API ConfigurationThe official Thinking guide documents Gemini 3.5 Flash effort levels, preservation of parts and thought signatures across turns, and API parameter boundaries.