One user ran Gemini 3.5 Flash an average of five times on roughly 10 saved evaluations of their own, and reported that it often performed below older Gemini versions on real-world tasks. This is a reminder to retest with your own task set before migrating instead of looking only at official leaderboards.
Suitable tasks: A pre-migration risk signal and a reference for local A/B evaluation of visual and Agent tasks.
Unsuitable tasks: Extrapolating “13th place” into a general ranking, or using the post as a substitute for auditable original benchmarks.
Applicable model version: The Gemini 3.5 Flash tested by the author; the specific API version and thinking configuration were not disclosed.
Applicable client, Agent, or API: The author used their own OpenMark benchmarking tool; OpenMark’s public service description supports custom tasks and comparisons across multiple models.
Recommended reasoning level and parameters: Not disclosed/unverifiable.
Task set: Roughly 10 evaluations saved by the author; the post emphasizes one visual test.
Number of runs: The post describes the visual evaluation result as the average of 5 runs, and says similar results appeared across roughly 10 benchmarks.
Comparison: Gemini 3.1 Pro, Gemini 3.1 Flash Lite, Gemini 3 Flash, and Gemini 3.5 Flash.
Tool: OpenMark; the post did not disclose the task YAML, grader, complete responses, or cost logs.
The task inputs, system prompt, images, model routing, temperature, token limit, and scoring rules were not disclosed/unverifiable.
The author emphasized that these were their own tasks and cautioned that Gemini results may depend heavily on prompt shape.
The post says Gemini 3.5 Flash ranked 13th in one visual evaluation, while Gemini 3.1 Pro and Gemini 3.1 Flash Lite ranked first and second, respectively, and Gemini 3 Flash also ranked above it.
The post says this visual result was an average across 5 runs, and reports a similar trend of older versions performing better across roughly 10 saved evaluations.
The post provides no raw score for each model, sample count, confidence interval, task text, or complete leaderboard, so it can only serve as a directional field report.
This report does not contradict official or specialized benchmarks claiming that “3.5 Flash has advantages in speed and tool tasks”: it measures a single user’s real-world task set, which notably includes visual tasks. The actionable conclusion is to fix the same prompt, inputs, and grader and run multiple A/B tests before migrating, then confirm whether the model is actually better in your own workflow.
A single user, an undisclosed task set, and an undisclosed grader make independent reproduction impossible.
“13th place” has no contextual sample, score, or statistical significance; it cannot be treated as a global model ranking.
The author says “roughly 10 evaluations” but provides no list; the share and difficulty of visual tasks are unknown.
OpenMark’s public homepage mainly describes the platform’s terms and features; this collection did not find the user’s specific evaluation data.
Export 10 sets of your own real-world tasks, fixing the text, images, tools, model versions, routing, and grader.
Compare at least Gemini 3.5 Flash, Gemini 3.1 Pro, Gemini 3.1 Flash Lite, and Gemini 3 Flash, running each 5 times.
Save each output, success criteria, score, token usage, latency, and failure reason; report the mean, standard deviation, and results by task segment.
Report visual, tool, coding, and knowledge tasks separately to prevent one extreme visual task from overwhelming the overall mean.
The author’s core reminder is “benchmark first rather than assume newer = better.” This is personal experience rather than a quantitative conclusion, but it is well suited as an acceptance gate for a version upgrade.
Gemini 3.5 Flash