Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Gemini 3.7 Flash · Community source · Editorial analysis

Gemini 3.7 Flash THOR Finding Triage Benchmark

This evidence note covers “Gemini 3.7 Flash THOR Finding Triage Benchmark” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceEditorial analysisEdited 2026-09-20

Test conditions

Test conditions
The THOR finding-triage test used 189 real-world findings; the author reports 72.5%, 100% Threat Capture, and 0% Critical Misses, while scoring details, repeats, and harness must be checked against the source benchmark.
Source boundary
Supports treating the result as one public external measurement for security triage, especially its false-positive filtering claim.
Unsupported claims
Does not support calling 72.5% a general security accuracy, performing unauthorized security actions, or claiming production protection.

Key data and applicable tasks

Summary

Reports a score of 72.5%, 100% Threat Capture, and 0% Critical Misses on 189 real-world THOR findings, and says its false-positive filtering outperformed the tested Qwen 3.7 Max, Kimi K3, and DeepSeek V4.

Note: This is a single public result from the author's self-built benchmark and should be reviewed alongside its scoring… This is a necessary excerpt; read the original source for full context.

Original article

The following preserves the body text extracted from the page during this visit; visible text from page navigation, the platform's automatic translation, comments, and other elements is retained as-is.


Florian Roth @cyb3rops Translated from English Show original

Google released Gemini 3.7 Flash today, so I immediately tested it against my THOR finding triage benchmark.

And then... we have a new #1. By a pretty wide margin.

Gemini 3.7 Flash scored 72.5%, with 100% Threat Capture and 0% Critical Misses across 189 real-world THOR findings.

Over the past few weeks, I also tested Qwen 3.7 Max, Kimi K3, and DeepSeek V4. Their scores were all significantly lower.

Interestingly, they weren't really failing at identifying actual threats. Their bigger problem was false positives: they escalated too many benign or suspicious findings to analysts instead of filtering them out.

And that's exactly where Gemini 3.7 Flash is surprisingly strong. It captures threats without drowning analysts in unnecessary reviews.

For this kind of security-incident triage, it is easily the best model I've tested so far.

A few months ago I wrote about this benchmark and its scoring methodology: https://cyb3rops.medium.com/why-i-built-my-own-llm-benchmark-for-thor-finding-triage-c8492e3997dc


Usage notes

This document is an archive of source material and does not represent an endorsement of the original article's conclusions by Tabbit or the maintainer of this document. When citing benchmark scores, prices, or model capabilities, return to the original article to confirm the version, test set, and date.

What this supports

  • Supports treating the result as one public external measurement for security triage, especially its false-positive filtering claim.

What this does not support

  • Does not support calling 72.5% a general security accuracy, performing unauthorized security actions, or claiming production protection.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X · Florian Roth (@cyb3rops) · Original publication date 2026-08-14 · Site edit date 2026-09-20

Open original source

Gemini 3.7 Flash

Compare Gemini 3.7 Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Gemini 3.7 Flash: What It Is, Access, and Where It Fits

A sourced Gemini 3.7 Flash overview covering the 3.6 upgrade, the 3.8 relationship, API limits, access routes, price timing, and practical risks.

Related reviews

Gemini 3.7 Flash benchmarkThis evidence note covers “Gemini 3.7 Flash benchmark” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Gemini 3.7 Flash Benchmarks, Pricing & SpeedThis evidence note covers “Gemini 3.7 Flash Benchmarks, Pricing & Speed” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Google DeepMind Gemini 3.7 Flash Model CardThis evidence note covers “Google DeepMind Gemini 3.7 Flash Model Card” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Gemini 3.7 Flash: Real Pricing, Speed, and Agent Boundaries (eesel)This evidence note covers “Gemini 3.7 Flash: Real Pricing, Speed, and Agent Boundaries (eesel)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Gemini 3.7 Flash image prompt — causal visual descriptionTurn “Gemini 3.7 Flash image prompt — causal visual description” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.Gemini 3.7 Flash with one prompt — voxel Japanese pagoda gardenTurn “Gemini 3.7 Flash with one prompt — voxel Japanese pagoda garden” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.How To Use The Latest Gemini Flash Model For SEO AutomationTurn “How To Use The Latest Gemini Flash Model For SEO Automation” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.Introducing Gemini 3.7 Flash — coding and agent workflowsTurn “Introducing Gemini 3.7 Flash — coding and agent workflows” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.