Gemini 3.7 Flash · Community source · Editorial analysis
This evidence note covers “Gemini 3.7 Flash THOR Finding Triage Benchmark” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
Reports a score of 72.5%, 100% Threat Capture, and 0% Critical Misses on 189 real-world THOR findings, and says its false-positive filtering outperformed the tested Qwen 3.7 Max, Kimi K3, and DeepSeek V4.
Note: This is a single public result from the author's self-built benchmark and should be reviewed alongside its scoring… This is a necessary excerpt; read the original source for full context.
The following preserves the body text extracted from the page during this visit; visible text from page navigation, the platform's automatic translation, comments, and other elements is retained as-is.
Florian Roth @cyb3rops Translated from English Show original
Google released Gemini 3.7 Flash today, so I immediately tested it against my THOR finding triage benchmark.
And then... we have a new #1. By a pretty wide margin.
Gemini 3.7 Flash scored 72.5%, with 100% Threat Capture and 0% Critical Misses across 189 real-world THOR findings.
Over the past few weeks, I also tested Qwen 3.7 Max, Kimi K3, and DeepSeek V4. Their scores were all significantly lower.
Interestingly, they weren't really failing at identifying actual threats. Their bigger problem was false positives: they escalated too many benign or suspicious findings to analysts instead of filtering them out.
And that's exactly where Gemini 3.7 Flash is surprisingly strong. It captures threats without drowning analysts in unnecessary reviews.
For this kind of security-incident triage, it is easily the best model I've tested so far.
A few months ago I wrote about this benchmark and its scoring methodology: https://cyb3rops.medium.com/why-i-built-my-own-llm-benchmark-for-thor-finding-triage-c8492e3997dc
This document is an archive of source material and does not represent an endorsement of the original article's conclusions by Tabbit or the maintainer of this document. When citing benchmark scores, prices, or model capabilities, return to the original article to confirm the version, test set, and date.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
X · Florian Roth (@cyb3rops) · Original publication date 2026-08-14 · Site edit date 2026-09-20
Open original sourceGemini 3.7 Flash
Download the Tabbit client to check model access