Reports a score of 72.5%, 100% Threat Capture, and 0% Critical Misses on 189 real-world THOR findings, and says its false-positive filtering outperformed the tested Qwen 3.7 Max, Kimi K3, and DeepSeek V4.
Note: This is a single public result from the author's self-built benchmark and should be reviewed alongside its scoring… This is a necessary excerpt; read the original source for full context.
The following preserves the body text extracted from the page during this visit; visible text from page navigation, the platform's automatic translation, comments, and other elements is retained as-is.
Florian Roth @cyb3rops Translated from English Show original
Google released Gemini 3.7 Flash today, so I immediately tested it against my THOR finding triage benchmark.
And then... we have a new #1. By a pretty wide margin.
Gemini 3.7 Flash scored 72.5%, with 100% Threat Capture and 0% Critical Misses across 189 real-world THOR findings.
Over the past few weeks, I also tested Qwen 3.7 Max, Kimi K3, and DeepSeek V4. Their scores were all significantly lower.
Interestingly, they weren't really failing at identifying actual threats. Their bigger problem was false positives: they escalated too many benign or suspicious findings to analysts instead of filtering them out.
And that's exactly where Gemini 3.7 Flash is surprisingly strong. It captures threats without drowning analysts in unnecessary reviews.
For this kind of security-incident triage, it is easily the best model I've tested so far.
A few months ago I wrote about this benchmark and its scoring methodology: https://cyb3rops.medium.com/why-i-built-my-own-llm-benchmark-for-thor-finding-triage-c8492e3997dc
This document is an archive of source material and does not represent an endorsement of the original article's conclusions by Tabbit or the maintainer of this document. When citing benchmark scores, prices, or model capabilities, return to the original article to confirm the version, test set, and date.
Gemini 3.7 Flash