Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Luna · Media / benchmark · Independent measurement

GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive

This evidence note covers “GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Test conditions
Semgrep ran a ten-repeat benchmark on production-app samples centered on IDOR, comparing a guided prompt with a tool-equipped harness; it reports roughly six-times lower cost per true positive for Luna with a marginal F1 trade-off, while exact chart values are not fully public.
Source boundary
Supports discussing precision/recall, F1, harness effects, and cost per true positive in security review.
Unsupported claims
Does not support treating CTF/harness results as bare-model capability, production protection, or authorization to test vulnerabilities.

Key data and applicable tasks

One-sentence takeaway

Semgrep's security evaluation found that Luna costs roughly 6 times less per true positive than heavier models, with only a marginal F1 sacrifice; however, precision and recall are strongly dependent on the harness, so raw-model scores must not be treated as production protection capability.

Test environment

  • Models: GPT-5.6 Luna, Terra, Sol, and GPT-5.5 as the control.

  • Task: Primarily IDOR detection, with mentions of vulnerability classes such as SSRF; the tests used representative production applications.

  • Repetitions: The article says that security researchers rigorously reviewed the true positives for each benchmark and ran it ten times.

  • Two conditions: The model on its own with a guided prompt, and a harness equipped with Semgrep tools, scaffolding, and a complete security workflow.

  • Metrics: recall (finding true positives), precision (the share of reported findings that are real vulnerabilities), and the F1 harmonic mean; cost per true positive was also counted.

Input/configuration

The article did not publish complete IDOR samples, prompts, model snapshots, per-run outputs, pricing bills, or the CSV underlying the charts; the page body provides the methodology, directional results, and cost multiples. The production applications, researcher review, and ten runs are clues for reproducing the method, but are insufficient for recomputing it item by item.

Results data

  • All GPT-5.6 reasoning levels outperformed GPT-5.5 on this IDOR benchmark.

  • Compared with GPT-5.5, GPT-5.6 improved true-positive recall by more than 2 times, but precision regressed, meaning more false positives.

  • Luna's cost per true positive is the standout result: the article says it is about 6 times cheaper than heavier models, while giving up only a marginal amount of F1.

  • The article presents OpenAI CTF's 96.7% as a "result with a harness," not as raw-model capability; after removing the preconfigured prompt and tool environment, one cannot assume that the figure would remain the same.

  • The article's IDOR charts are presented as images; the readable page text does not provide exact recall, precision, or F1 values for Luna, Terra, and Sol, so no unverified numbers from the charts are added here.

Conclusions

  • For high-volume, narrow-scope vulnerability triage that can be reviewed by security researchers, Luna may be a cost-effective first-pass filter.

  • The cost of high recall is lower precision; production systems should use static analysis, reachability checks, manual review, or a second model to filter false positives.

  • Tools and the harness are not ancillary: Semgrep's results clearly support the view that the combination of "model + tools + scaffolding" determines the effectiveness of the security workflow.

Limitations

  • Semgrep is a security product vendor, so the test tasks, toolchain, and commercial objectives may affect the results; this should be viewed as an independent engineering evaluation rather than a blind test.

  • The article discloses only directional multiples, not per-sample data, complete prompts, or the raw chart values, so the 6-times figure or the F1 difference cannot be independently recomputed.

  • IDOR and production-application security tests do not represent malware, supply-chain, compliance-audit, or other vulnerability categories.

  • CTFs and other public benchmarks may have training-data contamination or harness bias; the article itself also warns that models and harnesses must be evaluated together.

Reproduction steps

  1. Prepare de-identified representative applications and IDOR/SSRF tasks, first freezing the vulnerability labels and the security researchers' review rules.

  2. Fix Luna's model alias, reasoning effort, guided prompt, tool permissions, and timeout; run it at least ten times and save the complete outputs.

  3. Run GPT-5.5, Terra, or Sol as controls under the same conditions; then separately enable and disable the security-tool harness.

  4. For each run, calculate recall, precision, F1, false-positive count, manual review time, input/output tokens, and dollar cost.

  5. Report cost per true positive, and put the "raw model" and "complete security workflow" in separate tables to avoid attributing the harness gain to Luna.

Source excerpt or observation (for compliance short quote only)

Semgrep's core judgment is “Luna as a cost per true positive was a standout result,” but the same article emphasizes that the harness can significantly change the result.

What this supports

  • Supports discussing precision/recall, F1, harness effects, and cost per true positive in security review.

What this does not support

  • Does not support treating CTF/harness results as bare-model capability, production protection, or authorization to test vulnerabilities.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Semgrep Security Research · Brenden Noblitt, Erik Buchanan, Jayson DeLancey · Original publication date 2026-07-13 · Site edit date 2026-09-20

Open original source

GPT-5.6 Luna

Compare GPT-5.6 Luna in Tabbit

Download the Tabbit client to check model access

Related reviews

GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)This evidence note covers “GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based EvaluationThis evidence note covers “I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing RecommendationsThis evidence note covers “GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing Recommendations” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision TasksThis evidence note covers “GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Get Started with OpenAI GPT-5.6 on Amazon Bedrock: Reasoning, Tool Calling, and CachingAWS's article positions Luna as a high-throughput, low-latency model for classification, summarization, and routing, and uses Responses API examples to show how to set reasoning effort, call tools, carry the model's output in full into the next turn, and cache。GPT-5.6 Luna API Model Parameters and Cost ConfigurationOne-sentence takeaway The official model page confirms Luna's current API ID, pricing, reasoning tiers, tool surface, and rate limits. It can serve as a configuration baseline for high-throughput routing, but a single request above 272K tokens incurs a surchar。GPT-5.6 Prompting Guide: Luna's Work Contract and Model RoutingTurn “GPT-5.6 Prompting Guide: Luna's Work Contract and Model Routing” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.The Builder's Guide to GPT-5.6: Luna's Model Selection, Agent Orchestration, and CachingTurn “The Builder's Guide to GPT-5.6: Luna's Model Selection, Agent Orchestration, and Caching” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.