This evidence note covers “GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
BenchLM.ai · Read evidenceGPT-5.6 Luna · Reviews and evidence
Which GPT-5.6 Luna conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
This evidence note covers “I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Reddit, r/LLMDevs · Read evidenceThis evidence note covers “Agents on Rails: 8 Models, 21 Atomic Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
X · Read evidenceSelected evidence
GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)
This evidence note covers “GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- The BenchLM snapshot records 67.3/100, rank 23/218, Coding 73.0 (6/135), Agentic 43/130, and input/output/cache prices; speed and some categories lack sufficient evidence.
- Source boundary
- Supports a date-bounded reading of coding, agentic evidence, and rates; cached input must not be mixed with uncached input.
- Unsupported claims
- Does not support current pricing, a unified capability ranking, Tabbit account availability, or a production-cost forecast.
I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation
This evidence note covers “I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- The role-based sample covered strategy, repository execution, and two planted-bug repairs; model identities were visible, CLI latency was mixed in, and the author reported over-reasoning in some Luna high/max repairs.
- Source boundary
- Supports routing discussion by role rather than one total score, with the sample, latency, and human-scoring limits visible.
- Unsupported claims
- Does not support generalizing failure across coding tasks or inferring API-token costs.
Agents on Rails: 8 Models, 21 Atomic Tasks
This evidence note covers “Agents on Rails: 8 Models, 21 Atomic Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- Agents on Rails compared eight models across 21 atomic tasks; the X snippet reports 73% successful runs for Luna at default medium reasoning and about $0.90 across 63 runs, but the task list and scoring are not public.
- Source boundary
- Supports using the snippet as a retest lead about low-cost medium reasoning, not as a reproducible success rate.
- Unsupported claims
- Does not support calling 73% a general coding pass rate or substituting a search snippet for the original test.
GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive
This evidence note covers “GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- Semgrep ran a ten-repeat benchmark on production-app samples centered on IDOR, comparing a guided prompt with a tool-equipped harness; it reports roughly six-times lower cost per true positive for Luna with a marginal F1 trade-off, while exact chart values are not fully public.
- Source boundary
- Supports discussing precision/recall, F1, harness effects, and cost per true positive in security review.
- Unsupported claims
- Does not support treating CTF/harness results as bare-model capability, production protection, or authorization to test vulnerabilities.
All sources
All sources
GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)
This evidence note covers “GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- The BenchLM snapshot records 67.3/100, rank 23/218, Coding 73.0 (6/135), Agentic 43/130, and input/output/cache prices; speed and some categories lack sufficient evidence.
- Source boundary
- Supports a date-bounded reading of coding, agentic evidence, and rates; cached input must not be mixed with uncached input.
- Unsupported claims
- Does not support current pricing, a unified capability ranking, Tabbit account availability, or a production-cost forecast.
I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation
This evidence note covers “I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- The role-based sample covered strategy, repository execution, and two planted-bug repairs; model identities were visible, CLI latency was mixed in, and the author reported over-reasoning in some Luna high/max repairs.
- Source boundary
- Supports routing discussion by role rather than one total score, with the sample, latency, and human-scoring limits visible.
- Unsupported claims
- Does not support generalizing failure across coding tasks or inferring API-token costs.
Agents on Rails: 8 Models, 21 Atomic Tasks
This evidence note covers “Agents on Rails: 8 Models, 21 Atomic Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- Agents on Rails compared eight models across 21 atomic tasks; the X snippet reports 73% successful runs for Luna at default medium reasoning and about $0.90 across 63 runs, but the task list and scoring are not public.
- Source boundary
- Supports using the snippet as a retest lead about low-cost medium reasoning, not as a reproducible success rate.
- Unsupported claims
- Does not support calling 73% a general coding pass rate or substituting a search snippet for the original test.
GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive
This evidence note covers “GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Test conditions
- Semgrep ran a ten-repeat benchmark on production-app samples centered on IDOR, comparing a guided prompt with a tool-equipped harness; it reports roughly six-times lower cost per true positive for Luna with a marginal F1 trade-off, while exact chart values are not fully public.
- Source boundary
- Supports discussing precision/recall, F1, harness effects, and cost per true positive in security review.
- Unsupported claims
- Does not support treating CTF/harness results as bare-model capability, production protection, or authorization to test vulnerabilities.
GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing Recommendations
This evidence note covers “GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing Recommendations” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Model and version
- GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
- Provider / environment
- Reddit, r/OpenAI; the original conditions do not establish one controlled retest.
- Collection boundary
- The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
GPT-5.6 Luna Is Really Underrated: Codex User Experience
This evidence note covers “GPT-5.6 Luna Is Really Underrated: Codex User Experience” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Model and version
- GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
- Provider / environment
- Reddit, r/codex; the original conditions do not establish one controlled retest.
- Collection boundary
- The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
Thoughts after using GPT-5.6 Luna for 48 hours
This evidence note covers “Thoughts after using GPT-5.6 Luna for 48 hours” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Model and version
- GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
- Provider / environment
- Reddit, r/hermesagent; the original conditions do not establish one controlled retest.
- Collection boundary
- The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
GPT-5.6 Luna Max vs. Sol Medium: An X User's Real-World Cost Test
This evidence note covers “GPT-5.6 Luna Max vs. Sol Medium: An X User's Real-World Cost Test” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Model and version
- GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
- Provider / environment
- X; the original conditions do not establish one controlled retest.
- Collection boundary
- The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision Tasks
This evidence note covers “GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Model and version
- GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
- Provider / environment
- Reddit, r/GoogleGeminiAI; the original conditions do not establish one controlled retest.
- Collection boundary
- The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
GPT-5.6 Luna vs. DeepSeek V4 Flash: Cache Hits and Real-World Task Costs
This evidence note covers “GPT-5.6 Luna vs. DeepSeek V4 Flash: Cache Hits and Real-World Task Costs” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Model and version
- GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
- Provider / environment
- Reddit, r/DeepSeek; the original conditions do not establish one controlled retest.
- Collection boundary
- The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
5.6 Luna Extra High is the work horse I needed
This evidence note covers “5.6 Luna Extra High is the work horse I needed” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Model and version
- GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
- Provider / environment
- Reddit, r/codex; the original conditions do not establish one controlled retest.
- Collection boundary
- The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
GPT-5.6 Luna Reddit Codex Quota and Cache Cost: A Hands-on Measurement
This evidence note covers “GPT-5.6 Luna Reddit Codex Quota and Cache Cost: A Hands-on Measurement” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.
Unverified: the original source could not be rechecked.
- Model and version
- GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
- Provider / environment
- Reddit r/codex; the original conditions do not establish one controlled retest.
- Collection boundary
- The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
GPT-5.6 Luna
Compare GPT-5.6 Luna in Tabbit
Model access, features, and permissions depend on your current client account.