GPT-5.6 Luna

GPT-5.6 Luna · Reviews and evidence

Which GPT-5.6 Luna conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

This evidence note covers “GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

BenchLM.ai · Read evidence

This evidence note covers “I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Reddit, r/LLMDevs · Read evidence

This evidence note covers “Agents on Rails: 8 Models, 21 Atomic Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

X · Read evidence

Selected evidence

Media / benchmarkIndependent measurement

GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)

This evidence note covers “GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceBenchLM.ai
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
The BenchLM snapshot records 67.3/100, rank 23/218, Coding 73.0 (6/135), Agentic 43/130, and input/output/cache prices; speed and some categories lack sufficient evidence.
Source boundary
Supports a date-bounded reading of coding, agentic evidence, and rates; cached input must not be mixed with uncached input.
Unsupported claims
Does not support current pricing, a unified capability ranking, Tabbit account availability, or a production-cost forecast.
Information extractionReasoningCost
CommunityPersonal experience

I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation

This evidence note covers “I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit, r/LLMDevs
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
The role-based sample covered strategy, repository execution, and two planted-bug repairs; model identities were visible, CLI latency was mixed in, and the author reported over-reasoning in some Luna high/max repairs.
Source boundary
Supports routing discussion by role rather than one total score, with the sample, latency, and human-scoring limits visible.
Unsupported claims
Does not support generalizing failure across coding tasks or inferring API-token costs.
Information extractionReasoning
CommunityEditorial analysis

Agents on Rails: 8 Models, 21 Atomic Tasks

This evidence note covers “Agents on Rails: 8 Models, 21 Atomic Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceX
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
Agents on Rails compared eight models across 21 atomic tasks; the X snippet reports 73% successful runs for Luna at default medium reasoning and about $0.90 across 63 runs, but the task list and scoring are not public.
Source boundary
Supports using the snippet as a retest lead about low-cost medium reasoning, not as a reproducible success rate.
Unsupported claims
Does not support calling 73% a general coding pass rate or substituting a search snippet for the original test.
CodingAgent
Media / benchmarkIndependent measurement

GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive

This evidence note covers “GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceSemgrep Security Research
Published2026-07-13
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
Semgrep ran a ten-repeat benchmark on production-app samples centered on IDOR, comparing a guided prompt with a tool-equipped harness; it reports roughly six-times lower cost per true positive for Luna with a marginal F1 trade-off, while exact chart values are not fully public.
Source boundary
Supports discussing precision/recall, F1, harness effects, and cost per true positive in security review.
Unsupported claims
Does not support treating CTF/harness results as bare-model capability, production protection, or authorization to test vulnerabilities.
Information extractionReasoningCost

All sources

All sources

12 / 12
Media / benchmarkIndependent measurement

GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)

This evidence note covers “GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceBenchLM.ai
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
The BenchLM snapshot records 67.3/100, rank 23/218, Coding 73.0 (6/135), Agentic 43/130, and input/output/cache prices; speed and some categories lack sufficient evidence.
Source boundary
Supports a date-bounded reading of coding, agentic evidence, and rates; cached input must not be mixed with uncached input.
Unsupported claims
Does not support current pricing, a unified capability ranking, Tabbit account availability, or a production-cost forecast.
Information extractionReasoningCost
CommunityPersonal experience

I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation

This evidence note covers “I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit, r/LLMDevs
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
The role-based sample covered strategy, repository execution, and two planted-bug repairs; model identities were visible, CLI latency was mixed in, and the author reported over-reasoning in some Luna high/max repairs.
Source boundary
Supports routing discussion by role rather than one total score, with the sample, latency, and human-scoring limits visible.
Unsupported claims
Does not support generalizing failure across coding tasks or inferring API-token costs.
Information extractionReasoning
CommunityEditorial analysis

Agents on Rails: 8 Models, 21 Atomic Tasks

This evidence note covers “Agents on Rails: 8 Models, 21 Atomic Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceX
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
Agents on Rails compared eight models across 21 atomic tasks; the X snippet reports 73% successful runs for Luna at default medium reasoning and about $0.90 across 63 runs, but the task list and scoring are not public.
Source boundary
Supports using the snippet as a retest lead about low-cost medium reasoning, not as a reproducible success rate.
Unsupported claims
Does not support calling 73% a general coding pass rate or substituting a search snippet for the original test.
CodingAgent
Media / benchmarkIndependent measurement

GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive

This evidence note covers “GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceSemgrep Security Research
Published2026-07-13
Collected2026-08-17

Unverified: the original source could not be rechecked.

Test conditions
Semgrep ran a ten-repeat benchmark on production-app samples centered on IDOR, comparing a guided prompt with a tool-equipped harness; it reports roughly six-times lower cost per true positive for Luna with a marginal F1 trade-off, while exact chart values are not fully public.
Source boundary
Supports discussing precision/recall, F1, harness effects, and cost per true positive in security review.
Unsupported claims
Does not support treating CTF/harness results as bare-model capability, production protection, or authorization to test vulnerabilities.
Information extractionReasoningCost
OfficialVendor report

GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing Recommendations

This evidence note covers “GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing Recommendations” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit, r/OpenAI
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
Provider / environment
Reddit, r/OpenAI; the original conditions do not establish one controlled retest.
Collection boundary
The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
Information extractionReasoning
CommunityPersonal experience

GPT-5.6 Luna Is Really Underrated: Codex User Experience

This evidence note covers “GPT-5.6 Luna Is Really Underrated: Codex User Experience” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit, r/codex
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
Provider / environment
Reddit, r/codex; the original conditions do not establish one controlled retest.
Collection boundary
The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
Coding
CommunityPersonal experience

Thoughts after using GPT-5.6 Luna for 48 hours

This evidence note covers “Thoughts after using GPT-5.6 Luna for 48 hours” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit, r/hermesagent
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
Provider / environment
Reddit, r/hermesagent; the original conditions do not establish one controlled retest.
Collection boundary
The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
CodingAgent
CommunityEditorial analysis

GPT-5.6 Luna Max vs. Sol Medium: An X User's Real-World Cost Test

This evidence note covers “GPT-5.6 Luna Max vs. Sol Medium: An X User's Real-World Cost Test” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceX
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
Provider / environment
X; the original conditions do not establish one controlled retest.
Collection boundary
The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
Cost
CommunityPersonal experience

GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision Tasks

This evidence note covers “GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit, r/GoogleGeminiAI
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
Provider / environment
Reddit, r/GoogleGeminiAI; the original conditions do not establish one controlled retest.
Collection boundary
The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
Visual generationInformation extractionCost
CommunityPersonal experience

GPT-5.6 Luna vs. DeepSeek V4 Flash: Cache Hits and Real-World Task Costs

This evidence note covers “GPT-5.6 Luna vs. DeepSeek V4 Flash: Cache Hits and Real-World Task Costs” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit, r/DeepSeek
PublishedUnknown
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
Provider / environment
Reddit, r/DeepSeek; the original conditions do not establish one controlled retest.
Collection boundary
The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
apiCost
CommunityPersonal experience

5.6 Luna Extra High is the work horse I needed

This evidence note covers “5.6 Luna Extra High is the work horse I needed” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit, r/codex
Published2026-07-17
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
Provider / environment
Reddit, r/codex; the original conditions do not establish one controlled retest.
Collection boundary
The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
Coding
CommunityPersonal experience

GPT-5.6 Luna Reddit Codex Quota and Cache Cost: A Hands-on Measurement

This evidence note covers “GPT-5.6 Luna Reddit Codex Quota and Cache Cost: A Hands-on Measurement” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

SourceReddit r/codex
Published2026-08-01
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model and version
GPT-5.6 Luna; do not merge with other versions, reasoning tiers, or harnesses.
Provider / environment
Reddit r/codex; the original conditions do not establish one controlled retest.
Collection boundary
The source note was collected on 2026-08-17/18; the original page was not reopened this round, so dynamic facts remain unverified.
CodingapiCost

GPT-5.6 Luna

Compare GPT-5.6 Luna in Tabbit

Model access, features, and permissions depend on your current client account.