BenchLM rates GPT-5.6 Luna at 67.3/100, ranking it #23 among 218 models. Its strongest area in the public evidence is Coding: 73.0 points, #6/135; Agentic ranks #43/130. The page also lists $0.20 per million input tokens, $1.20 per million output tokens, and cached input at $0.020 per million tokens. BenchLM specifically warns that speed has not been independently measured, some categories do not have sufficient public evidence, and the ranking should not be understood separately from evidence coverage.
| Item | Page result |
|---|---|
| Overall score | 67.3/100 |
| Public ranking | #23/218 |
| Coding | 73.0, #6/135, 96th percentile |
| Agentic | 53.3, #43/130 |
| Knowledge | 81.6, #12/57 |
| Multimodal | 66.1, #17/32 |
| SWE-bench Pro | 62.7% |
| Terminal-Bench 2.0 | 84.7% |
| deepSWE | 67.2% |
| FrontierCode 1.1 Extended | 55.1% |
| cursorBench32 | 61.1% |
| API pricing | $0.20 input / $1.20 output per million tokens |
| Cached input | $0.020 per million tokens |
| Context | 1.05M tokens |
| Independent speed | Not measured |
The scenarios worth validating first are code generation, software development, and high-volume reasoning workloads.
The overall ranking should not be viewed in isolation: the page says 24 public benchmark rows have sources, but several tracked metrics are still blank.
The Coding results do not lead in every project: SWE-bench Pro is 17.6 percentage points below the best verified result recorded on the page, while Terminal-Bench 2.0 is 7.2 percentage points below GPT-5.6 Sol.
This is a compilation of third-party public information, not a retest in real Tabbit workflows. Before deployment, use your own task set to validate acceptance rate, latency, and total cost.
The following is the main visible body extracted this time through the Tabbit international page, with the original English retained; the page's navigation, collapsed controls, and unrelated footer have been omitted.
GPT-5.6 Luna
Current Released Jul 9, 2026 Proprietary Reasoning 1.05M context
Released Jul 9, 2026
DECISION READING
GPT-5.6 Luna scores 67.3 out of 100 and ranks #23 of 218. This profile shows 24 source-displayable benchmark rows; its strongest eligible category is Coding at #6. API pricing is $0.2 input and $1.2 output per million tokens, with cached input at $0.02.
Data as of August 15, 2026.
Strongest published evidence
Coding ranks #6. Particularly well-suited for software development and code generation tasks.
Validate before choosing
24 published rows leave some tracked benchmark slots empty. Independent runtime speed has not been measured.
Decision snapshot
Each value carries a field reference instead of floating alone. Markers compare this model with the current ranked and priced catalog; they are not absolute quality thresholds.
CAPABILITY
67.3/100 field median 58.2 #23 of 218 ranked models
PRICE
$0.20 input / $1.20 output input median $1 cached $0.020 · blended $0.70
SPEED
Not measured field median 94 tok/s Time to first token not measured
CONTEXT
1.05M tokens field median 200,000 Maximum output length is tracked separately
Eligible category ranks
Agentic: #43/130 Coding: #6/135 Reasoning: Not ranked Knowledge: #12/57 Math: Not ranked Multilingual: Not ranked Multimodal: #17/32 Instruction Following: Not ranked
How much of this is verified
Coverage is split by category so a strong number never hides a thin evidence base. Verified means the row is tied to a published source; provisional rows remain visible but separate.
Agentic: 7/7 verified Coding: 5/5 verified Reasoning: 2/2 verified Knowledge: 4/4 verified Math: 3/3 verified Multilingual: Not measured Multimodal: 2/2 verified Instruction Following: Not measured
Each documented value carries its source. Missing fields stay visible as not sourced or not published, rather than disappearing from the page.
API model ID: gpt-5.6-luna Context window: 1.05M Input modalities: text, image Output modalities: text Parameters: Not disclosed by the provider Availability: OpenAI Responses API Lifecycle: active Prompt caching: Published at $0.020 per million cached input tokens Self-host: Weights are not published
Category scores
Agentic: 53.3, #43 of 130, 67th percentile, 22% weight, 7 benchmarks, Verified Coding: 73.0, #6 of 135, 96th percentile, 20% weight, 5 benchmarks, Verified Reasoning: 57.8, not ranked, 17% weight, 2 benchmarks, Verified Knowledge: 81.6, #12 of 57, 80th percentile, 12% weight, 4 benchmarks, Verified Math: 96.9, not ranked, 5% weight, 3 benchmarks, Verified Multilingual: Not measured, 7% weight, 0 benchmarks Multimodal: 66.1, #17 of 32, 48th percentile, 12% weight, 2 benchmarks, Verified Instruction Following: Not measured, 5% weight, 0 benchmarks
Benchmark ledger — Coding
SWE-bench Pro: 62.7%. Best verified: Claude Mythos 5 at 80.3%; 17.6 behind; weighted 10%; provider exact, OpenAI: GPT-5.6.
Terminal-Bench 2.0: 84.7%. Best verified: GPT-5.6 Sol at 91.9%; 7.2 behind; display only; provider exact, OpenAI: GPT-5.6.
deepSWE: 67.2%. Best verified: GPT-5.6 Sol at 72.7%; 5.5 behind; display only; provider exact, OpenAI: GPT-5.6.
FrontierCode 1.1 Extended: 55.1%. Best verified: Claude Opus 5 at 63.6%; 8.5 behind; display only; provider exact, Cognition: GPT-5.6 models in Devin.
cursorBench32: 61.1%. Best verified: Grok 4.6 at 70.8%; 9.7 behind; display only; benchmark exact, Cursor evals: CursorBench 3.2.
How to read this profile
GPT-5.6 Luna ranks #23 of 218 on the public leaderboard with a score of 67.29/100. It does not yet have enough sourced coverage for a verified position.
GPT-5.6 Luna is a proprietary model with a 1.05M context window. It uses an explicit reasoning mode, which can improve complex problem solving while adding latency and token use.
GPT-5.6 Luna sits in the GPT-5.6 family with GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Cyber. 24 of 437 tracked benchmark slots currently have displayable evidence. Missing categories stay blank.
Its strongest eligible category is Coding at #6, while its lowest eligible position is Agentic at #43. Particularly well-suited for software development and code generation tasks.
GPT-5.6 Luna