Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Luna · Media / benchmark · Independent measurement

GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)

This evidence note covers “GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Test conditions
The BenchLM snapshot records 67.3/100, rank 23/218, Coding 73.0 (6/135), Agentic 43/130, and input/output/cache prices; speed and some categories lack sufficient evidence.
Source boundary
Supports a date-bounded reading of coding, agentic evidence, and rates; cached input must not be mixed with uncached input.
Unsupported claims
Does not support current pricing, a unified capability ranking, Tabbit account availability, or a production-cost forecast.

Key data and applicable tasks

Summary

BenchLM rates GPT-5.6 Luna at 67.3/100, ranking it #23 among 218 models. Its strongest area in the public evidence is Coding: 73.0 points, #6/135; Agentic ranks #43/130. The page also lists $0.20 per million input tokens, $1.20 per million output tokens, and cached input at $0.020 per million tokens. BenchLM specifically warns that speed has not been independently measured, some categories do not have sufficient public evidence, and the ranking should not be understood separately from evidence coverage.

Key results

ItemPage result
Overall score67.3/100
Public ranking#23/218
Coding73.0, #6/135, 96th percentile
Agentic53.3, #43/130
Knowledge81.6, #12/57
Multimodal66.1, #17/32
SWE-bench Pro62.7%
Terminal-Bench 2.084.7%
deepSWE67.2%
FrontierCode 1.1 Extended55.1%
cursorBench3261.1%
API pricing$0.20 input / $1.20 output per million tokens
Cached input$0.020 per million tokens
Context1.05M tokens
Independent speedNot measured

Assessment

  • The scenarios worth validating first are code generation, software development, and high-volume reasoning workloads.

  • The overall ranking should not be viewed in isolation: the page says 24 public benchmark rows have sources, but several tracked metrics are still blank.

  • The Coding results do not lead in every project: SWE-bench Pro is 17.6 percentage points below the best verified result recorded on the page, while Terminal-Bench 2.0 is 7.2 percentage points below GPT-5.6 Sol.

  • This is a compilation of third-party public information, not a retest in real Tabbit workflows. Before deployment, use your own task set to validate acceptance rate, latency, and total cost.

Original article

The following is the main visible body extracted this time through the Tabbit international page, with the original English retained; the page's navigation, collapsed controls, and unrelated footer have been omitted.


GPT-5.6 Luna

Current Released Jul 9, 2026 Proprietary Reasoning 1.05M context

Released Jul 9, 2026

DECISION READING

GPT-5.6 Luna scores 67.3 out of 100 and ranks #23 of 218. This profile shows 24 source-displayable benchmark rows; its strongest eligible category is Coding at #6. API pricing is $0.2 input and $1.2 output per million tokens, with cached input at $0.02.

Data as of August 15, 2026.

Strongest published evidence

Coding ranks #6. Particularly well-suited for software development and code generation tasks.

Validate before choosing

24 published rows leave some tracked benchmark slots empty. Independent runtime speed has not been measured.

Decision snapshot

Each value carries a field reference instead of floating alone. Markers compare this model with the current ranked and priced catalog; they are not absolute quality thresholds.

CAPABILITY

67.3/100 field median 58.2 #23 of 218 ranked models

PRICE

$0.20 input / $1.20 output input median $1 cached $0.020 · blended $0.70

SPEED

Not measured field median 94 tok/s Time to first token not measured

CONTEXT

1.05M tokens field median 200,000 Maximum output length is tracked separately

Eligible category ranks

Agentic: #43/130 Coding: #6/135 Reasoning: Not ranked Knowledge: #12/57 Math: Not ranked Multilingual: Not ranked Multimodal: #17/32 Instruction Following: Not ranked

How much of this is verified

Coverage is split by category so a strong number never hides a thin evidence base. Verified means the row is tied to a published source; provisional rows remain visible but separate.

Agentic: 7/7 verified Coding: 5/5 verified Reasoning: 2/2 verified Knowledge: 4/4 verified Math: 3/3 verified Multilingual: Not measured Multimodal: 2/2 verified Instruction Following: Not measured

Each documented value carries its source. Missing fields stay visible as not sourced or not published, rather than disappearing from the page.

API model ID: gpt-5.6-luna Context window: 1.05M Input modalities: text, image Output modalities: text Parameters: Not disclosed by the provider Availability: OpenAI Responses API Lifecycle: active Prompt caching: Published at $0.020 per million cached input tokens Self-host: Weights are not published

Category scores

Agentic: 53.3, #43 of 130, 67th percentile, 22% weight, 7 benchmarks, Verified Coding: 73.0, #6 of 135, 96th percentile, 20% weight, 5 benchmarks, Verified Reasoning: 57.8, not ranked, 17% weight, 2 benchmarks, Verified Knowledge: 81.6, #12 of 57, 80th percentile, 12% weight, 4 benchmarks, Verified Math: 96.9, not ranked, 5% weight, 3 benchmarks, Verified Multilingual: Not measured, 7% weight, 0 benchmarks Multimodal: 66.1, #17 of 32, 48th percentile, 12% weight, 2 benchmarks, Verified Instruction Following: Not measured, 5% weight, 0 benchmarks

Benchmark ledger — Coding

SWE-bench Pro: 62.7%. Best verified: Claude Mythos 5 at 80.3%; 17.6 behind; weighted 10%; provider exact, OpenAI: GPT-5.6.

Terminal-Bench 2.0: 84.7%. Best verified: GPT-5.6 Sol at 91.9%; 7.2 behind; display only; provider exact, OpenAI: GPT-5.6.

deepSWE: 67.2%. Best verified: GPT-5.6 Sol at 72.7%; 5.5 behind; display only; provider exact, OpenAI: GPT-5.6.

FrontierCode 1.1 Extended: 55.1%. Best verified: Claude Opus 5 at 63.6%; 8.5 behind; display only; provider exact, Cognition: GPT-5.6 models in Devin.

cursorBench32: 61.1%. Best verified: Grok 4.6 at 70.8%; 9.7 behind; display only; benchmark exact, Cursor evals: CursorBench 3.2.

How to read this profile

GPT-5.6 Luna ranks #23 of 218 on the public leaderboard with a score of 67.29/100. It does not yet have enough sourced coverage for a verified position.

GPT-5.6 Luna is a proprietary model with a 1.05M context window. It uses an explicit reasoning mode, which can improve complex problem solving while adding latency and token use.

GPT-5.6 Luna sits in the GPT-5.6 family with GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Cyber. 24 of 437 tracked benchmark slots currently have displayable evidence. Missing categories stay blank.

Its strongest eligible category is Coding at #6, while its lowest eligible position is Agentic at #43. Particularly well-suited for software development and code generation tasks.

What this supports

  • Supports a date-bounded reading of coding, agentic evidence, and rates; cached input must not be mixed with uncached input.

What this does not support

  • Does not support current pricing, a unified capability ranking, Tabbit account availability, or a production-cost forecast.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

BenchLM.ai · BenchLM.ai · Original publication date Unknown · Site edit date 2026-09-20

Open original source

GPT-5.6 Luna

Compare GPT-5.6 Luna in Tabbit

Download the Tabbit client to check model access

Related reviews

GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True PositiveThis evidence note covers “GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based EvaluationThis evidence note covers “I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing RecommendationsThis evidence note covers “GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing Recommendations” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision TasksThis evidence note covers “GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Get Started with OpenAI GPT-5.6 on Amazon Bedrock: Reasoning, Tool Calling, and CachingAWS's article positions Luna as a high-throughput, low-latency model for classification, summarization, and routing, and uses Responses API examples to show how to set reasoning effort, call tools, carry the model's output in full into the next turn, and cache。GPT-5.6 Luna API Model Parameters and Cost ConfigurationOne-sentence takeaway The official model page confirms Luna's current API ID, pricing, reasoning tiers, tool surface, and rate limits. It can serve as a configuration baseline for high-throughput routing, but a single request above 272K tokens incurs a surchar。GPT-5.6 Prompting Guide: Luna's Work Contract and Model RoutingTurn “GPT-5.6 Prompting Guide: Luna's Work Contract and Model Routing” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.The Builder's Guide to GPT-5.6: Luna's Model Selection, Agent Orchestration, and CachingTurn “The Builder's Guide to GPT-5.6: Luna's Model Selection, Agent Orchestration, and Caching” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.