Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Luna · Community source · Personal experience

I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation

This evidence note covers “I Benchmarked GPT-5.6 Sol/Luna/Terra by Role: Role-Based Evaluation” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Test conditions
The role-based sample covered strategy, repository execution, and two planted-bug repairs; model identities were visible, CLI latency was mixed in, and the author reported over-reasoning in some Luna high/max repairs.
Source boundary
Supports routing discussion by role rather than one total score, with the sample, latency, and human-scoring limits visible.
Unsupported claims
Does not support generalizing failure across coding tasks or inferring API-token costs.

Key data and applicable tasks

Summary

The author did not collapse the models into a single overall score, instead measuring them in roles such as strategic decision-making, repository execution, and code repair. The results show that Sol high is suited to the main strategic session, while Sol medium is faster and uses fewer resources; in this small-sample execution test, Luna high/max did not beat the existing routing and even showed signs of over-reasoning.

Results

  • In strategic tasks, Luna max scored 88, tying Terra max, but below Sol high at 94 and Sol max at 90.

  • Repository execution brief: Sol high scored 93 in 80.73 seconds with 1,818 reasoning tokens; Sol medium scored 91 in 70.66 seconds with 779 tokens.

  • In the small repair, Luna high went 4/4 in 52.59 seconds; Luna max went 4/4 in 121.34 seconds, which the author described as “heavily over-reasoned.”

  • The final routing did not give Luna a fixed slot; the author excluded Luna from the main-session, review, and implementation-rollback paths.

  • The author explicitly noted that the sample was small, the evaluator knew the model identities, CLI latency was mixed in, and subscription consumption was not equivalent to API token cost; the results cannot be generalized.

Original article

I ran a small role-based benchmark: three strategic decision memos, one repository-grounded execution brief, and two pla… This is a necessary excerpt; read the original source for full context.

What this supports

  • Supports routing discussion by role rather than one total score, with the sample, latency, and human-scoring limits visible.

What this does not support

  • Does not support generalizing failure across coding tasks or inferring API-token costs.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit, r/LLMDevs · u/petburiraja · Original publication date Unknown · Site edit date 2026-09-20

Open original source

GPT-5.6 Luna

Compare GPT-5.6 Luna in Tabbit

Download the Tabbit client to check model access

Related reviews

GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)This evidence note covers “GPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True PositiveThis evidence note covers “GPT-5.6 Luna Semgrep IDOR Security Benchmark and Cost per True Positive” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing RecommendationsThis evidence note covers “GPT-5.6 Sol, Terra, and Luna: Three-Tier Reddit Benchmarks and Routing Recommendations” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision TasksThis evidence note covers “GPT-5.6 Luna and Gemini 3.6 Flash: A Cost Counterexample in Document-Vision Tasks” under stated conditions; its version, sample, and runtime limits do not support a universal ranking or current production guarantee.Get Started with OpenAI GPT-5.6 on Amazon Bedrock: Reasoning, Tool Calling, and CachingAWS's article positions Luna as a high-throughput, low-latency model for classification, summarization, and routing, and uses Responses API examples to show how to set reasoning effort, call tools, carry the model's output in full into the next turn, and cache。Apply Occam’s Razor: Reducing Overengineering in Luna/Codex PromptsTurn “Apply Occam’s Razor: Reducing Overengineering in Luna/Codex Prompts” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.Reddit Codex: Diagnostic and Verification Prompt for Luna Subagent CompatibilityTurn “Reddit Codex: Diagnostic and Verification Prompt for Luna Subagent Compatibility” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.The Builder's Guide to GPT-5.6: Luna's Model Selection, Agent Orchestration, and CachingTurn “The Builder's Guide to GPT-5.6: Luna's Model Selection, Agent Orchestration, and Caching” into a bounded task entry with explicit inputs, runtime context, output format, and acceptance checks; confirm the model version and source limits before use.