Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Prompts and workflows

GPT-5.6 Luna · configuration

GPT-5.6 Luna API Model Parameters and Cost Configuration

One-sentence takeaway The official model page confirms Luna's current API ID, pricing, reasoning tiers, tool surface, and rate limits. It can serve as a configuration baseline for high-throughput routing, but a single request above 272K tokens incurs a surchar。

Source reviewed; not testedGPT-5.6 Luna API-compatible client; confirm model ID, tools, reasoning settings and limits before execution.

Prerequisites and inputs

  • API credentials
  • model ID
  • tool schema
  • test input
  • failure handling

Task outcome\n\nTurn “GPT-5.6 Luna API Model Parameters and Cost Configuration” into a checkable starting workflow; do not treat a showcase as general capability.\n\n## Prerequisites\n\n- Prepare the goal, inputs, runtime, and acceptance criteria.\n- Confirm the live GPT-5.6 Luna entry, tool permissions, and version.\n\n## Steps\n\n1. Run a minimal input first and record the model, tools, latency, and failure state.\n2. Check the output against the required format and acceptance criteria, then add missing constraints.\n3. Manually review facts, code, visual artifacts, and external actions.\n\n## Boundary\n\nThis guide is an editorial adaptation of the source observation; it does not claim that Tabbit has verified the source client capabilities.

Read the source research notes

One-sentence takeaway

The official model page confirms Luna's current API ID, pricing, reasoning tiers, tool surface, and rate limits. It can serve as a configuration baseline for high-throughput routing, but a single request above 272K tokens incurs a surcharge on the entire request.

Use cases

  • Suitable tasks: cost-sensitive, high-volume classification, extraction, summarization, tagging, lightweight code explanations, and verifiable pipeline steps.

  • Unsuitable tasks: audio/video, fine-tuning, and tasks requiring complex cross-file architectural judgment; the page describes capabilities and billing only and does not promise business accuracy.

  • Applicable model version: gpt-5.6-luna.

  • Applicable client, Agent, or API: OpenAI API; the Responses API supports the tools listed on the page, and Chat Completions is also available.

  • Recommended reasoning tiers and parameters: start batch tasks with none or low; use medium as the default tier; try high, xhigh, or max only when acceptance fails and the improvement can be quantified.

Ready-to-use content

This is a minimal configuration checklist organized from the official fields; the application must fill in authentication, SDK initialization, and business-tool implementation.

model = "gpt-5.6-luna"
reasoning.effort = "low"  # none | low | medium | high | xhigh | max
endpoint = "/v1/responses"

price_per_1M:
  input = 0.20
  cached_input = 0.02
  output = 1.20

The Responses API tools page marks web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search as Supported; function calling, structured outputs, and streaming outputs are also marked Supported.

Test/workflow steps

  1. Fix gpt-5.6-luna and low first, then select a batch of tasks with clear labels or test criteria. Record input/output tokens, elapsed time, pass rate, and retry count for each run.

  2. Run a small-sample A/B test with none, low, and medium on the same task set; record the pass-rate improvement together with the additional tokens and latency.

  3. Put static system instructions, tool definitions, and reference materials in a stable prefix, and place changing user input at the end to observe cache hits.

  4. Budget first for requests with more than 272K input tokens; the official rule bills the entire request at 2x input and 1.5x output rates. Test chunking when necessary.

  5. If the model's output will trigger file writes, message sending, payments, or data deletion, add human confirmation, a sandbox, or rollback; a low unit price does not change the risk of the action.

  6. After running for a week, evaluate by the cost of each passed and human-accepted result, rather than routing directly based on the per-token price.

Raw evidence and data

  • Positioning: OpenAI describes Luna as intended for cost-sensitive, high-volume workloads, roughly corresponding to the nano tier of the earlier GPT-5 family.

  • Reasoning: none, low, medium (default), high, xhigh, max.

  • Context/output: 1,050,000-token context window, 128,000 maximum output tokens; knowledge cutoff 2026-02-16.

  • Pricing: $0.20 input, $0.02 cached input, and $1.20 output (per 1M tokens).

  • Long inputs: when input exceeds 272K tokens, the full request is billed at 2x input and 1.5x output rates.

  • Modalities: text input/output and image input; audio and video are not supported.

  • API: streaming, function calling, and structured outputs are supported; fine-tuning is not supported.

  • Responses tools: web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search are all Supported.

  • Tier 1: 500 RPM, 500,000 TPM, and a 5,000,000 batch queue limit; see the account page for higher tiers.

Applicable boundaries

  • The model catalog page does not provide Luna's accuracy, first-pass success rate, or average latency for your business; do not treat “Fast” as an SLA.

  • The 272K rule applies a multiplier to the entire request, not only the excess; long-document tasks must test chunking and caching in practice.

  • Tool support is confirmed only in the Responses API tool table; individual tool calls still require handling permissions, failures, malicious tool output, and human confirmation.

  • The page displays an alias; to reproduce an experiment, record the collection date and the actually available snapshot. The current page does not disclose an independent snapshot name.

Source excerpt or observation (compliance short quote only)

The official positioning is “designed for cost-sensitive, high-volume workloads”; it indicates the intended direction, not low risk or zero rework.

Source and dates

OpenAI Developers · Source date: Not disclosed · Edited: 2026-09-20

Read the original source
Variable checklist

No required variables

Related prompts

Get Started with OpenAI GPT-5.6 on Amazon Bedrock: Reasoning, Tool Calling, and CachingThe Builder's Guide to GPT-5.6: Luna's Model Selection, Agent Orchestration, and CachingGPT-5.6 Prompting Guide: Luna's Work Contract and Model RoutingApply Occam’s Razor: Reducing Overengineering in Luna/Codex Prompts

Related reviews

GPT-5.6 Luna Reddit Codex Quota and Cache Cost: A Hands-on MeasurementGPT-5.6 Luna vs. DeepSeek V4 Flash: Cache Hits and Real-World Task CostsGPT-5.6 Luna Benchmarks & Pricing (Public Benchmarks & Pricing)Agents on Rails: 8 Models, 21 Atomic Tasks

GPT-5.6 Luna

Use GPT-5.6 Luna in Tabbit

Run this guide in the environment listed above. Downloading does not transfer the template or establish model availability for your account.