Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Prompts and workflows

GLM-5.3 · configuration

Configure three reasoning tiers across API protocols

Connect mandatory thinking, low/high/max, and three API protocols into a checkable integration path.

Source reviewed; not testedAIHubMix OpenAI-compatible API; Chat Completions, Responses, and Messages

Prerequisites and inputs

  • API key
  • model ID
  • reasoning_effort
  • tool schema

Complete templates

Editorial adaptation: auditable workflow

Tabbit editorial adaptation; not the original source prompt
Run one reversible GLM-5.3 task with these inputs:

- API key: {{API_KEY}}
- model ID: {{MODEL_ID}}
- reasoning_effort: {{REASONING_EFFORT}}
- tool schema: {{TOOL_SCHEMA}}

State the plan, environment, and acceptance criteria; get human approval, then separate source facts, model output, and items still to verify.

Replace before running: API key, model ID, reasoning_effort, tool schema

Steps

  1. Fix the provider route, model ID, and protocol; do not relabel aggregator coding-glm-5.3 as official glm-5.3.

  2. Start simple tasks at low, then use high or max for difficult Agent work; record actual output and reasoning tokens.

  3. Validate structured output and the business result after tools; never expose hidden reasoning.

Read the source research notes

Core content summary

GLM-5.3 is Z.ai's flagship model released on 2026-08-14. It uses exactly the same base as GLM-5.2 and is upgraded purely through post-training. Every "verified" conclusion in this article comes from real calls made through the AIHubMix API (using the Chat Completions, Responses, and Messages protocols) on 2026-08-14.

1. Model specification at a glance

ItemValue
Context window1M tokens (the official exact value is 1,048,576)
Maximum output128K (the measured limit is 131,072; exceeding it returns 400)
Input modalityText
ThinkingAlways on and cannot be disabled; reasoning_effort has three tiers: low / high / max (max by default)
Relationship to 5.2Same base, pure post-training; major gains in coding and long-horizon tasks + emergent cybersecurity capabilities

2. Key API differences from GLM-5.2 (most important)

ItemGLM-5.2GLM-5.3
thinking.typeenabled / disabled (can be turned off)enabled only (cannot be turned off)
reasoning_effortCompatible with 7-value mappingsThree tiers: low / high / max (max by default)
PositioningGeneral-purpose flagshipEnhanced coding and long-horizon Agent capabilities, with emergent network capabilities
  • Official migration recommendation: change applications that previously sent {"type":"disabled"} to {"type":"enabled"} and set reasoning_effort:"low".

  • Measured: sending disabled through AIHubMix still returns 200 and thinking still occurs (automatically converted according to the semantics of the official channel). If a client depends on "turning off thinking to save tokens," switch to reasoning_effort:"low".

  • Measured: an invalid enum value returns 200 and falls back to the default max; on the same arithmetic problem, thinking tokens were 27 with low versus 39 with max.

3. Calling the three protocols and reading thinking content

Chat Completions (thinking content is in the reasoning_content field, and in delta.reasoning_content for streaming):

from openai import OpenAI
client = OpenAI(base_url="https://aihubmix.com/v1", api_key="<KEY>")
completion = client.chat.completions.create(
    model="coding-glm-5.3",
    reasoning_effort="max",   # low / high / max, max by default
    extra_body={"thinking": {"type": "enabled"}},
    messages=[{"role": "user", "content": "Compute the square root of (17*23-19*11), rounded down. Digits only."}],
)
print(completion.choices[0].message.reasoning_content)
print(completion.choices[0].message.content)  # observed: "13"
  • Measured: usage.completion_tokens_details.reasoning_tokens reports thinking usage—for the same problem, low used 27 and max used 39.

Responses API (thinking content is a reasoning output item; the text is in summary_text within the summary array):

response = client.responses.create(model="coding-glm-5.3", input="What is the capital of France? City name only.")
# output item types: ["reasoning", "message"]
# reasoning item: {"type": "reasoning", "summary": [{"type": "summary_text", "text": "..."}]}
# usage.output_tokens_details.reasoning_tokens: 80
  • Measured: even without passing any reasoning parameters, a reasoning item is returned by default (no explicit opt-in is required).

Messages (Anthropic protocol) (thinking content is a native thinking content block):

client = Anthropic(api_key="<KEY>", base_url="https://aihubmix.com")
response = client.messages.create(model="coding-glm-5.3", max_tokens=4096,
    messages=[{"role": "user", "content": "What is the capital of France? City name only."}])
# content block types: ["thinking", "text"]

4. Tool calls and parallel tools

  • All three APIs were verified to work; the Responses API showed parallel tool calls within a single turn (the official declaration is supports_parallel_tool_calls: true).

  • Upstream limits: at most 128 functions in tools; the native tool_choice support is limited to auto.

  • In Chat Completions, tool_choice:"none" was measured to work (no further tool calls); in the Messages protocol, tool_choice:{type:"none"} was still observed to produce tool_use. To disable tools, remove the tools parameter directly, or use tool_choice:"none" with Chat Completions.

5. Structured output

  • response_format supports text and json_object; the upstream service does not provide a json_schema mode. When a strict schema is required, put the JSON Schema in the prompt and validate it on the client.

  • Measured: response_format={"type":"json_object"} returned valid JSON containing the requested key.

6. Other details

  • Context caching / automatic caching is available (on the AIHubMix side).

  • This article reflects the AIHubMix aggregation channel (coding-glm-5.3 is its preview route); the model ID on the official channel is glm-5.3, so keep the two distinct.

Key quotations from the original

"thinking.type no longer supports disabled — thinking cannot be turned off."

"If your client relied on 'turn off thinking to save tokens', switch to reasoning_effort: 'low'."

Source and dates

AIHubMix Blog (tutorial from an AI aggregation API provider) · Source date: 2026-08-14 · Edited: 2026-09-20

Read the original source
Variable checklist

Still to replace: 4

API keymodel IDreasoning_efforttool schema

Related prompts

Migrate GLM-5.3 thinking parametersChoose reasoning effort by task difficultyPlan before editing in ZCodeBuild staged coding tasks with explicit context

Related reviews

X (Twitter) @Rafa_Schwinger: Metal Kernel Review Task—GLM 5.3 xhigh 88/100 vs. Grok 4.6 86/100Z.ai Official Technical Blog: Frontier Coding and Emergent Cybersecurity Capabilities (Z.ai)GLM-5.3 Review: Advanced Cybersecurity Capabilities and Coding Gains (VentureBeat)GLM-5.3 Independent Benchmark: 91.25% on KingBench 3, Taking the Top Spot (MindStudio)

Read the full analysis

Overview · English

GLM-5.3 Explained: What Changed from GLM-5.2

GLM-5.3 keeps the GLM-5.2 base but adds post-training for longer coding and agent tasks. Compare the changes, access paths, costs, and open risks.

GLM-5.3

Use GLM-5.3 in Tabbit

Run this guide in the environment listed above. Downloading does not transfer the template or establish model availability for your account.