Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GLM-5.3 · Media / benchmark · Personal experience

Hands-on GLM-5.3: The Strongest Model in Its Size Class Is Back on Top After a Week of Fierce Competition (APPSO/Tencent News)

The hands-on report gives concrete web-generation and tool-loop observations, not a benchmark ranking.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkPersonal experienceEdited 2026-09-20

Test conditions

Test and source boundary
The author had early access and described tasks and clients; no reproducible controlled comparison.
Model and version
GLM-5.3; do not merge with GLM-5.2, other models, other reasoning tiers, or other harnesses
Collection date
2026-08-18; the original page was not reopened this round, so dynamic facts remain unverified

Key data and applicable tasks

Core content summary

APPSO received early access to GLM-5.3 for testing and evaluated it on tasks including 3D web generation, web application development, and macOS utility development. It described the model as “the strongest model in its size class and the first choice for developers who code.”

Background

  • The week of its release was a “fierce clash of the gods”: Grok 4.6 (boosted by Cursor data), the official release of DeepSeek V4 Pro, Gemini Flash, and GLM-5.3 all arrived in quick succession.

  • The official position: coding ability is comparable to Fable 5 and GPT-5.6 Sol; internal testers reported that it felt better than Kimi K3 and DeepSeek V4-Pro-0813.

  • Its performance across six benchmarks was solid, with GDPVal (introduced by OpenAI) taking first place; it also ranked near the top on AutomationBench and Agents' Last Exam.

  • Its parameter count is comparable to that of its predecessor (roughly a 700-billion-parameter/743B base model), unlike the “trillion-parameter” path taken by Kimi K3 (2.8T) and Qwen3.8-Max. Z.ai said: “We are efficiently advancing reinforcement learning on exactly the same base as GLM-5.2. We may still be far from reaching the intelligence ceiling of this base.”

  • Z.ai open-sourced Slime, a post-training framework, covering the complete post-training workflow for release-grade models. It has been used since GLM 4.5 and supports GLM, some Qwen and DeepSeek models, and Llama 3.

Performance on hands-on tasks

  • 3D interactive simulation of human blood circulation: The delivery was relatively rough (stick figures and stacked organs), but offered extensive customization, including view layers, physiological parameters such as heart rate and blood pressure, an optional hemorrhagic-shock scenario, and visible blood-flow directions. The assessment was: “It can be used in teaching, but it is not that attractive.”

  • Pyramid skateboard game (a DeepSeek Harness test, using GLM-5.3 + Claude Code at the highest reasoning depth): The game played well and the scene rendering was accurate. “If you do not operate it properly, you may really lose.” The downside was that the skateboard’s afterimage was too “eye-catching.”

  • Jellyfish lake 3D scene (millions of photorealistic jellyfish drifting in an emerald lake): “It is indeed quite beautiful.” The jellyfish were not rendered as simple circles as the human figures were; the prompt was quite long.

  • Claude Code compatibility issue: With automatic command approval enabled, it frequently reported “auto mode cannot determine the safety of Bash right now.” Third-party models have not yet been adapted to Claude Code’s automatic-mode mechanism. Switching to ZCode provided a better experience: like Codex, it could take screenshots, read the screen, analyze problems, and continue optimizing, while running for longer. A planet-colliding-with-Earth simulation was still repeatedly testing and checking in ZCode after more than an hour; the final effect was striking, with the crust cracking, lava erupting, and fragments forming a planetary ring.

Cybersecurity capabilities

  • Its performance on cybersecurity tasks was on par with Claude Mythos 5.

  • ExploitGym (the test OpenAI uses to attack Hugging Face when evaluating unreleased models): It completed 130 of 898 questions within the six-hour limit.

  • A public cybersecurity disclosure ledger (cvd.z.ai) records the model’s findings. The oldest vulnerability it found can be traced back roughly 40 years. “It had not been found for 40 years, and GLM-5.3 found it as soon as it took action.”

Access and pricing

  • It is now available in Zcode (Z.ai’s official agent application) and AutoClaw (a productivity tool), with full access for GLM Coding Plan users and subscriptions open. A reset occurred once on the afternoon of the release day.

  • Third-party platforms including WorkBuddy, QwenWork, and TraeWork opened early access. The API is expected to open next Tuesday, and the full weights are expected to be open-sourced within two weeks.

  • The API pricing on the official website is still listed for GLM-5.2; at the same size, GLM-5.3 is estimated not to change much in price.

Assessment

  • “The significance of GLM-5.3 … is that it has returned to the position that originally belonged to it: the strongest model in its size class and the first choice for developers who code.”

  • “K3 uses 2.8T [parameters], DeepSeek V4 uses 1.6T, while GLM 5.2 uses less than half as many parameters and achieves the same degree of intelligence.”

  • At a broader level, GLM-5.3 and this succession of new models are making the “shelf life of the strongest large model” shorter and shorter (GLM was overtaken by Kimi after 20 days, Kimi was overtaken by DeepSeek after 20 days, and so on).

What this supports

  • The hands-on report gives concrete web-generation and tool-loop observations, not a benchmark ranking. under the stated source conditions only.

What this does not support

  • Does not support extending early-access task observations into a reproducible model ranking.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Tencent News (republished from the official APPSO account) · APPSO · Original publication date 2026-08-14 · Site edit date 2026-09-20

Open original source

GLM-5.3

Compare GLM-5.3 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GLM-5.3 Explained: What Changed from GLM-5.2

GLM-5.3 keeps the GLM-5.2 base but adds post-training for longer coding and agent tasks. Compare the changes, access paths, costs, and open risks.

Related reviews

GLM-5.3 Review: Advanced Cybersecurity Capabilities and Coding Gains (VentureBeat)Separate cyber, coding, and migration claims in the launch-day report; vendor scores are not independent retests.GLM-5.3 Independent Benchmark: 91.25% on KingBench 3, Taking the Top Spot (MindStudio)A fixed prompt set supplies a bounded outside reference, not repeated retesting.GLM-5.3: BenchLM's Source-Verifiable Benchmark Ledger and "Not Ranked" ConclusionExact-source rows trace provider numbers; without retesting they should not become an overall rank.Z.ai Official Technical Blog: Frontier Coding and Emergent Cybersecurity Capabilities (Z.ai)The official release supports launch claims and conditional benchmark records, not a universal first-place conclusion.Compare models with one fixed Hermes taskReuse one visual task in a fixed Agent harness and separate framework effects from model effects.Build staged coding tasks with explicit contextTurn project context, goals, constraints, and acceptance criteria into a staged coding task.Plan before editing in ZCodeIn ZCode, inspect the project and approve a plan before using a small task to verify the edit-and-test loop.Separate API, benchmarks, and weight statusSeparate API, Coding Plan, benchmark conditions, and weight status instead of presenting an open-weight promise as a download.