Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GLM-5.1 · Media / benchmark · Editorial analysis

GLM-5.1: Artificial Analysis Independent Intelligence Index and Inference Throughput Benchmark

The Artificial Analysis GLM-5.1 Reasoning page collected 2026-08-20 records Intelligence Index 41 and 82.7 tokens/s, while noting verbosity and relatively high cost; this is an aggregated platform index.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Conditions
Version GLM-5.1 Reasoning; Artificial Analysis collected 2026-08-20; index aggregates evaluations, with full items, provider, repeats, and hardware incomplete.

Key data and applicable tasks

One-sentence takeaway

In third-party independent benchmark evaluations, GLM-5.1 (Reasoning) scored 41 on the Intelligence Index with a throughput of 82.7 tokens/s — placing it in the top 20% of its class and demonstrating high intelligence alongside fast generation speeds, though it tends to be more verbose and relatively expensive compared to other open-weight models.

Use cases

  • Suitable tasks: Agentic interactions, complex terminal-based coding tasks, and structured reasoning where both deep reasoning and high throughput speed are required.

  • Unsuitable tasks: Ultra-high-concurrency, budget-constrained scenarios with extreme sensitivity to per-token output costs or a strong preference for ultra-concise answers.

  • Applicable model version: GLM-5.1 (Reasoning).

  • Applicable clients, agents, or APIs: First-party Z.ai API and 8 major third-party API providers.

  • Recommended reasoning tier and parameters: Official default reasoning tier; enabling Prompt Caching is strongly recommended (delivering up to an 81% cache discount).

Test environment and input/configuration

  • Evaluation framework version: Artificial Analysis Intelligence Index v4.1.1.

  • Composite evaluation subsets: Integrates 9 independent benchmarks in total:

    1. GDPval-AA v2 (real-world work tasks)

    2. 𝜏³-Banking (tool calling and banking operations)

    3. Terminal-Bench v2.1 (agentic coding and terminal operations)

    4. SciCode (scientific computing coding)

    5. Humanity's Last Exam (HLE) (frontier reasoning and knowledge)

    6. GPQA Diamond (scientific logical reasoning)

    7. CritPt (deep physics reasoning)

    8. AA-Omniscience (knowledge reliability and non-hallucination rate)

    9. AA-LCR (long-context reasoning)

  • Model specifications: 744B total parameters, 40B active parameters per token (MoE architecture), and a 200k-token context window.

Results

Evaluation MetricGLM-5.1 (Reasoning)Peer MedianPeer Rank / Tier
Intelligence Index4127#19 / 107 (Tier 4/4)
Output Speed82.7 tokens/s67.0 tokens/s#20 / 107 (Tier 3/4)
Input Price$1.39 / 1M$0.30 / 1MAbove average
Output Price$4.40 / 1M$1.20 / 1MAbove average
Prompt Cache Discount81% ($0.30 / 1M)—Excellent
Evaluation Token Consumption (Verbosity)120M tokens100M tokensAbove average (Verbose)
Blended Cost per Task$0.30 / task—#25 / 107 (Tier 3/4)

Conclusion

Independent data from Artificial Analysis confirms that GLM-5.1 ranks among the leading tier of frontier large language models (scoring 41, well above the peer median of 27). Its output speed of 82.7 tokens/s delivers a responsive interactive user experience during long-horizon reasoning. However, it exhibits a tendency toward more verbose outputs (120M tokens evaluated), and its API pricing is on the higher side compared to other open-weight derived models. In production deployments, taking full advantage of Prompt Caching is strongly recommended to optimize inference costs.

Limitations

  • This evaluation reflects a composite weighted index for Reasoning mode rather than an isolated stress test for a single domain.

  • Pricing metrics are based on official list prices and provider medians at the time of testing; actual expenses may vary across API providers.

  • While the 200k context window comfortably accommodates the vast majority of tasks, it still has an upper bound compared to models offering 1M-token windows.

Reproduction steps

  1. Standardize input prompts in accordance with the Artificial Analysis benchmark protocol.

  2. Connect to individual API provider endpoints and record Time to First Token (TTFT) and token generation speed (tokens/s).

  3. Execute the Terminal-Bench v2.1, SciCode, and GPQA Diamond subsets, comparing weighted scores against output token lengths.

  4. Compare end-to-end API billing costs with Prompt Caching enabled versus disabled.

Original evidence and data

The official Artificial Analysis leaderboard publicly records GLM-5.1's Intelligence Index score of 41, output speed of 82.7 tokens/s, input pricing of $1.39 / 1M tokens, output pricing of $4.40 / 1M tokens, and its MoE architectural specifications featuring 744B total parameters with 40B activated parameters per token.

Source excerpt or observation (compliant short quote only)

Artificial Analysis notes in its summary: "GLM-5.1 (Reasoning) is amongst the leading models in intelligence, but particularly expensive when comparing to other open weight models of similar size. It's also faster than average, however somewhat verbose."

What this supports

  • Supports comparing intelligence, throughput, and verbosity within AA’s method.

What this does not support

  • Does not make the aggregate index fixed for a provider or production cost.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Artificial Analysis Research · Original publication date 2026-04-07 · Site edit date 2026-09-20

Open original source

GLM-5.1

Compare GLM-5.1 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GLM-5.1 Explained: Long-Horizon Agents, Access, and Cost

A sourced GLM-5.1 overview covering its 200K context, 8-hour execution claim, Z.AI pricing snapshot, deployment boundaries, and a cautious pilot path.

Related reviews

GLM-5.1: Serenities AI's Self-Reported Benchmarks and the Boundaries of Independent ValidationSerenities AI’s 2026-03-29 evaluation separates an early Claude Code self-reported 45.3 from a later SWE-Bench Pro 58.4 and warns they are not the same test; its setup must be read as reported.GLM-5.1: Z.ai's Official Long-Horizon Engineering Benchmarks and Reproduction ConditionsZ.AI’s 2026-04-07 material claims up to 8 hours of sustained execution, 58.4 on SWE-Bench Pro, and 3.6× geometric-mean speedup on KernelBench Level 3; results depend on OpenHands/Terminus/Claude Code harnesses.GLM-5.1: OpenCode Three-Model Industrial Webpage Benchmark and Real-World Capability BoundariesIn a single Reddit OpenCode industrial-dashboard comparison, GLM-5.1 produced the strongest visual UI and speed near DeepSeek V4 Pro but needed a second fix pass; Kubernetes YAML and 100k+ context exposed format/stability limits.GLM-5.1: Reddit LocalLLM Real-World Coding and Context ExperienceCommunity experiences describe GLM-5.1 as a cost-effective candidate for C++/everyday coding and long-running projects, but there is still significant disagreement over large monorepos, complex debugging, latency, and context stability. The provider and harness must be recorded..GLM-5.1: Long-horizon Agent and Claude Code ConfigurationGLM-5.1 should be configured as a “long-horizon engineering Agent”: provide ample context and output budget, clarify the role, tech stack, and acceptance criteria first, then let it loop through execution, compilation, testing, and iteration; in Claude Code, you can switch the model name directly to `GLM-5.1`..GLM-5.1: SGLang Heterogeneous Deployment and Interleaved Thinking ConfigurationLocal deployment of GLM-5.1 depends on the exact `transformers==5.3.0` version and SGLang parser configuration; coding Agent workflows must enable `Interleaved + Preserved Thinking` mode to prevent multi-turn forgetting..GLM-5.1: Claude Code Tool Discovery and System Role Compatibility WorkaroundWhen using GLM-5.1 in Claude Code or a multi-Agent framework, you must explicitly inject `tool_reference` parsing rules into the system prompt to prevent tool deadlocks, and intercept the `system` role in `messages[]` to avoid HTTP 422 errors..GLM-5.1: OpenCode Multi-Model Orchestration and Anti-Overthinking PromptEmbedding GLM-5.1 in a multi-model pipeline as a “high-value code executor,” together with a system prompt that enforces action, can effectively resolve overthinking deadlocks in Agents and YAML indentation defects..