Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaGLM-5.1

GLM-5.1: Artificial Analysis Independent Intelligence Index and Inference Throughput Benchmark

Original source

Artificial Analysis

AuthorArtificial Analysis Research

Source date2026-04-07

Tabbit curation2026-08-20

Read original

One-sentence takeaway

In third-party independent benchmark evaluations, GLM-5.1 (Reasoning) scored 41 on the Intelligence Index with a throughput of 82.7 tokens/s — placing it in the top 20% of its class and demonstrating high intelligence alongside fast generation speeds, though it tends to be more verbose and relatively expensive compared to other open-weight models.

Use cases

  • Suitable tasks: Agentic interactions, complex terminal-based coding tasks, and structured reasoning where both deep reasoning and high throughput speed are required.

  • Unsuitable tasks: Ultra-high-concurrency, budget-constrained scenarios with extreme sensitivity to per-token output costs or a strong preference for ultra-concise answers.

  • Applicable model version: GLM-5.1 (Reasoning).

  • Applicable clients, agents, or APIs: First-party Z.ai API and 8 major third-party API providers.

  • Recommended reasoning tier and parameters: Official default reasoning tier; enabling Prompt Caching is strongly recommended (delivering up to an 81% cache discount).

Test environment and input/configuration

  • Evaluation framework version: Artificial Analysis Intelligence Index v4.1.1.

  • Composite evaluation subsets: Integrates 9 independent benchmarks in total:

    1. GDPval-AA v2 (real-world work tasks)

    2. 𝜏³-Banking (tool calling and banking operations)

    3. Terminal-Bench v2.1 (agentic coding and terminal operations)

    4. SciCode (scientific computing coding)

    5. Humanity's Last Exam (HLE) (frontier reasoning and knowledge)

    6. GPQA Diamond (scientific logical reasoning)

    7. CritPt (deep physics reasoning)

    8. AA-Omniscience (knowledge reliability and non-hallucination rate)

    9. AA-LCR (long-context reasoning)

  • Model specifications: 744B total parameters, 40B active parameters per token (MoE architecture), and a 200k-token context window.

Results

Evaluation MetricGLM-5.1 (Reasoning)Peer MedianPeer Rank / Tier
Intelligence Index4127#19 / 107 (Tier 4/4)
Output Speed82.7 tokens/s67.0 tokens/s#20 / 107 (Tier 3/4)
Input Price$1.39 / 1M$0.30 / 1MAbove average
Output Price$4.40 / 1M$1.20 / 1MAbove average
Prompt Cache Discount81% ($0.30 / 1M)—Excellent
Evaluation Token Consumption (Verbosity)120M tokens100M tokensAbove average (Verbose)
Blended Cost per Task$0.30 / task—#25 / 107 (Tier 3/4)

Conclusion

Independent data from Artificial Analysis confirms that GLM-5.1 ranks among the leading tier of frontier large language models (scoring 41, well above the peer median of 27). Its output speed of 82.7 tokens/s delivers a responsive interactive user experience during long-horizon reasoning. However, it exhibits a tendency toward more verbose outputs (120M tokens evaluated), and its API pricing is on the higher side compared to other open-weight derived models. In production deployments, taking full advantage of Prompt Caching is strongly recommended to optimize inference costs.

Limitations

  • This evaluation reflects a composite weighted index for Reasoning mode rather than an isolated stress test for a single domain.

  • Pricing metrics are based on official list prices and provider medians at the time of testing; actual expenses may vary across API providers.

  • While the 200k context window comfortably accommodates the vast majority of tasks, it still has an upper bound compared to models offering 1M-token windows.

Reproduction steps

  1. Standardize input prompts in accordance with the Artificial Analysis benchmark protocol.

  2. Connect to individual API provider endpoints and record Time to First Token (TTFT) and token generation speed (tokens/s).

  3. Execute the Terminal-Bench v2.1, SciCode, and GPQA Diamond subsets, comparing weighted scores against output token lengths.

  4. Compare end-to-end API billing costs with Prompt Caching enabled versus disabled.

Original evidence and data

The official Artificial Analysis leaderboard publicly records GLM-5.1's Intelligence Index score of 41, output speed of 82.7 tokens/s, input pricing of $1.39 / 1M tokens, output pricing of $4.40 / 1M tokens, and its MoE architectural specifications featuring 744B total parameters with 40B activated parameters per token.

Source excerpt or observation (compliant short quote only)

Artificial Analysis notes in its summary: "GLM-5.1 (Reasoning) is amongst the leading models in intelligence, but particularly expensive when comparing to other open weight models of similar size. It's also faster than average, however somewhat verbose."

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GLM-5.1

Use and compare models in Tabbit

GLM-5.1

Related reviews

OfficialZ.ai2026-04-07

GLM-5.1: Z.ai's Official Long-Horizon Engineering Benchmarks and Reproduction Conditions

MediaSerenities AI2026-03-29

GLM-5.1: Serenities AI's Self-Reported Benchmarks and the Boundaries of Independent Validation

CommunityReddit r/LocalLLM

GLM-5.1: Reddit LocalLLM Real-World Coding and Context Experience

CommunityReddit r/opencodeCLI2026-05-15

GLM-5.1: OpenCode Three-Model Industrial Webpage Benchmark and Real-World Capability Boundaries

GLM-5.1

Related prompts

MediaZ.AI Developer Document / Z.ai2026-04-07

GLM-5.1: Long-horizon Agent and Claude Code Configuration