Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GLM-5.2 · Community source · Personal experience

r/LocalLLaMA Field Test: Running GLM 5.2 (744B MoE) Locally Without a GPU with the Colibrì Engine

u/LopsidedDot4557 shared a hands-on test of running GLM 5.2 (744B MoE) on a machine without a GPU, using colibrì, a pure-C engine that streams MoE experts from disk.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Model/version
GLM-5.2; source title “r/LocalLLaMA Field Test: Running GLM 5.2 (744B MoE) Locally Without a GPU with the Colibrì Engine”, with no cross-version merge.
Task/harness
Core content summary u/LopsidedDot4557 shared a hands-on test of running GLM 5.2 (744B MoE) on a machine without a GPU, using colibrì, a pure-C engine that streams MoE experts from disk。 The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count, repeats, and raw logs remain unknown.

Key data and applicable tasks

Core content summary

u/Lopsided_Dot_4557 shared a hands-on test of running GLM 5.2 (744B MoE) on a machine without a GPU, using colibrì, a pure-C engine that streams MoE experts from disk:

Principle and configuration

  • A 744B MoE activates only a small number of experts for each token: colibrì keeps the dense portion (about 10GB) resident in RAM and streams the experts selected by the router from disk as needed.

  • The complete int4 model occupies about 370GB on disk; it does not need to fit entirely in memory.

  • Test machine: a single machine with 132GB RAM, Ubuntu 22.04, and a local NVMe drive.

Measured performance (cold start → warm-up)

StageSpeedExpert hit rateRSS
First token after cold start~0.03 tok/s~21%-
After several rounds of short prompts~0.15 tok/s~65%-
Further warm-up~0.22 tok/s~71%~113GB
  • The hit rate rises with use: the engine pins the experts that are actually routed to, so it gets faster the longer it runs.

  • Full hands-on video: https://youtu.be/jxML3S5C-8Y

  • Additional comment (Apple M5 Max 128GB): CPU/default, 1.06 tok/s (RSS 21.8GB); after warm-up with Metal cu --ram 96, 1.11→1.83 tok/s; with --ram 110, up to about 2.06 tok/s.

  • u/Sleepybear2611 (who conducted a three-day hands-on test of 754B): the rise in hit rate depends mainly on having a large amount of RAM (at RSS 113GB, about 30% of expert storage is held, and the LRU converges); on a 31GB low-memory machine, coverage can only settle at ~40–60%, because a single generation touches 38% of all experts and the working set cannot fit in the small memory; the cold-start figures match the prediction from their “bytes/token ÷ disk bandwidth” model (about 11GB/token).

A killer use case from the comments (quoted verbatim from u/trej)

"For the past month, I've been running local GLM 5.2 Q4 non-stop on my workstation rig, on repeat asking the prompt **'R… This is a necessary excerpt; read the original source for full context.

(A local, continuously running audit prompt that can be copied directly; observed result: about one round every 3.5 days, 10–15 new bugs per round, 1–2 of them hallucinations, and occasional impressive findings.)

Key points from the community discussion

  • Opposition (u/Littlepharaoh): 0.5 tok/s is too slow; the same task can be completed in a few minutes with serverless Runpod or a cheap API, at a cost of a few cents. Running locally is suitable for people whose countries prohibit international payments or who do not want to hand over their codebase.

  • Support (u/SV_SV_SV): a local 24/7 background audit does not give the code or data to anyone else, so the philosophical and privacy value is real.

  • Ecosystem: u/misanthrophiccunt mentioned that colibrì is adding DeepSeek V4 Flash support; u/Refinery73 asked whether small or medium-sized MoE versions are available—the current engine was built specifically for the GLM series (proof of concept).

Evidence highlights and scope

  • What the conclusion covers: GLM 5.2 (744B MoE) can run on a machine with no GPU using only RAM + NVMe (pure CPU), making it suitable for “slow, background, long-running” code-audit tasks; its real throughput is far below that of cloud APIs, so it is not suitable for interactive development.

  • Parameter naming: The post calls it 744B (other sources, such as the GLM-5.3 release post, call it 743B; this may be a rounding difference); after int4 quantization, it is about 370GB.

  • Boundary: This is an experience report from a single-user environment with 132GB RAM; hit rate and throughput can drop substantially on machines with less memory. “10–15 bugs/3.5 days” is an observation from one user repeatedly running a single prompt, not a controlled benchmark.

  • Reproduction: The same order of magnitude can be reproduced with the colibrì engine, GLM-5.2 int4 weights, and the hardware described above; see the original post and video for the specific parameters.

What this supports

  • Supports the source-specific observation in “r/LocalLLaMA Field Test: Running GLM 5.2 (744B MoE) Locally Without a GPU with the Colibrì Engine”: Core content summary u/LopsidedDot4557 shared a hands-on test of running GLM 5.2 (744B MoE) on a machine without a GPU, using colibrì, a pure-C engine that streams MoE experts from disk。

What this does not support

  • Does not support a general capability or production-rate claim from “r/LocalLLaMA Field Test: Running GLM 5.2 (744B MoE) Locally Without a GPU with the Colibrì Engine”; the source lacks a controlled task set, provider snapshot, and repeated independent retest.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit r/LocalLLaMA · u/LopsidedDot4557 (main post); u/SVSVSV, u/trej, u/Sleepybear2611, u/Academic-Most6214, and others (comments) · Original publication date Unknown · Site edit date 2026-09-20

Open original source

GLM-5.2

Compare GLM-5.2 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GLM-5.2: What It Is, What It Costs, and Where It Fits

A sourced GLM-5.2 overview covering the June 2026 release, 1M context, open-weight deployment, API pricing boundaries, coding evidence and a safer pilot path.

Related reviews

Reddit Blind Code Review: GLM-5.2's Production-Readiness Score and Multi-Judge RecheckA Reddit VPS Manager blind review compared five models under one specification; Qwen 3.7 Plus first used a fixed 25-point rubric, followed by GPT Codex and Gemini 3.1 Pro rechecks; the sample is one project.GLM-5.2 Official Release Notes and Complete Benchmark Table (Z.ai Blog)Z.ai’s 2026-06-16 release positions GLM-5.2 as a 1M-context long-horizon flagship and reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-Bench Pro; it also discloses training-stage reward-hacking risk.NIST CAISI's Independent Capability Assessment of Z.ai GLM-5.2NIST CAISI published its assessment on 2026-07-17 after completing it on 2026-07-08: GLM-5.2 was similar to GPT-5.2 overall and Opus 4.6 on cyber capability, while safeguards were mixed for agentic exploits and biological questions.Semgrep IDOR Benchmark: GLM-5.2 Results with a Prompt-Only Setup in Security Code AuditingSemgrep’s 2026-06-22 IDOR benchmark held dataset, evaluation, and prompt constant: GLM-5.2 reached 39% F1 in a Pydantic AI prompt-only harness at about $0.17 per vulnerability; this is not a general cyber score.GLM-5.2 Official Documentation: Overview and API Quick Start (docs.z.ai)The official standard integration configuration for GLM-5.2 is: model name `glm-5.2`, a 1M context window / 128K maximum output, `thinking.type: enabled` + `reasoning_effort: max`, and `temperature: 1.0`. You can copy the curl / Python examples directly to make your first call and review the typical use cases identified by the official documentation..GLM-5.2 Thinking Mode Configuration: Default Thinking / Interleaved Thinking / Preserved Thinking / Turn-level Thinking (Official)The official documentation states that thinking is enabled by default for GLM-5.2 (as with GLM-5.1/5/4.7), and provides four thinking modes: default thinking, interleaved thinking (thinking between tool calls), preserved thinking (retaining reasoning content across turns with `clear_thinking: false`), and turn-level thinking (an independent switch for each turn). It also highlights a key constraint for Agent integrations: historical `reasoning_content` must be returned unchanged..Official Configuration Guide for Migrating from GLM-5.1 / GLM-5 / GLM-4.x to GLM-5.2The official GLM-5.2 migration checklist and parameter configuration: change the model ID to `glm-5.2`; use the default `temperature` of 1.0 or default `top_p` of 0.95 (tune only one of the two); enable thinking by default; use `high` or `max` for `reasoning_effort`; configure streaming and streaming tool calls (`stream=true` + `tool_stream=true`) as specified by the official guidance; and use the included Python migration example directly..Using GLM-5.2 (zai-glm-5-2) Through Mistral: Third-Party Hosting Configuration and PricingMistral now hosts GLM-5.2 as a third-party open model (Public Preview, model ID `zai-glm-5-2`, 1M context / 128k output, with no modifications), so it can be accessed directly across the Mistral ecosystem (including Vibe CLI) using that ID, at $1.4 / $0.14 (cached input) / $4.4 (output) per million tokens..