Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GLM-5.2 · Media / benchmark · Editorial analysis

Semgrep IDOR Benchmark: GLM-5.2 Results with a Prompt-Only Setup in Security Code Auditing

Semgrep’s 2026-06-22 IDOR benchmark held dataset, evaluation, and prompt constant: GLM-5.2 reached 39% F1 in a Pydantic AI prompt-only harness at about $0.17 per vulnerability; this is not a general cyber score.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Conditions
Version GLM-5.2; Semgrep IDOR dataset of real open-source applications; same prompt/Pydantic AI harness; single configuration, 39% F1 and about $0.17/vulnerability, not a large rerun.

Key data and applicable tasks

Test environment

  • Task: Detecting IDOR (insecure direct object references), using real open-source applications from a dataset used in earlier research.

  • Metrics: Calculate precision, recall, and F1 from known true positives, and record the cost of each true positive.

  • Fixed factors: The same IDOR dataset, evaluation method, and system prompt.

  • Variables: The model and the harness. GLM-5.2, MiniMax M3, and Kimi K2.7 Code received only the prompt and codebase in a simple Pydantic AI harness; Claude Code used the Claude Code SDK; Semgrep Multimodal used a separate dedicated harness for endpoint enumeration and targeted context.

Inputs / configuration

  • Model: GLM-5.2; Pydantic AI; a unified IDOR system prompt.

  • Publicly described prompting strategy: Provide a search strategy and the key points for identifying IDOR, but no endpoint-discovery scaffolding.

  • Run scope: One dataset / one experiment; the authors explicitly state that this is a harness comparison experiment, not a general model leaderboard.

Results

RankConfigurationHarnessF1
1Semgrep Multimodal + GPT-5.5Semgrep-specific61%
2Semgrep Multimodal + Opus 4.8Semgrep-specific53%
3GLM-5.2Pydantic AI (prompt only)39%
4Claude Code + Opus 4.6Claude Code SDK37%
5Claude Code + Opus 4.8/4.7Claude Code SDK28%
6MiniMax M3Pydantic AI (prompt only)23%
7Kimi K2.7 CodePydantic AI (prompt only)22%
8GPT-5.5Codex20%
10DeepSeek V4Pydantic AI (prompt only)17%
  • The cost for GLM-5.2 to find each true positive was approximately $0.17.

  • GLM-5.2's F1 was 7 percentage points higher than Claude Code's 32%, but it was still below Semgrep's dedicated pipeline with endpoint discovery.

Conclusion

With the same minimal prompt and a simple harness, GLM-5.2 achieved a higher F1 on IDOR detection than the Claude Code configuration tested here, reaching a usable result at lower cost. The larger gap came from whether the harness provided endpoint enumeration and targeted context, so model selection cannot be separated from the workflow scaffolding.

Limitations

  • One vulnerability category, one limited dataset, and one run; the authors explicitly note that other vulnerability types, such as SSRF, may produce a different ranking.

  • GLM-5.2 and Semgrep Multimodal did not use the same harness, so 39% and 61% cannot be treated as a pure model difference.

  • The blog does not publish the complete dataset, per-sample predictions, random seed, or full prompts; it is suitable for reviewing the method, but does not constitute a fully reproducible out-of-the-box setup.

Reproduction steps

  1. Obtain the same public IDOR dataset, and fix Pydantic AI, the model SDK, timeouts, retries, and output parsing.

  2. Use the unified system prompt to run GLM-5.2 and the other models separately, retaining every vulnerability judgment and cost log.

  3. Calculate precision, recall, and F1 from the known true positives; then add an endpoint-enumeration harness to separate the model contribution from the scaffolding contribution.

  4. Re-test on a second vulnerability category and across multiple random runs; do not generalize a one-off result into a general security ranking.

Original evidence and data

  • The Semgrep article reports a 39% F1 for GLM-5.2, 32% for Claude Code (in the body text), and a cost of approximately $0.17 per true positive.

  • The original also lists MiniMax M3 at 23%, Kimi K2.7 Code at 22%, GPT-5.5 Codex at 20%, and DeepSeek V4 at 17%, and explicitly states that the open-source models did not receive endpoint-discovery scaffolding.

Source excerpt or observation (compliant short quotation only)

  • The original frames the question as “how much do model capabilities and harness capabilities each contribute,” which is more useful for guiding agent design than a single leaderboard.

What this supports

  • Supports observing prompt-only versus harness effects on IDOR F1 and cost.

What this does not support

  • Does not generalize one IDOR result to SSRF, production audits, or all repositories.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Semgrep official blog · Semgrep Security Research and Engineering (author: Pablo Estrada) · Original publication date 2026-07 · Site edit date 2026-09-20

Open original source

GLM-5.2

Compare GLM-5.2 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GLM-5.2: What It Is, What It Costs, and Where It Fits

A sourced GLM-5.2 overview covering the June 2026 release, 1M context, open-weight deployment, API pricing boundaries, coding evidence and a safer pilot path.

Related reviews

NIST CAISI's Independent Capability Assessment of Z.ai GLM-5.2NIST CAISI published its assessment on 2026-07-17 after completing it on 2026-07-08: GLM-5.2 was similar to GPT-5.2 overall and Opus 4.6 on cyber capability, while safeguards were mixed for agentic exploits and biological questions.GLM-5.2 Official Release Notes and Complete Benchmark Table (Z.ai Blog)Z.ai’s 2026-06-16 release positions GLM-5.2 as a 1M-context long-horizon flagship and reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-Bench Pro; it also discloses training-stage reward-hacking risk.Reddit Blind Code Review: GLM-5.2's Production-Readiness Score and Multi-Judge RecheckA Reddit VPS Manager blind review compared five models under one specification; Qwen 3.7 Plus first used a fixed 25-point rubric, followed by GPT Codex and Gemini 3.1 Pro rechecks; the sample is one project.Evening-Truth's Complaints About Z.AI Coding Plan Response Quality and Quantization SuspicionsEvening-Truth created a dedicated page in the prompt library to complain about the response quality of the Z.AI Coding Plan.GLM-5.2 Official Documentation: Overview and API Quick Start (docs.z.ai)The official standard integration configuration for GLM-5.2 is: model name `glm-5.2`, a 1M context window / 128K maximum output, `thinking.type: enabled` + `reasoning_effort: max`, and `temperature: 1.0`. You can copy the curl / Python examples directly to make your first call and review the typical use cases identified by the official documentation..GLM-5.2 Thinking Mode Configuration: Default Thinking / Interleaved Thinking / Preserved Thinking / Turn-level Thinking (Official)The official documentation states that thinking is enabled by default for GLM-5.2 (as with GLM-5.1/5/4.7), and provides four thinking modes: default thinking, interleaved thinking (thinking between tool calls), preserved thinking (retaining reasoning content across turns with `clear_thinking: false`), and turn-level thinking (an independent switch for each turn). It also highlights a key constraint for Agent integrations: historical `reasoning_content` must be returned unchanged..Official Configuration Guide for Migrating from GLM-5.1 / GLM-5 / GLM-4.x to GLM-5.2The official GLM-5.2 migration checklist and parameter configuration: change the model ID to `glm-5.2`; use the default `temperature` of 1.0 or default `top_p` of 0.95 (tune only one of the two); enable thinking by default; use `high` or `max` for `reasoning_effort`; configure streaming and streaming tool calls (`stream=true` + `tool_stream=true`) as specified by the official guidance; and use the included Python migration example directly..Using GLM-5.2 (zai-glm-5-2) Through Mistral: Third-Party Hosting Configuration and PricingMistral now hosts GLM-5.2 as a third-party open model (Public Preview, model ID `zai-glm-5-2`, 1M context / 128k output, with no modifications), so it can be accessed directly across the Mistral ecosystem (including Vibe CLI) using that ID, at $1.4 / $0.14 (cached input) / $4.4 (output) per million tokens..