Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaGLM-5V Turbo

GLM-5V-Turbo Zero-Shot Reproducible Independent Evaluation of Visual Creativity Scoring

Original source

arXiv / Open-source Evaluation Study

Source date2026-06-30

Tabbit curation2026-09-08

Read original

One-sentence conclusion

In an independent zero-shot visual creativity scoring experiment involving 992 AI-generated images and 1,500 hand-drawn sketches, GLM-5V-Turbo showed a pronounced moderate-to-high correlation with human ratings (r=0.57 for AI-generated images, r=0.49 for hand-drawn sketches overall, and up to r=0.66 for a specific shape); enabling chain-of-thought reasoning (Reasoning) left overall score alignment essentially unchanged (Delta=-0.002 for AI-generated images and Delta=-0.018 for sketches), while substantially improving the interpretability of multimodal perception and quality judgments.

Suitable use cases

  • Suitable tasks: Preliminary screening of UI/image visual aesthetics and creativity, zero-shot multimodal review, analysis of image-text perception and quality interpretability, and structured commentary on design proposals.

  • Unsuitable tasks: Strict quantitative grading of extremely rudimentary, highly abstract sketches (for example, when the correlation for Shape 11 falls to 0.36); human expert review is still required in such cases.

  • Applicable model version: glm-5v-turbo (accessed through the OpenRouter interface).

  • Applicable client, Agent, or API: OpenRouter API, official Scoring App (https://review-visual-eval-scoring.hf.space).

  • Recommended reasoning level and parameters: The evaluation consistently used temperature = 0; enable reasoning when interpretable review rationale is needed, and disable reasoning to reduce latency when only fast discrete scores are required.

Test environment

  • Models compared:

    1. Gemini 3 Flash (gemini-3-flash-preview, called directly through the Google AI API)

    2. Gemma 4 31B IT (gemma-4-31b-it)

    3. GPT-5.4 Mini (gpt-5.4-mini-20260317)

    4. GLM-5V-Turbo (glm-5v-turbo)

    5. Kimi K2.5 (kimi-k2.5)

    6. Qwen 3.6 Plus (qwen3.6-plus)

  • Evaluation datasets:

    • Dataset 1: 992 AI-generated images (generated from real human creative prompts).

    • Dataset 2: 1,500 hand-drawn sketches, including different geometry-inspired shapes such as Shape 4, Shape 11, and Shape 12.

    • Annotation benchmark: All datasets have ground-truth creativity ratings on a 1–5 scale independently assigned by multiple human experts.

  • Evaluation tool: Open-source reproducible application Harness (https://review-visual-eval-scoring.hf.space).

Input/configuration

  • Task prompt: Input a single image and standardized scoring instructions, requiring the model to assign a creativity score on a scale of 1–5.

  • Parameter control: All requests uniformly set temperature = 0 to maximize the reproducibility and determinism of the test results.

  • Study 1 vs Study 2: Study 1 disabled extended-reasoning; Study 2 enabled reasoning on models that support thinking (GLM-5V-Turbo, Kimi K2.5, and Qwen 3.6 Plus), recording the complete chain of thought and final score.

Results

1. Study 1: Pearson correlation coefficient (r) between zero-shot creativity scores and human ratings

ModelAI-generated images (N=992)Hand-drawn sketches Shape 4Hand-drawn sketches Shape 11Hand-drawn sketches Shape 12Hand-drawn sketches overall (N=1500)
Gemini 3 Flash0.680.700.670.670.68
Gemma 4 31B IT0.660.590.570.510.55
Kimi K2.50.640.690.580.550.60
Qwen 3.6 Plus0.640.650.500.560.57
GPT-5.4 Mini0.610.300.260.340.29
GLM-5V-Turbo0.570.660.360.420.49

2. Study 2: Effect of chain-of-thought reasoning (Reasoning ON vs OFF) on alignment with human ratings

ModelTest datasetReasoning ON (r)Reasoning OFF (r)Change (Delta)
GLM-5V-TurboAI-generated images0.5700.572-0.002
GLM-5V-TurboHand-drawn sketches0.4440.462-0.018
Kimi K2.5AI-generated images0.6480.636+0.012
Kimi K2.5Hand-drawn sketches0.5440.550-0.005
Qwen 3.6 PlusAI-generated images0.5620.637-0.075
Qwen 3.6 PlusHand-drawn sketches0.4700.535-0.065

Conclusion

  1. Strong zero-shot visual-aesthetic capability: Without any fine-tuning on human ratings, GLM-5V-Turbo demonstrated clear aesthetic judgment on both complex AI-generated images and hand-drawn sketches; its overall sketch performance (r=0.49) substantially outperformed GPT-5.4 Mini (r=0.29).

  2. Stable chain-of-thought performance: Unlike Qwen 3.6 Plus, which showed a clear decline in alignment after chain-of-thought reasoning was enabled (-0.075 / -0.065), GLM-5V-Turbo's chain of thought had minimal impact on final numerical scores (-|Delta| <= 0.018), while the generated reasoning was highly interpretable across four dimensions: Perception, Originality, Quality, and Justification.

Limitations and reproduction steps

  • Limitation: The model's correlation declines somewhat on extremely abstract or very sparsely drawn hand-drawn patterns (Shape 11) (r=0.36); its understanding of geometric topology in extremely sparse linework still has room for improvement.

  • Reproduction steps:

    1. Visit the HuggingFace Space evaluation tool: https://review-visual-eval-scoring.hf.space;

    2. Configure an OpenRouter API Key and specify the model as glm-5v-turbo;

    3. Upload the 992-image standard set or a single test image, and set temperature = 0;

    4. Enable/disable extended_reasoning separately, export the scores, and calculate Pearson r against the human benchmark.

Original evidence and data

  • The paper provides a detailed 21-page report, 9 figures/tables, and a public code repository, documenting the complete API call identifier, temperature, token length, and detailed statistical data.

  • Figure 4 and Figure 5 in the paper present the original sentence-by-sentence chains of thought in full for GLM-5V-Turbo's image and hand-drawn sketch evaluations, confirming the completeness of its perception and rationale-generation structure.

Scope boundaries

  • This evaluation focuses on visual creativity and aesthetic review; it does not directly reflect code-generation or system-tool-calling success rates.

  • The evaluation is based on snapshot versions from April–June 2026; quantized deployments by different providers may produce minor score fluctuations.

Source excerpt or observation (for compliant short quotation only)

Paper conclusion: “For two of the three models (Kimi K2.5 and GLM-5v Turbo) reasoning was essentially neutral on both datasets... making model evaluations interpretable—showing what they attend to, how they balance originality vs. quality, and how they justify their ratings.” (arXiv:2606.29672 Study 2).

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GLM-5V Turbo

Use and compare models in Tabbit

GLM-5V Turbo

Related reviews

MediaPrimeAIcenter2026-04-02

GLM-5V-Turbo: Design-to-Code Benchmark and Task Boundaries

CommunityReddit r/ZaiGLM2026-06-18

GLM-5V-Turbo Reddit: Tool-Calling and Vision Failures in the Field

MediaarXiv / Z.AI & Tsinghua University2026-05-12

GLM-5V-Turbo Official Technical Report: Native Multimodal Agent Benchmarks and Hierarchical Optimization Architecture

GLM-5V Turbo

Related prompts

MediaZ.AI Developer Documentation

GLM-5V-Turbo Visual Localization and Design Mockup Recreation Prompt

MediaPrimeAIcenter2026-04-02

GLM-5V-Turbo: Vision-to-Code and OpenClaw Workflow

MediaarXiv / Z.AI & Tsinghua University2026-05-12

GLM-5V-Turbo Official Agent Framework Integration and Full-Stack Web Replication Workflow

CommunityX.com2026-06-24

GLM-5V-Turbo OpenCode Visual Delegation and Multi-Round Coding Workflow