Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Ox Alpha · Community source · Independent measurement

12-Prompt Stylometry Fingerprint Study: Ox Alpha's Similarity to GLM 5.3

Across 11 matched prompts, 7 reference models, and a deterministic 460-feature stylometry protocol, Ox Alpha was closest to GLM 5.3 on every prompt, but this indicates stylistic similarity under the test conditions rather than the identity of the weights or developer.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceIndependent measurementEdited 2026-09-20

Test conditions

Model/version
Ox Alpha; exact snapshot follows the source
Source date
2026-08-26
Method/client
Source-specific public post; client and provider conditions follow the source
Review state
Dynamic source not reopened on 2026-09-20; values remain unverified

Key data and applicable tasks

One-sentence takeaway

Across 11 matched prompts, 7 reference models, and a deterministic 460-feature stylometry protocol, Ox Alpha was closest to GLM 5.3 on every prompt, but this indicates stylistic similarity under the test conditions rather than the identity of the weights or developer.

Test environment

  • Entry point: OpenRouter Chat Completions API, 2026-08-25 (America/Los_Angeles).

  • Target: stealth/ox-alpha, with 11 matchable prompts; p12 was excluded from analysis because of upstream rate limiting/stalling.

  • Reference models: GLM 5.3, GLM 5.2, GLM 5, MiMo V2.5, DeepSeek V4 Flash, Gemini 3.7 Flash, and MiniMax M3.

  • Data scale: One output per source model per prompt; the balanced intersection contained 88 analysis documents, while the repository preserved 95 successfully collected final answers.

  • Features: 256 hashed character n-grams, 136 function-word rates, 23 structural features, 21 discourse-marker rates, 13 punctuation features, and 11 morphological features, for 460 total features.

  • Controls: Feature scaling used reference models only; the main prompt effect was removed by prompt; Ox Alpha was not included in scaling, prompt means, or the training centroid.

Raw results

RankReference modelMean distanceBootstrap winnerPrompt votes
1GLM 5.31.8094100.0%11/11
2GLM 5.21.93630.0%0/11
3Gemini 3.7 Flash1.99950.0%0/11
4GLM 52.00360.0%0/11
5MiMo V2.52.01160.0%0/11
6DeepSeek V4 Flash2.02990.0%0/11
7MiniMax M32.04230.0%0/11
  • GLM 5.3 won 100% of 4,000 prompt-level bootstrap resamples.

  • Leave-one-prompt-out accuracy for known models: 72.7%.

  • The distance gap between GLM 5.3 and the next-closest reference, GLM 5.2, was about 6.6%.

  • Experiment repository: https://github.com/ItsKaiwenDu/Ox-Alpha-Stylometry

  • Prompt battery: https://github.com/ItsKaiwenDu/Ox-Alpha-Stylometry/blob/main/prompts.md

  • Raw answers and prediction table: data/raw/ and results/predictions.csv in the repository.

Reproduction steps

  1. Pin the repository commit, create a new session for each prompt listed in prompts.md, and save only the final answer.

  2. Collect outputs from the reference models and Ox Alpha using the same OpenRouter model IDs, model list, and request settings.

  3. Run analyze.py validate on the data, then run analyze.py run to generate the report, distance table, confusion matrix, and images.

  4. Record routing provider, failures/rate limits, length truncation, and reasoning-parameter differences; do not remove anomalous samples.

  5. Use the pre-written p13–p30 prompts as a true held-out confirmation instead of repeating only the 11 prompts on which the result was observed.

Conclusion and applicability boundary

The evidence supports the statement that “under this candidate set and stylometric feature protocol, Ox Alpha exhibits GLM-5.3-like writing behavior.” It does not support “Ox Alpha has been confirmed as GLM 5.3/5.4,” because style can be affected by system prompts, post-processing, shared data, post-training, and the serving stack.

Limitations

  • The sample contains only 11 matched prompts, with one random generation per model/prompt.

  • Reasoning parameters were not fully consistent across reference models because provider support differed; some models used native/default reasoning.

  • The reference-model set is incomplete, and known-model validation accuracy was only 72.7%.

  • p13–p30 had not yet been completed, so the current result remains a screening study.

What this supports

  • Across 11 matched prompts, 7 reference models, and a deterministic 460-feature stylometry protocol, Ox Alpha was closest to GLM 5.3 on every prompt, but this indicates stylistic similarity under the test conditions rather than the identity of the weights or developer.

What this does not support

  • “12-Prompt Stylometry Fingerprint Study: Ox Alpha's Similarity to GLM 5.3” lacks a fully reproducible harness, repeats, or current-version snapshot (source date the source date); it cannot generalize to a unified rank, current price, or production performance.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit + GitHub experimental repository · u/Physical-Row960; repository author Kaiwen Du (ItsKaiwenDu) · Original publication date 2026-08-26 · Site edit date 2026-09-20

Open original source

Ox Alpha

Compare Ox Alpha in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Ox Alpha Explained: From Stealth Preview to GLM-5.3-Flash

Ox Alpha was the anonymous name for Z.ai GLM-5.3-Flash. Here are the verified specs, access boundaries, preview timeline and safe testing decision.

Related reviews

Cline Test: Ox Alpha and Fable Both Fixed a Real Repository Bug, with About 3x Less OutputCline says Ox Alpha and Fable both correctly fixed one real bug in its repository; Ox Alpha used about three times fewer output tokens and repeated less reasoning. This is an efficiency signal worth retesting, not a general capability ranking.Same Image, Same Prompt, Three Harnesses: Harness Effects on Ox AlphaLeMi compared the DeepSeek, OMP, and OpenCode frameworks using the same Ox Alpha, the same *Interstellar* Ranger RF-31D reference image, and the same prompt, suggesting that a harness may substantially affect results, but the post does not publish a quantitative score.OpenRouter Record: Ox Alpha Specs, Availability, and Anonymous ProviderOpenRouter describes Ox Alpha as a reasoning model for coding, sustained agentic work, and production workloads, listing free access, a 1,048,576-token context, up to 131,072 output tokens, and text/image/video input; the provider is only named Stealth, the developer remains anonymous, and the page publishes no reproducible capability benchmark.Jonathan Turner: Ox Alpha Fingerprint Comparisons and Identity BoundariesThe article compares Ox Alpha with public GLM, Gemini, DeepSeek, Kimi, and MiMo using tokenizer behavior, video-token budgets, and server errors, strongly pointing to a GLM-family model on a Z.ai serving stack while explicitly stopping short of naming a product or developer.Ox Alpha's ZCode + OpenRouter Configuration and Single-Prompt CaseArc's reusable experience is connecting Ox Alpha to ZCode through OpenRouter and using one fixed prompt for templated frontend experiments; this shows that the entry-point configuration is reproducible, but the undisclosed detailed promptOx Alpha + Three.js Nan Lian Garden: A Single-Prompt 3D Scene Example“Build a 3D version of Nan Lian Garden with Three.js” is a short task suitable for checking spatial layout, rendering stability, and the debugging loop, but the original post publishes only a task summary, not the full prompt.SVG structure smoke testA one-line SVG prompt for a dragon riding a bicycle is a quick way to check Ox Alpha's handling of structural relationships, physical plausibility, and executable SVG code.Custom-language long-task workflowGive Ox Alpha documentation for a custom language that cannot be in its training data, then implement the game and language feature in separate stages to test document reading, sustained coding, and regression verification.