Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
CommunityOx Alpha

12-Prompt Stylometry Fingerprint Study: Ox Alpha's Similarity to GLM 5.3

Original source

Reddit + GitHub experimental repository

Authoru/Physical-Row960; repository author Kaiwen Du (ItsKaiwenDu)

Source date2026-08-26

Tabbit curation2026-08-27

Read original

One-sentence takeaway

Across 11 matched prompts, 7 reference models, and a deterministic 460-feature stylometry protocol, Ox Alpha was closest to GLM 5.3 on every prompt, but this indicates stylistic similarity under the test conditions rather than the identity of the weights or developer.

Test environment

  • Entry point: OpenRouter Chat Completions API, 2026-08-25 (America/Los_Angeles).

  • Target: stealth/ox-alpha, with 11 matchable prompts; p12 was excluded from analysis because of upstream rate limiting/stalling.

  • Reference models: GLM 5.3, GLM 5.2, GLM 5, MiMo V2.5, DeepSeek V4 Flash, Gemini 3.7 Flash, and MiniMax M3.

  • Data scale: One output per source model per prompt; the balanced intersection contained 88 analysis documents, while the repository preserved 95 successfully collected final answers.

  • Features: 256 hashed character n-grams, 136 function-word rates, 23 structural features, 21 discourse-marker rates, 13 punctuation features, and 11 morphological features, for 460 total features.

  • Controls: Feature scaling used reference models only; the main prompt effect was removed by prompt; Ox Alpha was not included in scaling, prompt means, or the training centroid.

Raw results

RankReference modelMean distanceBootstrap winnerPrompt votes
1GLM 5.31.8094100.0%11/11
2GLM 5.21.93630.0%0/11
3Gemini 3.7 Flash1.99950.0%0/11
4GLM 52.00360.0%0/11
5MiMo V2.52.01160.0%0/11
6DeepSeek V4 Flash2.02990.0%0/11
7MiniMax M32.04230.0%0/11
  • GLM 5.3 won 100% of 4,000 prompt-level bootstrap resamples.

  • Leave-one-prompt-out accuracy for known models: 72.7%.

  • The distance gap between GLM 5.3 and the next-closest reference, GLM 5.2, was about 6.6%.

  • Experiment repository: https://github.com/ItsKaiwenDu/Ox-Alpha-Stylometry

  • Prompt battery: https://github.com/ItsKaiwenDu/Ox-Alpha-Stylometry/blob/main/prompts.md

  • Raw answers and prediction table: data/raw/ and results/predictions.csv in the repository.

Reproduction steps

  1. Pin the repository commit, create a new session for each prompt listed in prompts.md, and save only the final answer.

  2. Collect outputs from the reference models and Ox Alpha using the same OpenRouter model IDs, model list, and request settings.

  3. Run analyze.py validate on the data, then run analyze.py run to generate the report, distance table, confusion matrix, and images.

  4. Record routing provider, failures/rate limits, length truncation, and reasoning-parameter differences; do not remove anomalous samples.

  5. Use the pre-written p13–p30 prompts as a true held-out confirmation instead of repeating only the 11 prompts on which the result was observed.

Conclusion and applicability boundary

The evidence supports the statement that “under this candidate set and stylometric feature protocol, Ox Alpha exhibits GLM-5.3-like writing behavior.” It does not support “Ox Alpha has been confirmed as GLM 5.3/5.4,” because style can be affected by system prompts, post-processing, shared data, post-training, and the serving stack.

Limitations

  • The sample contains only 11 matched prompts, with one random generation per model/prompt.

  • Reasoning parameters were not fully consistent across reference models because provider support differed; some models used native/default reasoning.

  • The reference-model set is incomplete, and known-model validation accuracy was only 72.7%.

  • p13–p30 had not yet been completed, so the current result remains a screening study.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Ox Alpha

Use and compare models in Tabbit

Ox Alpha

Related reviews

OfficialOpenRouter2026-08-21

OpenRouter Record: Ox Alpha Specs, Availability, and Anonymous Provider

CommunityX2026-08-25

Cline Test: Ox Alpha and Fable Both Fixed a Real Repository Bug, with About 3x Less Output

CommunityX2026-08-25

OpenCode Official Observation: 26T Ox Alpha Tokens in Four Days

CommunityX2026-08-21

OpenCode Go Entry: Free Period and Load Feedback

Ox Alpha

Related prompts

CommunityReddit2026-08-21

SVG Visual Consistency Smoke Test: A Dragon Riding a Bicycle

CommunityReddit2026-08-22

Custom Language to Platform Game: A Long-Task Workflow

CommunityReddit2026-08-25

OpenCode Stalls and Upstream Errors: A Four-Step Troubleshooting Workflow

CommunityX2026-08-21

Ox Alpha on OpenCode: Long Context and Free Preview Configuration