Across 11 matched prompts, 7 reference models, and a deterministic 460-feature stylometry protocol, Ox Alpha was closest to GLM 5.3 on every prompt, but this indicates stylistic similarity under the test conditions rather than the identity of the weights or developer.
Entry point: OpenRouter Chat Completions API, 2026-08-25 (America/Los_Angeles).
Target: stealth/ox-alpha, with 11 matchable prompts; p12 was excluded from analysis because of upstream rate limiting/stalling.
Reference models: GLM 5.3, GLM 5.2, GLM 5, MiMo V2.5, DeepSeek V4 Flash, Gemini 3.7 Flash, and MiniMax M3.
Data scale: One output per source model per prompt; the balanced intersection contained 88 analysis documents, while the repository preserved 95 successfully collected final answers.
Features: 256 hashed character n-grams, 136 function-word rates, 23 structural features, 21 discourse-marker rates, 13 punctuation features, and 11 morphological features, for 460 total features.
Controls: Feature scaling used reference models only; the main prompt effect was removed by prompt; Ox Alpha was not included in scaling, prompt means, or the training centroid.
| Rank | Reference model | Mean distance | Bootstrap winner | Prompt votes |
|---|---|---|---|---|
| 1 | GLM 5.3 | 1.8094 | 100.0% | 11/11 |
| 2 | GLM 5.2 | 1.9363 | 0.0% | 0/11 |
| 3 | Gemini 3.7 Flash | 1.9995 | 0.0% | 0/11 |
| 4 | GLM 5 | 2.0036 | 0.0% | 0/11 |
| 5 | MiMo V2.5 | 2.0116 | 0.0% | 0/11 |
| 6 | DeepSeek V4 Flash | 2.0299 | 0.0% | 0/11 |
| 7 | MiniMax M3 | 2.0423 | 0.0% | 0/11 |
GLM 5.3 won 100% of 4,000 prompt-level bootstrap resamples.
Leave-one-prompt-out accuracy for known models: 72.7%.
The distance gap between GLM 5.3 and the next-closest reference, GLM 5.2, was about 6.6%.
Experiment repository: https://github.com/ItsKaiwenDu/Ox-Alpha-Stylometry
Prompt battery: https://github.com/ItsKaiwenDu/Ox-Alpha-Stylometry/blob/main/prompts.md
Raw answers and prediction table: data/raw/ and results/predictions.csv in the repository.
Pin the repository commit, create a new session for each prompt listed in prompts.md, and save only the final answer.
Collect outputs from the reference models and Ox Alpha using the same OpenRouter model IDs, model list, and request settings.
Run analyze.py validate on the data, then run analyze.py run to generate the report, distance table, confusion matrix, and images.
Record routing provider, failures/rate limits, length truncation, and reasoning-parameter differences; do not remove anomalous samples.
Use the pre-written p13–p30 prompts as a true held-out confirmation instead of repeating only the 11 prompts on which the result was observed.
The evidence supports the statement that “under this candidate set and stylometric feature protocol, Ox Alpha exhibits GLM-5.3-like writing behavior.” It does not support “Ox Alpha has been confirmed as GLM 5.3/5.4,” because style can be affected by system prompts, post-processing, shared data, post-training, and the serving stack.
The sample contains only 11 matched prompts, with one random generation per model/prompt.
Reasoning parameters were not fully consistent across reference models because provider support differed; some models used native/default reasoning.
The reference-model set is incomplete, and known-model validation accuracy was only 72.7%.
p13–p30 had not yet been completed, so the current result remains a screening study.
Ox Alpha