Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Ox Alpha · Community source · Independent measurement

Independent LiveCodeBench v6: Ox Alpha's Raw Pass@1

Under a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceIndependent measurementEdited 2026-09-20

Test conditions

Model/version
Ox Alpha
Source
https://www.reddit.com/r/LLMDevs/comments/1vv4hmb/ox_alpha_livecodebench_v6/
Collection/review
2026-09-20; the dynamic source was not reopened
Method and sample
Under a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.

Key data and applicable tasks

One-sentence takeaway

Under a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.

Test environment

  • Entry point: OpenRouter stealth/ox-alpha, the free preview available on 2026-08-21.

  • Dataset: LiveCodeBench code_generation_lite test6.jsonl, release_v6, with 175 problems.

  • Decoding: Greedy, temperature=0, one attempt per problem.

  • Model input: Problem statement + starter code, with an instruction to return a Python code block.

  • Execution: A new subprocess for each test case; stdin/stdout comparison with JSON normalization; function problems used argument parsing and a function-call wrapper.

  • Hardware: i5 / 8 GB RAM / no GPU; inference was 100% remote.

  • Reliability: API errors used exponential backoff, and each problem had a checkpoint so the run could resume.

Raw results

DifficultyPassedProblemsPass@1
Easy224351.2%
Medium165230.8%
Hard118013.8%
Overall4917528.0%
  • Generation failures: 0; all 175 attempts produced executable code.

  • Online raw report: https://xnasarx.github.io/ox-alpha-benchmarks/results/report.html

  • Reproducible experiment repository: https://github.com/xnasarx/ox-alpha-benchmarks

  • Raw per-problem results: results/lcb/latest.json in the repository.

Reproduction steps

  1. Clone the repository and pin the commit; inspect README.md, scripts/, oxbench/, and results/.

  2. Obtain the same LiveCodeBench release_v6 data and fix the problem order, prompt wrapper, temperature=0, and one attempt per problem.

  3. Request stealth/ox-alpha through OpenRouter, saving the model ID, provider, request time, response, errors, and checkpoint.

  4. Run the repository evaluator locally against the hidden tests, outputting per-problem passed/failed status, difficulty aggregates, and generation-failure count.

  5. When comparing with a known-version model, use the same problem set, wrapper, timeout, decoding, and execution environment.

Conclusion and applicability boundary

This is a more reproducible raw capability baseline than simply observing that the model “looks able to code”: it supports a 28.0% single-turn coding result for Ox Alpha under this protocol and shows a 13.8% pass rate on Hard problems. It does not represent real-world development ability with tools, long context, iterative repair, or an agent harness.

Limitations

  • This is one single-turn greedy result; it cannot estimate Pass@k or the sampling distribution.

  • The repository displays DeepSeek vendor results alongside it, but those figures do not use the same experimental protocol and cannot form a strict ranking.

  • The 20-second test timeout, temperature, and no-scaffold setup are suitable for reproducing the raw baseline, not necessarily the best-use configuration.

What this supports

  • Under a single-turn code-generation protocol with no agent, tools, scaffold, or sampling and with temperature=0, the experimental repository reports that Ox Alpha scored 49/175, or Pass@1=28.0%, on LiveCodeBench release_v6, with lower pass rates as difficulty increased.

What this does not support

  • “Independent LiveCodeBench v6: Ox Alpha's Raw Pass@1” lacks a fully reproducible harness, repeats, or current-version snapshot (source date the source date); it cannot generalize to a unified rank, current price, or production performance.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Reddit + GitHub experimental repository · u/nasarulislam; repository author xnasarx · Original publication date 2026-08-22 · Site edit date 2026-09-20

Open original source

Ox Alpha

Compare Ox Alpha in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Ox Alpha Explained: From Stealth Preview to GLM-5.3-Flash

Ox Alpha was the anonymous name for Z.ai GLM-5.3-Flash. Here are the verified specs, access boundaries, preview timeline and safe testing decision.

Related reviews

Same Image, Same Prompt, Three Harnesses: Harness Effects on Ox AlphaLeMi compared the DeepSeek, OMP, and OpenCode frameworks using the same Ox Alpha, the same *Interstellar* Ranger RF-31D reference image, and the same prompt, suggesting that a harness may substantially affect results, but the post does not publish a quantitative score.OpenRouter Record: Ox Alpha Specs, Availability, and Anonymous ProviderOpenRouter describes Ox Alpha as a reasoning model for coding, sustained agentic work, and production workloads, listing free access, a 1,048,576-token context, up to 131,072 output tokens, and text/image/video input; the provider is only named Stealth, the developer remains anonymous, and the page publishes no reproducible capability benchmark.OpenCode Official Observation: 26T Ox Alpha Tokens in Four DaysOpenCode reports that Ox Alpha processed 26T tokens in four days, showing heavy real-world use of the preview but saying nothing by itself about model quality, individual quotas, or availability.OpenCode Go Entry: Free Period and Load FeedbackOpenCode announced Ox Alpha on OpenCode Go for six days of near-unlimited free use outside Go usage; public replies also report mid-run stops, roughly 20 tokens/s, and overload, so convenience and service stability must be evaluated separately.Ox Alpha on OpenCode: Long Context and Free Preview ConfigurationOpenCode presented Ox Alpha as a one-week free stealth preview with 1M context, multimodality, and zero data retention, making it useful for long-task prototypes but not a long-term pricing or SLA commitment.OpenCode 1.18.21: Automatic Retries for Ox Alpha StopsOpenCode recommends upgrading to 1.18.21 when Ox Alpha produces network errors; the release automatically retries unknown stops, improving client resilience but not repairing provider outages or rate limits.Ox Alpha Same-Session Typecheck Audit and Custom-Instruction WritebackWhen Ox Alpha continues to report errors after repeated typechecks in OpenCode, ask it to audit the errors in the same session, then write verified repair principles back into custom instructions to create a project-specific feedback loopCustom-language long-task workflowGive Ox Alpha documentation for a custom language that cannot be in its training data, then implement the game and language feature in separate stages to test document reading, sustained coding, and regression verification.