As of August 21, 2026, Ox Alpha is a real but still unexplained AI model available through OpenRouter. It offers a 1,048,576-token context window, text/image/video input and temporary free access. Early community tests look unusually strong. We have also tried Ox Alpha ourselves at Botchi, and our first impression is that it is very promising. That first-hand observation is useful context, not a substitute for a controlled benchmark or production evidence. SMEs should still treat it as a promising preview—not a proven production dependency—until its identity, privacy terms, reliability, benchmarks and long-term pricing are clearer.
What changed
On August 20, OpenRouter listed a new model under the identifier stealth/ox-alpha. The page describes it as a reasoning model for coding, sustained agentic work and production workloads, including long-horizon software engineering and workflows that combine text with visual context.
The public specification is unusually ambitious:
- Context window: 1,048,576 tokens
- Maximum output: 131,072 tokens
- Inputs: text, images and video
- Output: text
- Tool use: function calling through
toolsandtool_choice - Pricing: free on OpenRouter at the time of writing
OpenRouter says it routes requests to a single third-party provider and is not the model's developer, owner or provider. That provider is shown only as “Stealth” and has chosen to remain anonymous during the preview.
OpenCode, the coding-agent platform, promoted the same model on X as free for the next week, with generous or near-unlimited usage and claimed capacity of 100 trillion tokens per day. Those are platform claims, not an independently audited quota or a guarantee that every account will receive identical access.
Why developers are paying attention
The combination matters more than any single specification.
A million-token context window can make it easier to work across a large repository, a long technical document, a collection of screenshots or a multi-step agent session without repeatedly compressing context. Image and video input also expand the range of tasks: interface inspection, visual QA, document analysis and workflows where the model needs to connect code with what a user sees.
Free access lowers the cost of experimentation. For a small business, that means a team can compare the model against its current assistant before committing budget. It does not mean the total cost of using it in production is zero: engineering time, review time, failed runs, latency, storage, monitoring and future price changes still count.
The model is also appearing at the right moment for long-running coding agents. Instead of asking for a snippet, a developer can ask an agent to inspect a repository, plan a change, use tools, test the result and revise it. That is the kind of workflow where context length, tool reliability and persistence matter more than a good single-turn answer.
What early tests actually show
The strongest claims so far come from individual users, not an official benchmark programme.
| Test or reaction | Reported result | What it does—and does not—show |
|---|---|---|
| Ben Davis on a 10-task DeepSWE subset | 8/10, reported as 80%, versus 65% for Fable and 52% for GPT-5.6 Sol | A striking early coding signal, but the sample is tiny and one near-miss involved a scoring judgement about the letter “x”. It is not a stable leaderboard result. |
| Robin Ebers' video-editing test | On clean footage, 7 of 9 finished runs were 90–97% similar to an Opus 5 High edit and roughly 7× faster | A promising custom multimodal test. Four of 20 runs failed, including three that hung, and messy footage produced less consistent results. |
| Florian Roth's threat-triage test | No real threats missed, but 45.4% of benign cases were sent for analyst review | High recall came with poor precision for that task. Roth reported Gemini was clearly ahead. This is a useful warning against generalising from one impressive workflow. |
| Patrick Collison's first reaction | “Very impressive” after trying it through an agent interface | Evidence that an experienced user found it compelling, but not a reproducible evaluation. |
The pattern is more useful than the headline score: Ox Alpha may be excellent at some long-horizon, multimodal or agentic tasks, while still being slow, inconsistent or overly cautious in others. That is exactly why a small business should test its own work rather than copy a ranking from X.
OpenRouter's public listing does not provide official capability benchmarks for Ox Alpha. The most responsible description today is promising, community-tested and unverified at frontier scale.
Who built Ox Alpha?
Nobody has confirmed it publicly.
The community has proposed Zhipu/Z.ai's GLM family, Xiaomi's MiMo and several other labs. Some researchers have pointed to tokenizer behaviour, error messages, video-token patterns and response style. Those observations may be useful clues, but they are not an announcement from the developer and they do not establish the model's architecture or training history.
For an SME buyer, the identity question matters for practical reasons rather than trivia. It affects expectations about support, jurisdiction, data processing, model continuity, documentation, incident response and what happens after the preview ends. Until the provider is named, those questions remain open.
The privacy wording needs careful reading
The official messages do not use identical language.
OpenRouter's Ox Alpha page says that prompts and completions are retained by the provider but are not used for training. OpenCode's launch post says “Zero Data Retention”. Those statements may refer to different layers or definitions, but the available public material does not explain the difference.
That is enough reason not to send confidential customer data, source code, credentials, personal data or regulated information into the preview until the applicable route and terms are clear. “Not used for training” and “not retained” are not automatically the same guarantee.
What this means for SMEs
Ox Alpha is important even if it turns out to be a short-lived preview. It shows how quickly a capable model can appear, attract real workloads and compete on context, multimodality and agentic performance before its creator is publicly known.
For SMEs, that creates an opportunity and a planning problem. The opportunity is cheaper experimentation with difficult workflows. The planning problem is dependency risk: a free anonymous model may change limits, disappear, reveal a different privacy policy or become expensive after the preview.
The sensible approach is to design the workflow so the model can be replaced. Keep prompts, test cases, evaluation criteria, tool permissions and approval steps under your control. Treat the model as one component—not as the process itself.
Botchi does not currently offer Ox Alpha, so this article is not an announcement of Botchi access or integration. The broader lesson still applies to any SME evaluating AI: choose a measurable workflow, compare models on the work that matters and put governance around the actions that can cause harm.
Frequently asked questions
Is Ox Alpha a confirmed GLM or MiMo model?
No. Those are community theories based on informal fingerprinting and observed behaviour. OpenRouter identifies the provider only as Stealth, and no lab has publicly claimed Ox Alpha in the sources reviewed for this article.
Is Ox Alpha really free?
OpenRouter currently lists zero prompt and completion-token pricing, while OpenCode advertised a roughly one-week free preview. Treat that as temporary access, not a permanent commercial offer.
Is Ox Alpha ready for production?
There is not enough public evidence to say. Its specifications and early tests are promising, but official benchmarks, long-term pricing, provider identity, retention terms and reliability expectations are still unresolved.
What should a small business test first?
Start with a bounded, reversible task such as codebase exploration, long-document analysis, visual QA or draft generation. Compare it with your current model using the same inputs and measure completion, corrections, latency, failures and operator time.
Put AI to work in your SME
If you are deciding which AI workflow to pilot—not simply which model is trending—book a call with Botchi or email hello@botchi.ai. Botchi does not provide Ox Alpha today, but we can still discuss the broader question: which AI process is worth measuring, governing and improving first?
Sources
- OpenRouter: Ox Alpha — API Pricing & Providers
- OpenRouter: Ox Alpha launch announcement on X
- OpenCode: Ox Alpha preview announcement on X
- Ben Davis: 10-task DeepSWE test on X
- Robin Ebers: custom video-editing test on X
- Florian Roth: threat-triage test on X
- AGTPinsights: benchmark caveats and community reaction on X
- Patrick Collison: first reaction on X
- OpenRouter Stealth Model Terms
