Ox Alpha Debuts on OpenRouter: Mystery 1M-Context Model Sparks Unverified Community Benchmark Claims


On August 20, 2026, developers monitoring OpenRouter and OpenCode's real-time model indexes noticed an unexpected new entry listed under the provider alias "Stealth": Ox Alpha (stealth/ox-alpha). Arriving without a corporate announcement, research whitepaper, or product marketing campaign, the model debuted offering an expansive 1,048,576-token context window, native multimodal input for text, image, and video, tool-calling support, and a promotional price tag of zero dollars. Within 48 hours of listing, developer communities began routing repository-scale refactoring tasks through the endpoint and posting their own informal benchmark numbers — unverified claims worth examining carefully, not results to take at face value.
Model Vitals
| Metric / Specification | Listed Value |
|---|---|
| Architecture / Size | Undisclosed (Estimated Frontier-Scale MoE) |
| Context Window | 1,048,576 tokens (~1.05M tokens) |
| Max Output Limit | 131,072 tokens |
| Modalities | Text, Image, Video, Structured Function / Tool Calling |
| Provider Access | OpenRouter (stealth/ox-alpha), OpenCode Zen |
| Release Date | August 20, 2026 |
| Pricing | $0.00 / 1M Input & Output (Promotional Free Preview) |
| Deployment Mode | Cloud API (Proprietary / Managed Hosting) |
Benchmark Analysis & Community Evaluations
No official technical report, model card, or benchmark disclosure accompanies Ox Alpha. Everything below comes from unverified, self-reported community testing on developer forums and social media — not from Ox Alpha's own maker (unknown), and not yet reproduced by an independent evaluator like Artificial Analysis or LM Arena. Treat every number in this section as a claim, not a confirmed result, until independent benchmarking catches up.
Software Engineering & DeepSWE Evaluation
Several developers on X and r/LocalLLaMA have posted their own small-sample DeepSWE runs against Ox Alpha, with one frequently cited unverified result showing an 80.0% pass rate on a 10-task sample — compared to the same poster's self-run 65.0% for Claude Fable 5 and 52.0% for [GPT-5.6 Sol](/models/gpt-5-6-sol) under the same scaffolding. A 10-task sample is far too small to be statistically meaningful, and no one has published the task set, transcripts, or methodology needed to reproduce it. On repository-level issue resolution, posters reported the model kept shared TypeScript definitions consistent across multi-package monorepos, though again, this is anecdotal and unverified.
1M-Token Long-Context Ingestion
With its listed 1,048,576-token context window, developers testing full-codebase ingestion have reported strong recall on long-context retrieval passes — successfully loading entire multi-module applications alongside API schema documentation in a single prompt without the attention degradation that often affects mid-context segments. These reports are informal and haven't been benchmarked against a standard long-context eval.
Inference Throughput & Serving Scale
Several OpenCode Zen users have reported streaming speeds averaging 60-85 tokens per second in their own sessions. A claim circulating in developer channels that OpenCode's infrastructure can serve "up to 100 trillion tokens per day" during the promotional period has not been traced to an official OpenCode statement or blog post — it should be treated as unconfirmed until a named source surfaces.
Community Reactions & Sourcing Disclosures
The appearance of Ox Alpha has prompted widespread discussion across developer channels:
- Attribution & Speculation: Independent developers on Reddit (r/LocalLLaMA) and developer forums noted prompt-response formatting patterns resembling modern Chinese open-weight architectures, though no lab has verified connection to the model.
- Developer Feedback: Full-stack engineers using OpenCode Zen reported successful one-shot migrations of complex Next.js App Router applications, praising the model's adherence to React Server Component boundaries.
- Enterprise Caveats: Security and compliance practitioners highlighted that the absence of a named corporate entity, data retention guarantees, or formal service level agreements means the model remains suitable primarily for non-sensitive testing and prototype development.
Try It Yourself: Structured Prompt Blueprints
Developers evaluating Ox Alpha on OpenRouter or OpenCode can test its long-context capabilities with these structured prompts:
Prompt 1: Full-Repository Architecture Audit (1M Context)
You are a principal software architect.
I am providing the entire source code of a multi-package TypeScript Monorepo in this prompt.
Tasks:
1. Map all cross-package circular dependencies and architectural boundary violations.
2. Identify uncleaned event listeners or memory leaks in asynchronous worker modules.
3. Propose a step-by-step refactoring plan to consolidate database client connections.
4. Output updated source files for the affected modules only.
Prompt 2: Autonomous Terminal Debugger & Tool Loop
You are an autonomous software engineering agent with access to bash tools.
Given the failing test suite output below:
- Analyze the stack trace and isolate root causes in asynchronous race conditions.
- Generate minimal mock fixtures to reproduce the failure locally.
- Formulate the exact code diff that fixes all failing assertions without altering public API contracts.
Industry Implications
The launch of Ox Alpha illustrates the rise of stealth pre-release testing within commercial AI aggregation platforms. By deploying unannounced models onto platforms like OpenRouter, model developers can stress-test inference infrastructure and gather natural query telemetry before unveiling formal brand identities and pricing models.


