Z.ai says Ox Alpha was GLM-5.3-Flash, a 320B open-weights model
Z.ai has confirmed that Ox Alpha, the anonymous coding model that had developers guessing for a week, was a preview of GLM-5.3-Flash. The Chinese lab published the weights on Hugging Face under an MIT licence on 26 August 2026.
The model card puts it at 320 billion total parameters with 18 billion active per token, and calls it the first natively multimodal model in the GLM-5 series. It takes text, images and video, with a context window of just over 1 million tokens. On OpenRouter it lists at $0.15 per million input tokens and $0.50 per million output tokens, currently halved by a Z.ai discount that runs to 9 September 2026.
Z.ai says the stealth run was deliberate, and that it was gathering feedback rather than hiding.
Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback.
Z.ai, GLM-5.3-Flash model overview
That settles a question the developer community couldn’t answer on its own. SiliconANGLE reported on 23 August that the model had appeared free on OpenRouter three days earlier with no named provider, and that OpenCode advertised capacity for 100 trillion tokens a day. The developer unclecode matched Ox Alpha to GLM-5.3 infrastructure on six of nine fingerprint probes, while warning that matching fingerprints prove shared infrastructure rather than identity.
Z.ai’s own figures put the model well ahead of GLM-5.2 and just short of Anthropic’s Claude Opus 4.8 on an in-house coding test. These are the lab’s numbers, not ours, and no independent leaderboard has confirmed them yet.
| Benchmark | GLM-5.3-Flash | Comparison |
|---|---|---|
| DeepSWE v1.1 | 63.4 | 46.2 (GLM-5.2) |
| AutomationBench | 48.8 | 26.2 (GLM-5.2) |
| Z.ai Code Bench v1.0, max effort | 29.0 | 29.5 (Claude Opus 4.8) |
The lab also claims 57 on the Artificial Analysis Intelligence Index v4.1.1, at $0.045 per task on the discounted price. Most of that cost story sits in the architecture. Z.ai says this is the first model in the series to combine sparse and linear attention, cutting attention compute by 3.0x and KV cache size by 4.4x against GLM-5.3. Against the older GLM-4.5 it nearly halves both the active parameter count and the layer count, 18 billion against 32 billion and 45 layers against 92.
The part that reaches beyond price is where all of this ran. Z.ai says the entire Ox Alpha week was served on Chinese AI chips, across what it describes as tens of thousands of domestically developed accelerators. It reports a 3x end-to-end serving improvement over its own first attempt on that hardware, and a per-token cost comparable to mainstream Nvidia GPUs. That’s the lab describing its own cluster, and there’s no outside audit of any of it.
Not everyone is pleased about a release like this. Semafor reported on 25 August that the coming open-weight release had reopened the argument over whether frontier cyber capability should be downloadable at all, with some researchers warning about hacking ability handed out free and others arguing it lets defenders keep pace. Z.ai had already held back the GLM-5.3 weights over that same concern.
Business Insider reported that Chinese social media users had tied the Ox name to a low-budget animated film that went viral this summer, and that Z.ai had run the same anonymous playbook before under the name Pony Alpha. TechCrunch reads the launch as more pressure on the pricing of expensive frontier labs, which is the usual calculation behind an open weights release. The test now is whether independent evaluations land near Z.ai’s numbers, and whether the Chinese-chip serving claim survives contact with people running the weights themselves.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
