News / Model Watch · Mystery Model Watch
Ox Alpha Revealed: Zhipu Confirms It's GLM-5.3-Flash — Official Release and Open Weights Tonight
An anonymous model lands on OpenRouter with 1M context, image+video input, and a free week — then out-tests Claude Fable 5 in early samples, pulls Stripe's CEO into the comments, gets a wink-wink non-claim from Google that the internet ruthlessly mocks, and absorbs every forensic probe the community can throw at it. Chinese media calls it 牛来. Nobody has claimed it. Everything below is attributed, dated, and unconfirmed where noted.
Update (Aug 26, evening): mystery solved — Zhipu confirmed to Bloomberg that Ox Alpha is GLM-5.3-Flash, with the official release and open weights landing tonight. Zhipu's Zixuan Li: "Ox Alpha was an early version of GLM-5.3-Flash. The official release delivers stronger performance and significantly better stability." The forensic crowd called it days ago; the Pony Alpha script replayed to the letter. Details in The reveal section below.
Ox Alpha is GLM-5.3-Flash — officially confirmed by Zhipu to Bloomberg on Aug 26, with the official release and open weights landing tonight via Z.ai. The stealth run (OpenRouter/OpenCode since Aug 20-21, 1M context, text+image+video, free week) was an early public stress test: "frontier intelligence, flash cost," community-measured at 80% vs Fable 5's 65% on a 10-task DeepSWE sample, and per SemiAnalysis, its 100T-tokens/day capacity runs on Chinese chips. Zhipu says the official release is stronger and more stable than the stealth build.
The week it happened, verified
| When | What happened · source |
|---|---|
| Aug 20-21 | Ox Alpha appears on OpenRouter + OpenCode Zen — 1M context, multimodal, free ~1 week, "100T tokens/day" platform capacity (capacity, not parameters). OpenRouter: "routing only; anonymous third party." (36氪/APPSO) |
| Aug 21-22 | Silicon Valley piles in: DeepSWE's author measures 63% on his agentic suite; Ben Davis's 10-task sample: Ox 80% / Fable 5 65% / GPT-5.6 Sol 52%; Stripe CEO Patrick Collison "相当惊艳"; OpenRouter/T3 Chat/Box CEOs amplify. (36氪) |
| Aug 22-23 | Forensics converge on GLM: tokenizer counts vs GLM-5.3 differ by a constant 75 tokens across 25 prompts (one hidden system prompt's worth); video-token consumption matches GLM-5V-Turbo exactly; serving-layer errors match Z.ai's dialect. (TechTimes 08-23; 36氪) |
| Aug 23-24 | Google's non-claim: DeepMind researchers hint "Gemini" and "trust the process" — widely mocked (谷歌抢着认领被全网笑惨, 36氪). HN's verdict thread: "Ox-Alpha Is GLM" (26 pts, still community attribution only). |
| Aug 24 | 36氪 reports the free window has "less than four days" left; Chinese nickname 牛来 goes mainstream. Still zero official claims. |
The reveal (Aug 26): Bloomberg-confirmed, weights tonight
- Zhipu confirmed to Bloomberg (Aug 26, Wednesday): "the Ox Alpha model is a new iteration of its GLM series" — specifically GLM-5.3-Flash — and said it will release the weights tonight, in response to Bloomberg News' queries.
- Zhipu's Zixuan Li (X, @zixuanli_): "Ox Alpha was an early version of GLM-5.3-Flash. The official release delivers stronger performance and significantly better stability. Huge thanks to @opencode and @OpenRouter for making it available to the community."
- The shock detail (SemiAnalysis): the stealth run's claimed 100T tokens/day was served on Chinese chips — the free week was, among other things, a domestic-accelerator stress test at frontier scale.
- First reactions: Andrew Curran — "the Pareto Frontier has been redrawn" (announcement + benchmarks already up); community summary: "GLM-5.3 Flash outperforms GLM-5.2 at everything"; positioning line making the rounds: "frontier intelligence, flash cost."
- The script, replayed. February's Pony Alpha (anonymous → claimed by Zhipu as GLM-5 on day ~6) predicted this exactly: stealth release → free traffic → benchmarks → reveal. The tokenizer forensics that called it GLM days before the confirmation were right.
- Weights delivered (Aug 27 verification): zai-org/GLM-5.3-Flash (+ BF16) is live on Hugging Face. Official card facts: 320B total / 18B active, the first natively multimodal model in the GLM-5 series, outperforming GLM-5.2 across benchmarks at one-tenth the price (≈ $0.14/$0.44 per 1M vs 5.2's $1.40/$4.40) while approaching Claude Opus 4.8 on coding and agentic suites — with llama.cpp / Ollama / LM Studio quantizations for local serving. HN's reception: 809 points on the model, 407 on the reveal.
The pre-reveal forensics — and how right they were
- Tokenizer fingerprint. Across 25 controlled prompts, Ox Alpha's token counts sit exactly 75 tokens above GLM-5.3's — the signature of a constant hidden system prompt on the same tokenizer (36氪's technical recap).
- Video-token signature. Four controlled video tests matched GLM-5V-Turbo's frame-rate/duration/resolution token math line by line; other vendors' models diverged.
- It out-multimedials public GLM-5.3. The released 5.3 is text-only via API — so if this is Zhipu's, it's an unreleased multimodal variant (community shorthand: a GLM-5.5-class workhorse). APPSO's side-by-side found the reasoning chain "very close to GLM-5.3" — and one number-theory test Ox solved that GLM-5.3-max didn't finish.
- The prior. Four of four previous anonymous OpenRouter releases were Chinese labs: Pony Alpha → GLM-5 (claimed by Zhipu on day 6), Hunter/Healer Alpha → Xiaomi MiMo-V2, Elephant Alpha → Ant Ling-2.6-flash, Owl Alpha → Meituan LongCat-2.0.
- The limits (pre-reveal). As of Aug 25 none of this was official — Google hinted without claiming, Zhipu was silent. On Aug 26 the confirmation landed and validated every forensic thread above.
Practical notes for API builders
- From stealth to official, tonight. The free window (~Aug 27-28) now resolves into the official GLM-5.3-Flash release with open weights via Z.ai — the zero-cost benchmark window ends exactly as the real product arrives.
- Treat it as a stealth stress test. No SLA, no deprecation notice, no attribution — fine for evals and curiosities, wrong place for production traffic or sensitive data.
- Weights + API landing tonight via Z.ai — watch zai-org on Hugging Face and docs.z.ai; our GLM-5.3 coverage tracks the full family, and tomorrow's run carries verified release specs.
Primary sources
- Bloomberg 确认(经 @kimmonismus/Techmeme 转述)— 智谱周三确认 Ox Alpha 为 GLM 系新迭代,今晚放权重
- Techmeme 聚合 — GLM-5.3-Flash 官宣反应(HN/r/singularity/r/LocalLLaMA + SemiAnalysis 国产芯片帖 + Zixuan Li 致谢)
- 36氪 — 神秘「牛来」掀翻硅谷,谷歌抢着认领被全网笑惨(实测细节与取证汇总)
- 36氪/APPSO — 神秘「牛来」模型刷屏,实测到底牛不牛(tokenizer/视频取证 + 数论实测)
- TechTimes — serving-layer forensics (Zhipu class names, error dialect, 30/30 tokenizer probes)(2026-08-23)
- OrcaRouter — Ox Alpha likely GLM-5.3 (匿名模型发布史记分卡)
- Trending Topics — Ox Alpha: A Mysterious New AI Model Aims to Win Developers Over
- Hacker News — "Ox-Alpha Is GLM" 讨论帖(社区归因,非官方)
FAQ (2026)
What is Ox Alpha?
An anonymous stealth model on OpenRouter/OpenCode since Aug 20-21: 1M context, text+image+video, free ~1 week. No lab has claimed it. Chinese media nickname: 牛来.
How good?
Small-sample but striking: 80% vs Fable 5's 65% on a 10-task DeepSWE run; 63% on the DeepSWE author's suite. Stripe's CEO: "相当惊艳" (per 36氪).
Who built it?
Confirmed Aug 26: Zhipu — Ox Alpha was an early GLM-5.3-Flash (Bloomberg; Zhipu's Zixuan Li on X). The tokenizer +75-offset and video-signature forensics called it days early.
Did Google claim it?
DeepMind researchers hinted hard ("Gemini", "trust the process") and got mocked for it — no confirmation from anyone.
Still free?
The stealth free window resolves into tonight's official GLM-5.3-Flash release with open weights via Z.ai — free preview ends as the real product lands.
Link to GLM-5.3 weights?
Resolved and verified: GLM-5.3-Flash weights are live on zai-org (Aug 27 check) — 320B/18B-active, first natively multimodal GLM-5-series model, ~1/10th of GLM-5.2's price. The ~Aug-28 window landed early for Flash.