ChinaModelAPI

News / Model Watch · Chinese frontier

Official release GLM-5.3 Zhipu / Z.ai Flagship LLM
2026-08-14 · Model Watch · updated 2026-08-22 · API GA'd Aug 19 — see the dedicated API GA brief

GLM-5.3 Officially Released: Post-Training Scaling, Emergent Cyber, Weights in 2 Weeks

Z.ai officially announced GLM-5.3 on August 14 — same base model as GLM-5.2, with every gain coming from scaled post-training. It's live for all GLM Coding Plan subscribers right now, ships with an API breaking change on thinking parameters, finds real vulnerabilities at a rate that surprised its own creators, and goes open-weights in ~two weeks. Updated Aug 22: the API went GA on Aug 19 (pricing $1.40/$4.40 per 1M, same as 5.2) — full timeline, platform roll call, and builder notes now in the dedicated API GA brief. Full numbers, the migration warning, and overseas reaction below.

Direct answer

GLM-5.3 is officially out (Aug 14, 2026). The official announcement states: "Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training." Z.ai calls it "the most capable open-weights model for coding" (a 50% jump on its in-house Z.ai Code Bench) and state-of-the-art on CyberGym for vulnerability discovery. Rollout status: GLM Coding Plan — live for all subscribers (points-based quota; 50% off-peak discount); API — documented with a required migration (three effort levels low/high/max, and thinking.type: "disabled" is gone — old requests will fail); weights — expected ~Aug 28 after safety hardening (not on Hugging Face as of Aug 19). API pricing is now published: $1.40 input / $4.40 output per 1M tokens ($0.26 cached input) — identical to GLM-5.2, and OpenRouter lists glm-5.3 at the same prices (Aug 18). The first independent eval is in: Artificial Analysis index 60 (rank 8/182). For production stacks on relays: stay on glm-5.2 until 5.3 is generally available and healthy upstream.

Verified status at a glance

ItemStatus (per official sources, 2026-08-14)
AnnouncementOfficial blog post live: z.ai/blog/glm-5.3 ("Today we are releasing GLM-5.3")
GLM Coding PlanRolled out to all subscribers; points-based quota; off-peak 50% (peak = weekdays 14:00–18:00 UTC+8)
APIModel page + docs live; glm-5.3 with thinking effort low / high / max (default max); disabling thinking no longer supported. Listed on OpenRouter at the same price since Aug 18
Open weights"in two weeks after launch, once safety evaluation and hardening are complete" — as of Aug 19 not yet on Hugging Face (zai-org); expected ~Aug 28
Context / output1M context · 128K max output · text in / text out (no native vision)
API pricing$1.40 input / $0.26 cached input / $4.40 output per 1M (docs.z.ai international USD, verified Aug 19) — same as GLM-5.2. Coding Plan: points-based, 50% off-peak
Changelog entriesdocs.bigmodel.cn model page live; the release-notes/changelog pages hadn't caught up at the time of writing

Rollout order this time was blog + Coding Plan first, changelogs trailing — the mirror image of a docs-first rollout. If your signal is "is it in the changelog yet", you'll be late.

The numbers (vendor-reported, from the official table)

All scores below are Z.ai's own reported numbers from the announcement. Selected rows; the full table in the blog covers ~20 benchmarks.

BenchmarkGLM-5.3GLM-5.2Kimi K3Opus 4.8Fable 5 / GPT-5.6 Sol
Terminal Bench 2.188.281.088.385.088.0 / 88.8
Terminal Bench 3.028.34.617.421.133.7 / 34.6
DeepSWE v1.166.946.267.558.069.7 / 72.7
SWE-Marathon v1.142.519.448.148.833.1 / 42.5
Agents' Last Exam (CLI)28.523.827.625.723.8 / 28.6
HLE w/ Tools62.554.759.857.963.9 / 64.5
GDPval-AA v217691508168215881743 / 1730
CyberGym84.577.280.078.183.8 / 83.6
ExploitBench54.424.432.240.078.0 / 76.5
  • Honest read of the table: GLM-5.3 roughly matches or beats Kimi K3 and DeepSeek-V4-Pro-0813 (62.7 DeepSWE) across coding/agentic rows, sits at or near Opus 4.8, and still trails Claude Fable 5 and GPT-5.6 Sol on the hardest rows (Terminal Bench 3.0, DeepSWE, ExploitBench). Z.ai itself says the gap to the closed frontier is widest exactly where it improved most.
  • Token efficiency: on the in-house Z.ai Code Bench, at High effort 5.3 hits 31.4% with ~50K output tokens, surpassing Opus 4.8's 29.5% at 120K; Fable 5 still leads at 39.5% (Max effort).
  • Vendor claims, labeled as such: "most capable open-weights model for coding", open-source SOTA on Terminal Bench 3.0 and ALE, ~50% coding-experience improvement. Independent verification arrives with the weights in ~2 weeks.

First independent check: Artificial Analysis (Aug 18)

Artificial Analysis added GLM-5.3 on August 18 — the first non-vendor evaluation. From its model page (read Aug 19):

  • Intelligence Index v4.1.1: 60 — rank 8 of 182 (median 35); listed as proprietary until weights ship; 84.7 tok/s, 1.88s time-to-first-token, 1M context.
  • Where that lands: at max effort, AA-tracked numbers compiled in the HN benchmarks thread put GLM-5.3 at 59.5 vs Kimi K3 59.7, GPT-5.6 Sol 60.9, Claude Opus 5 61.5 — parity with K3, just under the closed frontier, up ~7 points over GLM-5.2's 53.0.
  • The caveat — verbosity: AA explicitly flags GLM-5.3 as "very verbose": it emitted 170M tokens across the index run vs a 72M median, and ~41k output tokens per task vs ~16.9k for GPT-5.6 Sol. Token efficiency, not intelligence, is the cost lever to watch.
  • Not yet on LMArena as of Aug 19 (checked directly).

The cyber story is real-world, not just benchmarks

The most unusual part of the release: Z.ai ran GLM-5.2/5.3 against real codebases with several security teams in China. After expert review and deduplication, the model identified 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high-severity issues (the announcement's severity chart shows 107 Critical + 990 High) — spanning kernels, browser engines, open-source infrastructure, web apps, and network protocols. The oldest flaw dates to 1981; on average a vulnerability had lived 26.6 years before discovery.

  • Findings flow into a public Z.ai Security Disclosure Ledger — per unite.ai's coverage, 53 were public with CVEs at announcement and 2,383 remain under embargo (Linux kernel use-after-free, WebKit/Safari memory bugs, FreeBSD parameter validation among them).
  • On exploitation-depth benchmarks the pattern is consistent: big jumps from 5.2 (ExploitGym 29→105 tasks @2h), but Mythos 5 and GPT-5.6 Sol remain well ahead. This is a strong defensive/security-review capability, not an offense frontier.
  • Z.ai's own framing: capability "developed faster than we expected" as post-training scaled — hence the two-week safety hardening window before weights ship.

API builders: one breaking change to plan for

  • Migration required: GLM-5.3 supports thinking effort low / high / max (default max, recommended for coding) — and thinking.type: "disabled" is no longer supported. The official docs are blunt: change it to enabled and set reasoning_effort before switching the model ID, "otherwise, the request will fail."
  • Coding Plan economics: the plan is now points-based (input, cached input, and output counted separately); weekdays 14:00–18:00 UTC+8 are peak, everything else — including weekends — consumes 50% of standard points. ZCode adds a 98%+ cache hit rate and a 1.5x quota boost through Aug 31.
  • Through relays like ChinaModelAPI: production stays on glm-5.2 for now; we'll add the 5.3 model ID once the upstream API is generally available and healthy. When you migrate, budget for the mandatory thinking-parameter change above.
  • Same base, same envelope: 1M context / 128K output / same capability surface (thinking, streaming, function calling, context cache, structured output, MCP) — so migration is a model-ID + params swap plus regression runs on your agent loops.
  • Still text-only: vision was the top community ask in Zhipu's feedback round; it didn't make this release.

Overseas reaction (updated Aug 19)

  • VentureBeat ran the deepest English write-up (Aug 14): repeats the official table with a caution that Z.ai Code Bench is a vendor-private benchmark, reports a Z.ai developer advocate's claim that GLM-5.3 already found a "serious vulnerability" in Cursor (Cursor hadn't confirmed at press time), and cites Reuters on the "trusted access" controls behind the two-week weights delay. It also lists Coding Plan promotional tiers (Lite $12.60/mo, Pro $56, Max $117.60) — labeled promotional; Hacker News buyers report standard $18/$80 pricing.
  • The Decoder covered it as "Zhipu AI releases GLM-5.3, claims it's the strongest open weights coding model" — same-base post-training, biggest gains on agentic tasks, weights in two weeks, usable via Coding Plan / ZCode / Claude Code / OpenCode.
  • Interconnects (Nathan Lambert) published the most substantive independent analysis: "GLM-5.3: How Chinese labs keep stride with the frontier" — the scores "are the real deal", he's skeptical of distillation explanations, credits faster release cadence and a narrower text/coding scope, notes the model is reportedly ~750B parameters (roughly a third of Kimi K3's size) and cites reports of ~$1B ARR for Z.ai.
  • Reuters (per search excerpts; paywalled for direct read): GLM-5.3 "nears Anthropic's Mythos 5 in cyber defence tests" — CyberGym 84.5% vs 83.8% — with the weights delay tied to testing "trusted access" controls for sensitive cyber capabilities.
  • Hacker News: the launch thread hit 1,167 points / 582 comments; a follow-up AA-benchmarks thread (Aug 18) and an OpenRouter-availability post (Aug 19) followed.
  • Reddit r/LocalLLaMA has a "GLM 5.3 Released" thread linking the official blog, with the headline detail that weights land "in two weeks".
  • Reddit r/ZaiGLM users on Lite plans reported GLM-5.2 "responding very differently" in the days before announcement — consistent with a staged rollout behind the 5.2 label rather than a big-bang switch.
  • unite.ai covered the launch under "a cyber capability that outgrew its training", highlighting the 2,436-vulnerability disclosure program and the end-of-August weights window for independent verification. (Note: unite.ai labels the piece AI-generated, editor-reviewed.)
  • Laurie Voss (LinkedIn) — a widely read commentator on open-weight models — framed 5.2-as-go-to with "GLM 5.3 open weights in development", putting the GLM line ~6–9 months ahead of the open-weight pack in his telling.

Community read after five days

  • Positive: security researchers describe it as the first model that "agrees to and fully executes serious red-team scenarios" (WordPress plugin 0-days, RCE, kernel-exploit adaptation — HN, first-hand); teams are migrating from Claude subscriptions over refusal friction on security tooling; the AA score (60) matched release-day tester impressions, and visible reasoning tokens let you kill a runaway trace early.
  • Negative: verbosity drives token spend (AA's "very verbose" flag); debates over Coding Plan value vs API pricing ("GLM only wins at API price"); skepticism that the two-week "safety hardening" window will quietly nerf the cyber capability; and no multimodal — still the top ask for web-dev workflows.

Primary sources

FAQ

Is GLM-5.3 released?

Yes — officially announced Aug 14, 2026 on the Z.ai blog. Coding Plan users have it now; weights follow in ~2 weeks.

GLM-5.3 pricing?

Now published: $1.40 input / $4.40 output per 1M tokens ($0.26 cached) on docs.z.ai — same as GLM-5.2; OpenRouter matches. Coding Plan is points-based with 50% off-peak discounts (peak = weekdays 14:00–18:00 UTC+8).

Will my glm-5.2 calls break?

No — 5.2 stays GA, open-weight, and routed as usual. But when you do move to 5.3, thinking.type: "disabled" requests will fail; migrate params first.

Does 5.3 add vision?

No — text in / text out. Vision was the top community request but is not in this release.

Related guides