News / Model Watch · Speed tier listing
GLM-5.3-FlashX API: Z.ai's Speed Tier, Officially Priced at $0.37/$1.25
No launch event, no benchmark day, no X thread from Z.ai: GLM-5.3-FlashX simply appeared in the official docs and pricing page as the speed tier of GLM-5.3-Flash, at $0.37 in / $1.25 out per 1M tokens with a 1M-token context and reasoning that cannot be switched off. That is enough to establish a real, purchasable API model, and not enough to answer the two questions buyers actually have: how much faster it is, and whether the weights follow the Flash line's open release. This brief holds that line carefully, and places the listing in the budget-tier fight that reshaped this week.
What is established: Z.ai's pricing page lists glm-5.3-flashx at $0.37 / $0.075 cached / $1.25 per 1M (storage limited-time free); API docs give 1M context, 128K max output, text+image+video input, reasoning always on. Third-party routes are live: zai/glm-5.3-flashx on the Vercel AI Gateway and OpenRouter at the same rates.
What is not: no independent benchmark has scored FlashX (Artificial Analysis still lists only Flash, at 41.9 intelligence / 113.9 tok/s), no separate weights release is verified, and the "up to 200 tok/s" figure circulating is a third-party tracker claim, not a measured result.
Buyer takeaway: roughly 2.5x Flash's price for a promised speed tier; buy it for latency-sensitive interactive agents on the official endpoint, benchmark before committing volume, and do not plan self-hosting on it.
The facts, as verified (Sep 24)
- The official rate card. Z.ai's pricing page lists GLM-5.3-FlashX at $0.37 input / $0.075 cached input / $1.25 output per 1M tokens, cached-input storage limited-time free, directly beside GLM-5.3-Flash ($0.15/$0.50) and GLM-5.3 ($1.40/$4.40). This is a provider-published product with a purchasable price, a clear step up in evidence from the API-model-code sighting it started as.
- Documented specs. 1M-token context (1,048,576), 128K max output, native multimodal input spanning text, images and video, and reasoning that stays on: a lower reasoning_effort does not disable it, and Z.ai's thinking control is documented for the FlashX tier specifically.
- API-only positioning. Third-party coverage of the Z.ai Coding Plan reports GLM-5.3-Flash added to the subscription with 3x quota while FlashX stays outside it. Subscription buyers do not have FlashX; API buyers do.
- Third-party availability, same day. The Vercel AI Gateway lists
zai/glm-5.3-flashx, OpenRouter carries it at $0.37/$1.25 with the 1M window, and independent resellers have it in their catalogs. Identical rates across official and gateway channels suggest the listing, not a promotional price. - The Flash base underneath. GLM-5.3-Flash is documented by Z.ai as the first open-source frontier model combining sparse and linear attention, 320B total parameters with 18B activated. Artificial Analysis scores Flash at 41.9 on its Intelligence Index v4.3 (GLM-5.3: 44.9) with 113.9 tok/s output, placing it near the top of the open-weights field alongside Kimi K3 and Qwen3.8.
- What remains unverified, explicitly. FlashX has no Artificial Analysis score yet, no published independent latency measurement, and no confirmed separate weights release. The "up to 200 tok/s" figure comes from a third-party model tracker, not from Z.ai or a measured benchmark. As-of statements, anchored to Sep 24; this page gets revised when the sources move.
Prices are provider-published per-1M figures read from Z.ai's pricing page and docs on Sep 23-24, 2026. Speed claims are attributed to their sources; no benchmark purchase or API test was performed for this article.
Context: the budget tier just got crowded
FlashX landed in the same week DeepSeek made V4.1 Flash official at $0.15/$0.60 per 1M with MIT open weights (our launch coverage), while the Western top tier repriced around it: Opus 5.5 at $4/$20 and Grok 4.7 holding $2/$6 (tri-model brief). Blended 3:1, the ladder now reads Flash about $0.24, FlashX about $0.59, GLM-5.3 $2.15, Grok 4.7 about $3, Opus 5.5 about $8 per 1M.
That makes FlashX's fight internal as much as external. Against DeepSeek's flash it concedes price and self-host freedom (MIT weights versus no verified FlashX weights); against its own sibling Flash it must prove the 2.5x premium buys real latency. The one structural card it holds over both: 1M context with multimodal input and always-on reasoning in a speed-tier SKU, a combination neither competitor packages at that price.
The flash-tier price war this extends started earlier in the month; see the Gemini 3.8 Flash analysis for the 2027 price-trap angle that still applies, and the GLM API guide for endpoint, model-id and integration detail (FlashX section included).
What this means for builders
The honest evaluation protocol for a speed tier without independent speed data: run one real request through the official endpoint and read the usage line, measure first-token and output latency on your own prompt mix, and compare against Flash and DeepSeek V4.1 Flash on cost-per-completed-task, not price-per-token. Reasoning always on means the thinking bills, so effort tuning belongs in the cost model from day one.
Route shape: interactive coding agents and chat-with-latency-SLA products are the natural FlashX candidates, where 2.5x over Flash is cheap against UX gains. Batch classification and throughput work have no reason to leave the $0.15/$0.60 tier. A model id on a pricing page is not a routing promise from any relay, including this site's future service: confirm the exact alias, price and region in your provider dashboard before production traffic.
Watch three signals for the follow-up: an Artificial Analysis FlashX score (which would put a real number on the speed claim), a weights release (which would change the self-host math), and Coding Plan inclusion (which would change the subscription math). Any of them moves, and this page gets an update entry.
Access from outside China — and the usual disclaimer
This site tracks Chinese model APIs for international builders: official endpoints, pricing and open-weights status. The GLM API guide covers Z.ai's OpenAI-compatible endpoint and the full model-id lineup; the comparison table holds the cross-vendor ladder.
Independence disclaimer: ChinaModelAPI is an independent information site, not affiliated with Z.ai, DeepSeek or any model vendor. Figures here are provider-published or explicitly attributed to the independent outlets and trackers that reported them; the speed claim is third-party and unverified; nothing on this page is a routing or availability promise.
Primary sources
- Z.ai docs — pricing page: FlashX $0.37/$1.25, Flash $0.15/$0.50, GLM-5.3 $1.40/$4.40 per 1M
- Z.ai developer docs — GLM-5.3-Flash/FlashX overview: 320B/18B, sparse+linear attention, specs
- Vercel AI Gateway — zai/glm-5.3-flashx listing
- Artificial Analysis — open-source model comparison (Flash scored 41.9 / 113.9 tok/s; FlashX not yet scored)
- 站内对照:DeepSeek V4.1 Flash 官宣文 · Flash 层价格战分析 · 三模型同周对比简报 · GLM API 指南
FAQ (2026)
What is GLM-5.3-FlashX?
The speed tier of GLM-5.3-Flash in Z.ai's official docs and pricing page: $0.37/$1.25 per 1M, 1M context, 128K max output, text+image+video input, reasoning always on. API-only.
How much does GLM-5.3-FlashX cost?
$0.37 in / $0.075 cached / $1.25 out per 1M (official page). Scale: 2.5x Flash ($0.15/$0.50), about a quarter of GLM-5.3 ($1.40/$4.40), an order of magnitude under Western flagships.
What are GLM-5.3-FlashX's specifications?
1M-token context (1,048,576), 128K max output, native multimodal input (text/image/video), reasoning always on: lower reasoning_effort does not switch it off.
Is GLM-5.3-FlashX included in the Z.ai Coding Plan?
No. Third-party coverage reports Flash was added to the Coding Plan with 3x quota while FlashX remains API-only. Subscription capacity and API capacity are separate as of Sep 24, 2026.
Is GLM-5.3-FlashX an open-weight model?
Flash is documented open (320B/18B, sparse+linear attention). No separate FlashX weights release is verified as of Sep 24, 2026. Treat it as an API-served variant; check the repo and license before self-hosting plans.
GLM-5.3-FlashX or DeepSeek V4.1 Flash?
Price: DeepSeek ($0.15/$0.60, MIT weights). FlashX counters with speed-tier positioning, 1M context and multimodal input. Interactive agents may justify 2.5x; batch rarely does. Benchmark both.