ChinaModelAPI

News / Model Watch · Speed tier listing

Official rate card · API-only 1M context reasoning always on no independent benchmarks yet
2026-09-24 listing check UTC+8 · Z.ai docs + pricing page + third-party gateways

GLM-5.3-FlashX API: Z.ai's Speed Tier, Officially Priced at $0.37/$1.25

No launch event, no benchmark day, no X thread from Z.ai: GLM-5.3-FlashX simply appeared in the official docs and pricing page as the speed tier of GLM-5.3-Flash, at $0.37 in / $1.25 out per 1M tokens with a 1M-token context and reasoning that cannot be switched off. That is enough to establish a real, purchasable API model, and not enough to answer the two questions buyers actually have: how much faster it is, and whether the weights follow the Flash line's open release. This brief holds that line carefully, and places the listing in the budget-tier fight that reshaped this week.

Direct answer

What is established: Z.ai's pricing page lists glm-5.3-flashx at $0.37 / $0.075 cached / $1.25 per 1M (storage limited-time free); API docs give 1M context, 128K max output, text+image+video input, reasoning always on. Third-party routes are live: zai/glm-5.3-flashx on the Vercel AI Gateway and OpenRouter at the same rates. What is not: no independent benchmark has scored FlashX (Artificial Analysis still lists only Flash, at 41.9 intelligence / 113.9 tok/s), no separate weights release is verified, and the "up to 200 tok/s" figure circulating is a third-party tracker claim, not a measured result. Buyer takeaway: roughly 2.5x Flash's price for a promised speed tier; buy it for latency-sensitive interactive agents on the official endpoint, benchmark before committing volume, and do not plan self-hosting on it.

The facts, as verified (Sep 24)

  • The official rate card. Z.ai's pricing page lists GLM-5.3-FlashX at $0.37 input / $0.075 cached input / $1.25 output per 1M tokens, cached-input storage limited-time free, directly beside GLM-5.3-Flash ($0.15/$0.50) and GLM-5.3 ($1.40/$4.40). This is a provider-published product with a purchasable price, a clear step up in evidence from the API-model-code sighting it started as.
  • Documented specs. 1M-token context (1,048,576), 128K max output, native multimodal input spanning text, images and video, and reasoning that stays on: a lower reasoning_effort does not disable it, and Z.ai's thinking control is documented for the FlashX tier specifically.
  • API-only positioning. Third-party coverage of the Z.ai Coding Plan reports GLM-5.3-Flash added to the subscription with 3x quota while FlashX stays outside it. Subscription buyers do not have FlashX; API buyers do.
  • Third-party availability, same day. The Vercel AI Gateway lists zai/glm-5.3-flashx, OpenRouter carries it at $0.37/$1.25 with the 1M window, and independent resellers have it in their catalogs. Identical rates across official and gateway channels suggest the listing, not a promotional price.
  • The Flash base underneath. GLM-5.3-Flash is documented by Z.ai as the first open-source frontier model combining sparse and linear attention, 320B total parameters with 18B activated. Artificial Analysis scores Flash at 41.9 on its Intelligence Index v4.3 (GLM-5.3: 44.9) with 113.9 tok/s output, placing it near the top of the open-weights field alongside Kimi K3 and Qwen3.8.
  • What remains unverified, explicitly. FlashX has no Artificial Analysis score yet, no published independent latency measurement, and no confirmed separate weights release. The "up to 200 tok/s" figure comes from a third-party model tracker, not from Z.ai or a measured benchmark. As-of statements, anchored to Sep 24; this page gets revised when the sources move.

Prices are provider-published per-1M figures read from Z.ai's pricing page and docs on Sep 23-24, 2026. Speed claims are attributed to their sources; no benchmark purchase or API test was performed for this article.

Context: the budget tier just got crowded

FlashX landed in the same week DeepSeek made V4.1 Flash official at $0.15/$0.60 per 1M with MIT open weights (our launch coverage), while the Western top tier repriced around it: Opus 5.5 at $4/$20 and Grok 4.7 holding $2/$6 (tri-model brief). Blended 3:1, the ladder now reads Flash about $0.24, FlashX about $0.59, GLM-5.3 $2.15, Grok 4.7 about $3, Opus 5.5 about $8 per 1M.

That makes FlashX's fight internal as much as external. Against DeepSeek's flash it concedes price and self-host freedom (MIT weights versus no verified FlashX weights); against its own sibling Flash it must prove the 2.5x premium buys real latency. The one structural card it holds over both: 1M context with multimodal input and always-on reasoning in a speed-tier SKU, a combination neither competitor packages at that price.

The flash-tier price war this extends started earlier in the month; see the Gemini 3.8 Flash analysis for the 2027 price-trap angle that still applies, and the GLM API guide for endpoint, model-id and integration detail (FlashX section included).

What this means for builders

The honest evaluation protocol for a speed tier without independent speed data: run one real request through the official endpoint and read the usage line, measure first-token and output latency on your own prompt mix, and compare against Flash and DeepSeek V4.1 Flash on cost-per-completed-task, not price-per-token. Reasoning always on means the thinking bills, so effort tuning belongs in the cost model from day one.

Route shape: interactive coding agents and chat-with-latency-SLA products are the natural FlashX candidates, where 2.5x over Flash is cheap against UX gains. Batch classification and throughput work have no reason to leave the $0.15/$0.60 tier. A model id on a pricing page is not a routing promise from any relay, including this site's future service: confirm the exact alias, price and region in your provider dashboard before production traffic.

Watch three signals for the follow-up: an Artificial Analysis FlashX score (which would put a real number on the speed claim), a weights release (which would change the self-host math), and Coding Plan inclusion (which would change the subscription math). Any of them moves, and this page gets an update entry.

Access from outside China — and the usual disclaimer

This site tracks Chinese model APIs for international builders: official endpoints, pricing and open-weights status. The GLM API guide covers Z.ai's OpenAI-compatible endpoint and the full model-id lineup; the comparison table holds the cross-vendor ladder.

Independence disclaimer: ChinaModelAPI is an independent information site, not affiliated with Z.ai, DeepSeek or any model vendor. Figures here are provider-published or explicitly attributed to the independent outlets and trackers that reported them; the speed claim is third-party and unverified; nothing on this page is a routing or availability promise.

Primary sources

FAQ (2026)

What is GLM-5.3-FlashX?

The speed tier of GLM-5.3-Flash in Z.ai's official docs and pricing page: $0.37/$1.25 per 1M, 1M context, 128K max output, text+image+video input, reasoning always on. API-only.

How much does GLM-5.3-FlashX cost?

$0.37 in / $0.075 cached / $1.25 out per 1M (official page). Scale: 2.5x Flash ($0.15/$0.50), about a quarter of GLM-5.3 ($1.40/$4.40), an order of magnitude under Western flagships.

What are GLM-5.3-FlashX's specifications?

1M-token context (1,048,576), 128K max output, native multimodal input (text/image/video), reasoning always on: lower reasoning_effort does not switch it off.

Is GLM-5.3-FlashX included in the Z.ai Coding Plan?

No. Third-party coverage reports Flash was added to the Coding Plan with 3x quota while FlashX remains API-only. Subscription capacity and API capacity are separate as of Sep 24, 2026.

Is GLM-5.3-FlashX an open-weight model?

Flash is documented open (320B/18B, sparse+linear attention). No separate FlashX weights release is verified as of Sep 24, 2026. Treat it as an API-served variant; check the repo and license before self-hosting plans.

GLM-5.3-FlashX or DeepSeek V4.1 Flash?

Price: DeepSeek ($0.15/$0.60, MIT weights). FlashX counters with speed-tier positioning, 1M context and multimodal input. Interactive agents may justify 2.5x; batch rarely does. Benchmark both.

Related guides