News / Model Watch · Model releases
Opus 5.5, Grok 4.7 and GLM-5.3-FlashX: What Actually Changed for API Builders
Three model names landed in the same week, and they look like one launch wave. The evidence is different for each: Claude Opus 5.5 (Sep 22) and Grok 4.7 (Sep 21) are full launches with published rate cards and vendor benchmarks, while GLM-5.3-FlashX is an officially listed, officially priced Z.ai API model with no launch event at all. This brief separates what each source establishes, what stays unverified, and where the September price ladder now sits for builders routing Chinese-model traffic.
What shipped: Anthropic launched Claude Opus 5.5 on Sep 22, 2026 at $4 / $20 per 1M tokens (cache reads $0.20), API id claude-opus-5-5, live on Claude Platform, AWS, Google Cloud and Azure. xAI launched Grok 4.7 on Sep 21 at $2 / $6 per 1M below a 200k-token prompt threshold (doubling above), API id grok-4.7, 500k context.
What was listed, not launched: Z.ai's official docs and pricing page carry glm-5.3-flashx, a speed variant of GLM-5.3-Flash at $0.37 / $1.25 per 1M with 1M context. That listing establishes an official API model and its price; it does not by itself prove a separate open-weight release or independently verified speed gains.
Builder takeaway: two Western flagships repriced the top tier inside 48 hours while the Chinese speed tier held a sub-dollar blended cost. Verify any specific route in the provider dashboard before moving production traffic.
The three announcements, as verified (Sep 24)
- Claude Opus 5.5 (Sep 22), first of the 5.5 family. Standard API rates $4 input / $20 output per 1M tokens, cache reads $0.20, cache writes $5 (Opus 5 was $5 / $25 / $0.50 / $6.25). A fast mode runs $8 / $40 with up to 2.5x speed, and Anthropic says standard output is more than 30% faster than Opus 5. Positioning: near-Fable-5.1 performance at 40% lower run cost than Opus 5 (a vendor claim, not our test). The announcement does not state a context window. Sonnet 5.5 and Haiku 5.5 are announced as following.
- Opus 5.5 vendor benchmarks. Terminal-Bench 4.0: 66.4% at xhigh effort vs Opus 5 at 52.3%, Fable 5.1 at 55.8%, GPT-6 Astra at 57.9%, GPT-5.6 Sol at 37.3%. GDPval-AA v2.1: 1846 Elo (Fable 5.1 1735, Opus 5 1708). Humanity's Last Exam 67.7% with tools; CursorBench 4.0 57.8%.
- Grok 4.7 (Sep 21), built for coding and knowledge work. 500k context window, multimodal input, knowledge cutoff May 2026. Documented rates: $2 input / $0.50 cached / $6 output per 1M below 200k prompt tokens, doubling to $4 / $1 / $12 above; docs.x.ai additionally notes a 1.1x regional-processing multiplier (US-endpoint scope). Grok 4.7 Fast (2x rates) exists only inside Cursor and Grok Build. Rates match Grok 4.6's card unchanged.
- Grok 4.7 benchmarks, with an independent counterweight. xAI-reported: CursorBench 4.0 46.3%, Terminal-Bench 4.0 38.0% (xhigh), well under Opus 5.5's 57.8% / 66.4%. Artificial Analysis's independent Terminal-Bench 4.0 run scored Grok 4.7 at 26%. The Hacker News thread also notes the launch slipped about two weeks past the teased date, shipping the day before Opus 5.5.
- GLM-5.3-FlashX — officially listed, quietly. Z.ai's pricing page lists
glm-5.3-flashxat $0.37 / $1.25 per 1M (cached input $0.075, storage limited-time free) beside GLM-5.3-Flash at $0.15 / $0.50 and GLM-5.3 at $1.40 / $4.40. Documented specs: 1M-token context, 128K max output, text+image+video input, reasoning always on (a lower reasoning effort does not switch it off). Third-party routezai/glm-5.3-flashxis live on the Vercel AI Gateway and OpenRouter at the same rates.
All prices are provider-published per-1M-token figures read from official pages on Sep 23-24, 2026. No benchmark or API purchase was performed for this article; benchmark numbers are labeled by who ran them.
The 48-hour flagship window
The timing is the frame: Grok 4.7 on Sep 21, Opus 5.5 on Sep 22, and in the same window OpenAI cut GPT-6 API prices 50% with the Sol and Luna models (1113-point HN thread). Three vendors, three separate benchmark suites, and, as independent roundups point out, Terminal-Bench 4.0 is the only benchmark all three actually reported, which makes it the least-bad common yardstick this month.
The Opus 5.5 reception went mainstream fast: the launch thread reached 1152 points on Hacker News, and Anthropic's claim of a 680,000-line code migration completed in under a day circulated widely. A more quotable datapoint came from a developer patch-review harness posted in the thread: Opus 5.5 found 8 of 14 known issues for $15.40, where Fable 5.1 found 7 for $66.34 and Opus 5 found 6 for $15.19: near-flagship accuracy at mid-tier cost, in one real (n=1) test.
Grok 4.7's read is value, not peak: at $2 / $6 it prices at roughly a third of Opus 5.5's output rate, and cursor-side analyses put its xhigh scores within fractions of Opus 5 max at about half the task cost. If your work is high-volume coding that does not need the top agentic tier, that is the slot it is built for. Our previous coverage of the Grok 4.6 launch holds the baseline this builds on.
GLM-5.3-FlashX: what the listing establishes, and what it doesn't
- It is official, with a price. This is no longer a gray-stage model-code sighting: the Z.ai pricing page and API docs name the model, its rates ($0.37 / $1.25 per 1M) and its specs (1M context, 128K max output, multimodal input, reasoning always on). The original version of this brief treated FlashX as "an API model code"; the official pricing entry is what upgrades it to a documented product.
- Where it sits in the GLM lineup. GLM-5.3-Flash is the open frontier baseline: 320B total parameters, 18B activated, the first open-source frontier model to combine sparse and linear attention, per Z.ai's docs. FlashX is the speed tier on top of it — roughly 2.5x Flash's price for lower latency. Third-party trackers claim up to 200 tok/s; Artificial Analysis has not scored FlashX yet (Flash measures 41.9 intelligence / 113.9 tok/s on their v4.3 board), so treat speed claims as unverified until an independent run lands.
- API-only, for now. Third-party coverage of the Z.ai Coding Plan notes GLM-5.3-Flash was added to the subscription with 3x quota while FlashX stays API-only. If you buy capacity through the Coding Plan rather than the API, FlashX is not in your plan as of Sep 24.
- What remains unverified. No separate FlashX open-weight release is confirmed; no independent benchmark of FlashX exists yet; and nothing in the docs ties FlashX to a specific relay's routing. As-of statements, not permanent facts — this page gets revised when the sources move, and the timestamp above is the anchor.
What this means for Chinese-model buyers
Blended at a typical 3:1 input-output ratio, the September ladder reads like this: GLM-5.3-Flash about $0.24 per 1M blended, GLM-5.3-FlashX about $0.59, GLM-5.3 $2.15, Grok 4.7 about $3, Opus 5.5 about $8. The Chinese speed tier remains an order of magnitude below Western flagships, and that gap is the whole reason the FlashX listing matters even without a launch event.
The real contest is inside the budget tier. DeepSeek's V4.1 Flash went official this same week at $0.15 / $0.60 per 1M with MIT open weights (our launch coverage), which undercuts FlashX on price and beats it on freedom-to-self-host. FlashX's counter is latency and the 1M context window with reasoning always on. If your workload is interactive coding agents, that trade can be worth 2.5x; if it is batch throughput, it likely is not — the same logic as the flash-tier price war analysis from earlier this month, one tier up.
Practical checklist before switching production traffic: confirm the exact model alias in your provider's dashboard (listings and routes do not always move together), check the region you are served from, run one real request and read the usage line, and only then move volume. A model id on a pricing page is not a routing promise from any relay, including this site's future service.
Access from outside China — and the usual disclaimer
This site tracks Chinese model APIs for international builders — official endpoints, pricing and open-weights status — with Western launches covered as comparison baselines. The GLM API guide below covers Z.ai's OpenAI-compatible endpoint, model ids and pricing in detail; the comparison page holds the current cross-vendor table.
Independence disclaimer: ChinaModelAPI is an independent information site, not affiliated with Anthropic, xAI, Z.ai or any model vendor. Figures here are provider-published or explicitly attributed to the independent outlets that ran them; performance claims are vendor claims unless marked otherwise, and nothing on this page is a routing or availability promise.
Primary sources
- Anthropic — Claude Opus 5.5 announcement (Sep 22, 2026): pricing, fast mode, benchmarks, cloud availability
- xAI — "Introducing Grok 4.7" (Sep 21, 2026)
- docs.x.ai — API pricing: grok-4.7 rates below/above the 200k prompt threshold, 1.1x regional multiplier
- Z.ai docs — pricing page: GLM-5.3-FlashX $0.37/$1.25, GLM-5.3-Flash $0.15/$0.50, GLM-5.3 $1.40/$4.40 per 1M
- Vercel AI Gateway — zai/glm-5.3-flashx listing
- Hacker News — Grok 4.7 thread (weights +40% vs 4.6 at same price, launch timing)
- Kingy.ai — Opus 5.5 vs Grok 4.7 benchmark and price comparison (incl. AA's independent Terminal-Bench run)
- explainx.ai — Opus 5.5 launch reaction (HN patch-review cost test, GPT-6 Sol context)
- 站内对照:DeepSeek V4.1 Flash 官宣文 · Grok 4.6 上线文 · Flash 层价格战分析 · GLM-5.3-Flash 夜间免费事件
FAQ (2026)
Are Claude Opus 5.5 and Grok 4.7 officially available?
Yes. Opus 5.5 (Sep 22, id claude-opus-5-5, $4/$20 per 1M) is on the Claude Platform, AWS, Google Cloud and Azure. Grok 4.7 (Sep 21, id grok-4.7, $2/$6 below 200k prompt tokens) is live in the xAI API.
What is GLM-5.3-FlashX?
A speed-focused variant of GLM-5.3-Flash in Z.ai's official docs and pricing page: $0.37/$1.25 per 1M, 1M context, 128K max output, text+image+video input, reasoning always on. API-only, not in the Coding Plan.
How much does GLM-5.3-FlashX cost?
$0.37 input / $0.075 cached / $1.25 output per 1M (official page). Scale: 2.5x GLM-5.3-Flash ($0.15/$0.50), about a quarter of GLM-5.3 ($1.40/$4.40), and a small fraction of Grok or Opus rates.
Is GLM-5.3-FlashX an open-weight model?
GLM-5.3-Flash is documented as open (320B/18B, sparse+linear attention). No separate FlashX weights release is verified as of Sep 24, 2026 — treat it as an API-served variant and check the license before self-hosting plans.
How do Opus 5.5, Grok 4.7 and GLM-5.3-FlashX compare on price?
Per 1M tokens: Opus 5.5 $4/$20, Grok 4.7 $2/$6 (short context), FlashX $0.37/$1.25. Blended 3:1, roughly $8 vs $3 vs $0.59 — the Chinese speed tier is an order of magnitude cheaper.
Can ChinaModelAPI route these three models?
This is a news and comparison brief, not a routing promise. Claude and Grok are Western baselines here; verify any Chinese-model route's alias, price and region in the provider dashboard before production use.