ChinaModelAPI

News / Model Watch · Model launch

Official launch · Sep 21 500K context same rates as 4.6 xhigh token-use caveat
2026-09-21 launch · rechecked 2026-09-24 UTC+8 · xAI launch post + docs.x.ai + Artificial Analysis

Grok 4.7 API: Coding-Frontier Rates Unchanged at $2/$6 per 1M

xAI shipped Grok 4.7 on September 21, a day before Anthropic's Opus 5.5 and roughly two weeks later than teased. The pitch is value, not peak: $2 / $6 per 1M tokens below a 200K-prompt threshold, unchanged from Grok 4.6, on a 500K-token multimodal-input window. Independent scoring puts it in the top four labs but not at the frontier, and the number that matters for agent builders is not on the rate card: at xhigh effort it burns roughly 81,000 output tokens per benchmark task, about 2.25x Grok 4.6. This brief separates the published rates from the measured behavior.

Direct answer

What shipped: Grok 4.7, live Sep 21, 2026 in the xAI API, id grok-4.7, built for coding and knowledge work; 500K context, text+image input, May 2026 cutoff, effort ladder low to xhigh (high default). The rate card: $2 in / $0.50 cached / $6 out per 1M below 200K prompt tokens, doubling to $4 / $1 / $12 above; a 1.1x regional multiplier applies on US-regional endpoints (docs phrase the effective rates as $2.20 / $0.55 / $6.60). Grok 4.7 Fast (2x rates) exists only inside Cursor and Grok Build. Buyer takeaway: roughly a third of Opus 5.5's output price for clearly sub-frontier agentic scores; budget for verbosity at xhigh, and benchmark cost-per-completed-task, not price-per-token, before routing agent loops to it.

The facts, as verified (Sep 24)

  • Rates identical to Grok 4.6, threshold and all. Below 200K prompt tokens: $2 / $0.50 cached / $6 per 1M; at or above 200K: $4 / $1 / $12. docs.x.ai additionally documents a 1.1x multiplier for US-regional processing scope, applied to input, output and cached tokens after caching discounts. Grok 4.6's card was the same, so this launch holds price while claiming capability.
  • Specs. 500K-token context window, multimodal (text and image) input, text output, knowledge cutoff May 2026. Reasoning effort runs low / medium / high (default) / xhigh. Grok 4.7 Fast, the same model on faster infrastructure at exactly 2x rates ($4/$12 below threshold), is available only through Cursor and Grok Build, not the public API, and xAI has published no tokens-per-second figure for it.
  • Benchmarks, vendor side. xAI-reported: CursorBench 4.0 46.3% and Terminal-Bench 4.0 38.0% at xhigh effort. Most of xAI's own comparison charts pit Grok 4.7 at xhigh against Grok 4.6 at high, an apples-to-oranges framing worth noting when reading them.
  • Benchmarks, independent side. Artificial Analysis scores Grok 4.7 at 46 on its Intelligence Index v4.3 (Grok 4.6: 44), enough to place xAI in the top four AI labs, and reports its Coding Agent Index overtaking GPT-5.6 Sol. AA measured the hallucination rate improving to 29% from 34% with accuracy broadly unchanged (47% vs 48%). A third-party-cited independent Terminal-Bench rerun scored it near 26%.
  • The token-use catch. Independent measurement of the standard intelligence benchmark suite puts Grok 4.7 (xhigh) at roughly 81,000 output tokens per task versus 36,000 for Grok 4.6, about 125% more, and nearly triple a leading OpenAI comparison model. On output-billed agentic loops, that can erase the per-token savings entirely.
  • Speed, two readings. Artificial Analysis measured roughly 188 output tokens/s for long prompts (about 7.1 minutes per index task); another independent harness ranks it slow at about 39.5 t/s (152nd of 202 models) counting full generation. The harnesses measure different things; treat both as unverified until you run your own.
  • Launch texture. The Hacker News thread notes Grok 4.7 carries roughly 40% more weights than Grok 4.6 at the same price, and that the launch slipped about two weeks past Musk's early-September teaser, landing the day before Opus 5.5.

Prices are provider-published per-1M figures from docs.x.ai read on Sep 23-24, 2026. Benchmark and speed figures are labeled by who ran them; no benchmark purchase or API test was performed for this article.

Context: the value slot in a repriced week

Grok 4.7 opened the 48-hour window that Anthropic's Opus 5.5 (Sep 22, our launch brief) and OpenAI's 50% GPT-6 Sol/Luna cuts closed. On the only common yardstick, Terminal-Bench 4.0, the gap is not close: 38.0% xAI-reported against Opus 5.5's 66.4%. But the price gap runs the other way, and Cursor-side analyses put Grok xhigh within fractions of Opus-class models on their internal sessions at roughly half the task cost.

The slot it occupies is crowded from the Chinese side too: at exactly $2/$6, Grok 4.7 shares its rate card with Qwen3.8-Max and MiMo-V2.5 Pro, while GLM-5.3 undercuts it at $1.40/$4.40 and the speed tiers sit an order of magnitude lower. Our Grok 4.6 launch coverage holds the baseline this builds on; the week's combined view is in the tri-model brief.

One practical note for anyone comparing against xAI's marketing: the "twice as fast, half the price of comparable models" line refers to the rate card and the Fast variant, while independent token-use measurements point the other way on real task cost. Route on measured cost per completed task, not on either headline.

What this means for Chinese-model buyers

Blended 3:1, Grok 4.7 costs about $3 per 1M tokens against roughly $0.59 for GLM-5.3-FlashX (our FlashX brief) and near $0.24 for DeepSeek V4.1 Flash. As with Opus 5.5, this is not a cost play for anyone already routing Chinese models; it is a capability purchase for workloads that need Western-frontier behavior at a mid price.

Two comparisons actually matter for our readers. First, against Qwen3.8-Max at the identical $2/$6: same card, open weights on the Qwen side, so self-hosting teams get a hedge Grok cannot offer. Second, the verbosity pattern: the xhigh token-use finding is a general reminder that reasoning-tier models everywhere (GLM included, with thinking always on in FlashX) bill the thinking, so effort-level tuning belongs in every cost model. The full ladder is in the cross-vendor comparison table.

If you evaluate Grok 4.7 for repository-scale work: the 200K doubling threshold means a 500K-context session bills every token at $4/$12, and the US-regional multiplier can add another 10%. Quote your real prompt distribution before comparing cards.

Access from outside China — and the usual disclaimer

This site tracks Chinese model APIs for international builders, with Western launches covered as comparison baselines. Grok 4.7 is available through the xAI API and, in its Fast variant, through Cursor and Grok Build.

Independence disclaimer: ChinaModelAPI is an independent information site, not affiliated with xAI or any model vendor. Figures here are provider-published or explicitly attributed to the independent outlets that measured them; benchmark claims are labeled by who ran them, and nothing on this page is a routing or availability promise.

Primary sources

FAQ (2026)

What is Grok 4.7 and when did it launch?

xAI's model for coding and knowledge work, released Sep 21, 2026 (about two weeks past the teased date). API id grok-4.7, 500K context, multimodal input, May 2026 cutoff.

How much does the Grok 4.7 API cost?

$2/$0.50/$6 per 1M below 200K prompt tokens, doubling above; same card as Grok 4.6. A 1.1x US-regional multiplier can apply. Fast variant: 2x rates, Cursor/Grok Build only.

What context window does Grok 4.7 have?

500K tokens, text+image input. Note the pricing threshold: prompts at or above 200K tokens bill at the doubled $4/$12 card, so full-window sessions cost double.

How does Grok 4.7 benchmark?

xAI: CursorBench 46.3%, TB4 38.0% (xhigh). AA independent: Intelligence Index 46 (top-4 labs), hallucination 29% vs 34%. A third-party TB4 rerun scored it near 26%.

Is Grok 4.7 fast, and what about token use?

AA measured ~188 t/s answer output; another harness ranks it slow at ~39.5 t/s. The cost lever is verbosity: ~81K output tokens/task at xhigh vs 36K for 4.6, which can erase per-token savings.

How does Grok 4.7 compare with Chinese models on price?

Blended ~$3 per 1M vs FlashX ~$0.59 and GLM-5.3 $2.15; it shares $2/$6 with Qwen3.8-Max and MiMo-V2.5 Pro. Covered as a Western baseline, not a routing promise.

Related guides