ChinaModelAPI

News / Model Watch · Model launch

Official launch · Sep 22 1M context thinking always on -20% vs Opus 5
2026-09-22 launch · rechecked 2026-09-24 UTC+8 · Anthropic announcement + platform docs

Claude Opus 5.5 API: Fable-Class Output at $4/$20 per 1M Tokens

Anthropic's first Claude 5.5-family model arrived September 22 with a simple pitch: near-Fable-5.1 performance at 40% less run cost than Opus 5. The rate card is a clean 20% cut on every standard line ($4 in / $20 out, cache reads down 60% to $0.20), the 1M-token context and 128K max output carry over, and one behavioral break matters more than any price: thinking is now always on and cannot be disabled. This brief separates the published facts from the vendor benchmark story, and sets the launch against what Chinese-model buyers actually pay.

Direct answer

What shipped: Claude Opus 5.5, live Sep 22, 2026 on the Claude Platform, Bedrock, Google Cloud and Microsoft clouds, API id claude-opus-5-5. The rate card: $4 / $20 per 1M tokens standard; cache read $0.20, cache write $5 (5-min) / $8 (1-hour); Batch API half price at $2/$10; fast mode $8/$40 at up to 2.5x speed. Specs: 1M context, 128K max output (300K on Batches beta), text+image input, June 2026 cutoff, thinking adaptive and always on with default effort now medium. Buyer takeaway: near-frontier quality at mid-tier cost, but always-on thinking changes your token math; blended 3:1, Opus 5.5 runs about $8 per 1M against roughly $0.59 for GLM-5.3-FlashX, so it is a quality purchase, not a cost saving.

The facts, as verified (Sep 24)

  • Standard pricing, every line cut 20%. Input $4 (was $5), output $20 (was $25), 5-minute cache write $5 (was $6.25). Cache reads drop hardest: $0.20 per 1M, down 60% from Opus 5's $0.50, which matters most for the long-horizon agent loops this class of model is bought for. Batch API runs at $2/$10.
  • Context and output carry over. 1M-token context window and 128K max output, default on with no beta header, per Anthropic's platform docs; the Batch API beta allows 300K output tokens. Input is text plus images; knowledge cutoff June 2026. Anthropic commits to retirement no sooner than September 22, 2027.
  • Thinking is always on, a real breaking change. The mode cannot be switched off, and the default effort level moves from high to medium on a low-to-max ladder. Billing is output-token based, so existing prompts that assumed controllable thinking will price differently even when the code is untouched. Anthropic's docs flag four documented breaking changes for code running on Opus 5.
  • Vendor benchmarks, labeled as such. Terminal-Bench 4.0: 66.4% at xhigh effort against Opus 5 at 52.3%, Fable 5.1 at 55.8%, GPT-6 Astra at 57.9%, GPT-5.6 Sol at 37.3%. GDPval-AA v2.1: 1846 Elo (Fable 5.1 1735, Opus 5 1708). CursorBench 4.0: 57.8%. Humanity's Last Exam: 67.7% with tools. All are Anthropic-run numbers; no independent Terminal-Bench rerun has appeared yet.
  • One real-world cost datapoint from the field. A developer patch-review harness posted in the launch's Hacker News thread (1152 points) measured Opus 5.5 finding 8 of 14 known issues for $15.40, where Fable 5.1 found 7 for $66.34 and Opus 5 found 6 for $15.19: near-frontier accuracy at mid-tier cost in one n=1 test. Anthropic also cites a 680,000-line code migration completed in under a day.
  • Safeguards and data terms. Opus 5.5 ships with preserved thinking, the anti-distillation measure introduced with Fable 5.1, applying to API accounts created on or after August 31, 2026. Zero data retention remains available, and EU AI Act watermarking measures carry over. On the safety side, Anthropic reports about 85% fewer containment-boundary circumvention attempts than Opus 5 / Mythos 5.1, with cyber-flagged requests falling back to Opus 4.8 and biology/frontier-LLM requests to Opus 5.

All prices are provider-published per-1M figures read from Anthropic's announcement and platform pricing pages on Sep 22-24, 2026. Benchmark figures are vendor-run unless attributed otherwise; no benchmark purchase was made for this article.

Context: the 48-hour repricing window

Opus 5.5 landed one day after xAI's Grok 4.7 (Sep 21, our launch brief) and in the same window as OpenAI's 50% GPT-6 price cuts with Sol and Luna, so the top of the Western market repriced inside two days. Terminal-Bench 4.0 is the only benchmark all three vendors reported, which makes it the least-bad common yardstick this month; on it, Opus 5.5's 66.4% sits well clear of Grok 4.7's 38.0% (xAI-reported) and everything below.

The internal ladder matters too: Fable 5.1 keeps the frontier slot at $10/$50, Opus 5.5 takes the "most workloads" recommendation at $4/$20, and Sonnet 5 holds $2/$10. For teams already on Claude, the migration question is not whether to move from Opus 5 (same slot, cheaper, faster output) but how always-on thinking reshapes spend: medium default effort plus 30% faster standard output is a real throughput change, not just a price cut.

For the distillation dispute backdrop that surrounds everything Anthropic ships this quarter, see our coverage of the September threat report naming Alibaba, Moonshot and Xiaomi; the combined three-model view of this week is in the tri-model brief.

What this means for Chinese-model buyers

Blended at a typical 3:1 input-output ratio, Opus 5.5 costs about $8 per 1M tokens. The Chinese speed tier sits an order of magnitude below it: GLM-5.3-FlashX at roughly $0.59 blended (our FlashX brief), GLM-5.3 flagship at $2.15, DeepSeek V4.1 Flash near $0.24. Nobody routing Chinese-model traffic switches to Opus 5.5 to save money.

What the launch does change is the benchmark anchor. When the Western "good enough for most work" tier posts TB4 66.4%, the honest question for a Chinese-model route is where your workload actually sits: interactive coding agents and batch classification rarely need frontier agentic scores, and the cache-read price war (Opus 5.5 at $0.20) shows even Anthropic now competes hardest on long-horizon loop cost, the exact territory where GLM-5.3-Flash's $0.03 cached input and DeepSeek's off-peak rates already live. The current cross-vendor comparison table holds the full ladder.

Practical note for teams evaluating both sides: preserved thinking applies only to Claude API accounts created on or after August 31, 2026, so older accounts see different behavior than new ones; check your account vintage before comparing distillation-resistance claims.

Access from outside China — and the usual disclaimer

This site tracks Chinese model APIs for international builders, with Western launches covered as comparison baselines. Opus 5.5 is available through Anthropic and the three hyperscalers directly; regional terms, throughput and billing vary by channel, so verify the contract you actually use.

Independence disclaimer: ChinaModelAPI is an independent information site, not affiliated with Anthropic or any model vendor. Figures here are provider-published or explicitly attributed to independent outlets; performance claims are vendor claims unless marked otherwise, and nothing on this page is a routing or availability promise.

Primary sources

FAQ (2026)

What is Claude Opus 5.5 and when did it launch?

The first Claude 5.5-family model, live Sep 22, 2026 on the Claude Platform, Bedrock, Google Cloud and Microsoft clouds. API id claude-opus-5-5; Sonnet 5.5 and Haiku 5.5 to follow.

How much does Claude Opus 5.5 cost?

$4 / $20 per 1M standard; cache read $0.20, writes $5 (5-min) / $8 (1-hour); Batch API $2/$10; fast mode $8/$40 at up to 2.5x speed. Every standard line is 20% under Opus 5.

What context window and output limits does Opus 5.5 have?

1M context and 128K max output by default (300K on the Batch API beta). Text+image input, June 2026 knowledge cutoff, retirement committed no sooner than Sep 22, 2027.

Can thinking mode be disabled on Opus 5.5?

No. Thinking is adaptive and always on, and default effort moves from high to medium. Output-token billing changes even for untouched prompts; Anthropic documents four breaking changes versus Opus 5.

How does Opus 5.5 compare with Fable 5.1 and Opus 5?

Positioned at Fable-5.1-level performance at 40% lower run cost than Opus 5: TB4 66.4% vs 52.3% (Opus 5) and 55.8% (Fable 5.1); GDPval-AA 1846 Elo. Fable 5.1 keeps the frontier slot at $10/$50.

Can ChinaModelAPI route Claude Opus 5.5?

This is a news and comparison brief, not a routing promise. Opus 5.5 is a Western baseline here; verify any Chinese-model route's alias, price and region in the provider dashboard before production use.

Related guides