ChinaModelAPI

News / Model Watch · Coding-Tool Economics

Pricing shift cache reads −75% agent costs −45% China gap: 6-18x on cache
2026-09-02 · Model Watch · Anthropic newsroom announcement Sep 1, priced against official CN rate cards

Claude Fable 5.1 Cuts Cache Prices 75%: The Frontier Cost War Arrives — and Chinese Model APIs Are Still 6-18x Cheaper

The morning after Labor Day, Anthropic shipped Claude Fable 5.1 (plus the safeguard-restricted Mythos 5.1) — and the headline isn't intelligence, it's economics: cache reads down 75% to $0.25/1M, typical workloads ~25% cheaper, agentic workloads up to ~45% cheaper. That's a direct answer to the "agentic cost" problem this summer's quota-turbulence cycle exposed. Here's what shipped, and what the same economics look like on the Chinese model API shelf.

Direct answer

Claude Fable 5.1 is live (Sep 1): $10/$50 per 1M in/out unchanged, cache reads cut 75% to $0.25/1M — worth up to ~45% off for agent workloads that re-feed context. Mythos 5.1 is the trusted-access twin; biology safeguards fire 85% less on benign queries; Fable can now help identify (not exploit) vulnerabilities; output carries a statistical watermark everywhere. The China comparison as of Sep 2: DeepSeek V4-Flash cache hits are ~18x cheaper ($0.014), V4-Pro ~6x ($0.044), GLM-5.3 fresh input is 1/7th the price ($1.40) — the cut narrows the gap by a constant, not an order of magnitude.

What shipped (Sep 1)

  • Two configurations, one model. Fable 5.1 is generally available — API id claude-fable-5-1, on AWS, Google Cloud and Azure from day one. Mythos 5.1 is the same base with looser safeguards for trusted-access programs (cybersecurity, life sciences).
  • The pricing story. Base rates hold at $10/$50 per 1M. Cache reads — where an agent re-sends context it already processed — drop 75% to $0.25. Anthropic's math: typical workloads −25%, highly agentic workloads up to −45%.
  • Friction reduction. New biology safeguards trigger 85% less often on benign elementary biology/medical questions. Fable 5.1 is cleared for vulnerability identification; exploit generation and pentesting stay on Opus-class models.
  • Enterprise privacy architecture. Enterprise Frontier Safeguards stores customer data in customer-controlled cloud infrastructure — zero-data-retention while automated monitoring continues. Built with 100+ customers, phased rollout from this fall.
  • Provenance. Fable/Mythos output carries Anthropic's statistical text watermark on every platform, not just in the EU.

Sources: Anthropic newsroom announcement (Sep 1, title and availability verified directly); pricing and safeguard details per Unite.AI and ModemGuides coverage citing Anthropic's What's-new page and platform docs, with Bloomberg coverage framing the same commercial read. As-of Sep 2, 2026.

Priced against the Chinese shelf (per 1M tokens, as of Sep 2)

ModelInputOutputCached input
Claude Fable 5.1$10.00$50.00$0.25 (−75%)
GLM-5.3 (Z.ai GA)$1.40$4.40
DeepSeek V4-Flash (peak)$0.44$1.32$0.014
DeepSeek V4-Pro (peak)$1.32$3.96$0.044

CN rates from official docs as verified on earlier ChinaModelAPI briefings (GLM-5.3 GA, V4-Pro release); DeepSeek off-peak rates are ~50% lower still. GLM cached-input rate not yet verified — omitted rather than guessed.

The read for API builders

Anthropic's own launch framing — the competition moving to "who can complete the largest useful unit of work at the lowest acceptable cost and risk" — is a concession that cost is now the battleground. Fable 5.1's cache cut is the first big Western move aimed squarely at agentic economics. But do the arithmetic: a 75% cut on $1 cache reads lands at $0.25; DeepSeek's peak-hour cache hits were already at $0.014 — an 18x gap that survives the cut. Fresh-token gaps run 7x (vs GLM-5.3 input) to 23x (vs V4-Flash output). The practical split hasn't changed as of Sep 2: frontier intelligence with a newly cheaper context loop on one side; transparent per-token pricing, no usage windows, and off-peak discounts on the other. The interesting question for the next quarter is whether the cache-price war goes another round — and whether Chinese labs respond at all, given they're not the ones under margin pressure.

Primary sources

FAQ (2026)

What shipped Sep 1?

Claude Fable 5.1 (GA, claude-fable-5-1, on AWS/GCP/Azure) + Mythos 5.1 (trusted-access twin for cyber/bio work). Anthropic's most advanced coding/knowledge models.

Pricing changes?

$10/$50 per 1M in/out unchanged; cache reads −75% to $0.25. Typical workloads ~−25%, highly agentic ones up to ~−45%.

vs Chinese APIs (Sep 2)?

Cache: Fable $0.25 vs DeepSeek V4-Flash $0.014 (~18x) and V4-Pro $0.044 (~6x). Fresh input: GLM-5.3 $1.40 (7x less), V4-Flash $0.44 (23x less).

Safeguard changes?

Biology guardrails fire 85% less on benign queries; Fable cleared for vulnerability identification (exploits stay Opus-class); Enterprise Frontier Safeguards zero-retention rolls out this fall.

Watermarked?

Yes — statistical text watermark on all Fable/Mythos output, every platform (not EU-only). Anthropic published the technical explainer Aug 14.

Should agentic stacks switch?

Gap narrowed, ranking intact: budget-sensitive pipelines still land on DeepSeek/GLM per-token economics; frontier-intelligence loops just got a cheaper context re-feed on Fable.

Related guides