ChinaModelAPI
News · Model Watch

Chinese AI model updates for API builders

Short, dated briefings on Chinese frontier model changes, open weights, local hardware fit, and OpenAI-compatible access. Official docs first; estimates and vendor benchmarks are labeled.

Distillation dispute · Sep 11

Anthropic names Alibaba, Moonshot AI and Xiaomi in Claude distillation report

151M+ exchanges attributed to Alibaba, reasoning-trace reconstruction by Moonshot, Xiaomi replays user sessions — allegations with zero API impact today. Builder read inside.

DeepSeek · Sep 10

DeepSeek-V4.1-Flash official launch: deepseek-flash is live

New-architecture multimodal flash: 1M context, 890 bytes/token KV, $0.15/$0.60 per 1M off-peak, MIT weights — V4 Pro routes to it from Sep 14.

Channel · Sep 5

Tmall AI Space Station fills up: Kimi, MiniMax, DeepSeek join — V4 day cards from ¥4.8

Five vendors now share Tmall's token shelf. A channel event, not a price event — day-card math vs API rates for builders.

GLM · Sep 5

GLM-5.3-Flash runs free nightly (Sep 3–20) inside ZCode

Zhipu's build events: nightly 23:00–09:00 unlimited flash usage, 300M-token weekend drops — off-peak-free era deepens, API rates unchanged.

Infrastructure · Sep 5

DeepSeek plans 160K+ Huawei Ascend 950DT cluster

Bloomberg exclusive: a ~1GW data center in Ulanqab, Inner Mongolia — the domestic-silicon capacity story under V4 API economics.

Sovereign AI milestone · Sep 5

Saudi Arabia's HUMAIN-M3 is built on MiniMax M3

The 428B MoE marketed as "100% Saudi" runs on Chinese open weights, post-trained on 1T+ Arabic tokens. National-base era begins.

Ecosystem launch · Sep 4

Qwen3.8-27B on Cerebras: 1,500 tokens/s

The Apache-2.0 27B streams at ~1,500 tok/s on wafer-scale inference — $0.99/$1.49 per 1M, preview tier. HN 392-pt reception.

Monetization milestone · Sep 4

Zhipu opens a Tmall flagship for GLM tokens

GLM Coding Plan from ¥118/month; Tmall AI 空间站 follows. Channel news, not a rate change.

Anti-distillation thinking-block lock watermark everywhere
2026-09-05

Anthropic Writes Anti-Distillation Into Its API — the Stated Suspect Is "Industrial Scale," and China's Answer Was Already on the Shelf

New Claude API accounts can't edit around thinking blocks (Aug 31+), Anthropic's launch post calls distillation industrial-scale, and The New Stack voices the China suspicion aloud. The open-weight shelf makes parts of this fight look obsolete.

Read →
Frontier launch GPT-6 Astra $10/$50 per 1M
2026-09-03

GPT-6 Astra Launches the "AGI Era" at Fable Prices — the Coding Pack Didn't Move, and China's Open-Weight Shelf Starts 20x Cheaper

OpenAI shipped GPT-6 Astra (Sep 3): OSWorld 72.6 pct, alignment escape 0 pct, staged rollout, Fable-tier pricing — and no clear coding-leadership claim. What the AGI-era flagship means vs GLM-5.3, Kimi K3 and the MIT open-weight shelf.

Read →
Price war Gemini 3.8 Flash DeepSWE 73.7%
2026-09-02

Gemini 3.8 Flash Walks Into the China Flash-Tier Price War — Cheapest Per Intelligence Unit, With a 2027 Price Trap

Google's third Flash in six weeks: DeepSWE 73.7% at $0.75/$3.75 intro ($1.50/$7.50 in 2027). AA: cheapest per intelligence unit — but per-task burn runs 40% above 3.7 Flash. Priced against DeepSeek V4-Flash and GLM-5.3-Flash.

Read →
Pricing shift Fable 5.1 cache −75% China gap 6-18x
2026-09-02

Claude Fable 5.1 Cuts Cache Prices 75% — Chinese Model APIs Are Still 6-18x Cheaper

Anthropic's Sep-1 Fable 5.1 makes agentic workloads up to 45% cheaper (cache reads $0.25/1M) — and DeepSeek V4 cache hits remain ~18x cheaper still. The frontier cost war, priced against official CN rate cards.

Read →
Reported (Reuters) Kimi K3 30% revenue share Azure / AWS / GCP
2026-09-01

Moonshot's 30% Ask: Kimi K3 Open-Source Model Seeks Hyperscaler Revenue Share

Reuters: Moonshot AI negotiating with Microsoft Azure, Amazon AWS, and Google Cloud for K3 hosting, seeking up to 30% revenue share. Early-stage, nothing signed — but the first Chinese-lab-to-US-hyperscaler revenue deal would reverse the software-licensing flow. Chinese media calls it 开源模型「海外收租」.

Read →
Open weights Vision-Exp MIT · ungated 1M context
2026-08-31

DeepSeek Opens the Vision-Exp Weights: The 1M-Context Vision Agent Goes Self-Hostable, MIT and Ungated

Ten days after the API launch, deepseek-ai posted the weights on Hugging Face — MIT, ungated, 49 shards, card benchmarks matching the API changelog 1:1. Every live DeepSeek API model now has MIT open weights.

Read →
Usage limits Codex 5h back Claude Code −25%
2026-08-30

Codex Gets Its 5-Hour Limit Back, Claude Code Cuts 25% — The Metered-API Counterpoint

OpenAI restored Codex's 5-hour cap Aug 25 after a reset-heavy August (three full resets: Aug 8/13/21); Anthropic cuts Claude Code limits from Sep 14. The verified timeline, the quota mechanics, and why pay-per-token China model APIs have no windows at all.

Read →
Open weights GLM-5.3 flagship 141 shards + BF16
2026-08-29

GLM-5.3 Open Weights Land On Schedule: The Post-Training-Scaling Flagship Goes Downloadable

zai-org/GLM-5.3 went live the night of Aug 28 — 141 safetensors shards + BF16, ungated, custom glm-5.3 license — closing the ~two-week safety-review window promised on Aug 14. Specs, license notes and what it completes on the open-weight shelf.

Read →
Official countdown Qwen3.8-Flash-Next 125B + 51B / A6B Qwen4 arch preview
2026-08-26

Qwen3.8-Flash-Next: Alibaba Ships the Qwen4 Architecture Early — Open Weights Tonight

ModelScope countdown: open-weight multimodal MoE landing tonight 23:00 Beijing — 125B main + 51B N-gram embeddings with 6B active per token, first taste of Qwen4's GDN + Qwen Sparse Attention. Confirmed-vs-unconfirmed breakdown, why 6B-active matters for local runners, and what to watch when the repo drops.

Read →
Hardware guideMac mini M6M5 ProLocal LLM
2026-08-26

Mac mini M6 vs M5 Pro for Local LLMs: 32GB or 64GB?

Apple's M6 has the newer accelerator story, but M5 Pro doubles the memory ceiling and reaches 307GB/s bandwidth. A no-hype buyer's guide with official specs, Apple-test labels, Qwen3.8-27B fit, and the M4 upgrade decision.

Read →
Hardware guideM5 MaxM5 Ultra512GB
2026-08-26

Mac Studio M5 Max vs M5 Ultra for Local LLMs

M5 Max 128GB is the sensible workstation; M5 Ultra 256GB/512GB is for named models that cannot fit elsewhere. Memory math for DeepSeek V4, GLM-5, Qwen3.8, and Kimi K3, plus the limits of Apple's cluster claim.

Read →
Officially confirmed Ox Alpha / 牛来 1M context · free week = GLM-5.3-Flash
2026-08-25 · Updated 2026-08-26

Ox Alpha Revealed: Zhipu Confirms GLM-5.3-Flash — Official Release and Open Weights Tonight

Mystery solved: Zhipu confirmed to Bloomberg that Ox Alpha is GLM-5.3-Flash, with the official release and open weights tonight (Aug 26). The stealth run's 100T tokens/day ran on Chinese chips (SemiAnalysis); the Pony Alpha script replayed to the letter. Full saga, forensics, and reveal inside.

Read →
Adoption milestone Harvey Tenet Kimi K3 base Research preview
2026-08-24

OpenAI-Backed Harvey Builds Its First Model on Kimi K3 Open Weights

Legal-tech leader Harvey (backed by OpenAI/Sequoia/a16z) announced Harvey Tenet — a Kimi K3 open-weight base post-trained with Fireworks for long-horizon legal work. ~2× tasks vs base K3 on Harvey's LAB at open-source cost; research preview, no published weights. The clearest Western adoption signal yet for Chinese open weights.

Read →
Pricing update Weekend = all-day off-peak Effective Aug 23
2026-08-23

DeepSeek Drops Weekend Peak Pricing: Saturdays & Sundays Now All-Day Off-Peak

From Aug 23 00:00 Beijing time, weekends bill entirely at off-peak rates — v4-pro output ¥13.5/M all day vs ¥27 weekday peak. Less than a week after the peak/off-peak mechanism landed. Full weekend rate card, scheduling math, and sources.

Read →
Official API launch HappyHorse 1.1 Video editing No.2 arena
2026-08-23

HappyHorse 1.1 Goes Global: Full API on Model Studio, Video Editing, and a Climb Past Seedance

Alibaba opened HappyHorse 1.1's full capabilities via API on Model Studio — new video-editing mode, integrated audio at no extra cost, 40% two-week launch discount, $0.0988/s (720p) on OpenRouter, and a No.2 global video-arena rank ahead of Seedance and Sora. Platforms, pricing table, and sources inside.

Read →
Official API GA GLM-5.3 $1.40/$4.40 Weights Aug 28
2026-08-22

GLM-5.3 API Goes GA: $1.40/$4.40, One-Day Slip, and an Ecosystem That Didn't Wait

Zhipu opened general API access in the early hours of Aug 19 — one day late, priced same as GLM-5.2, AA Index 60 (open-source co-#1). Verified launch-week timeline, pricing table vs Qwen3.8/K3/V4-Flash, platform roll call (ZCode, PhanRouter, supercomputing internet), and what changes before weights land Aug 28.

Read →
Product launch MiniMax Design built on H3
2026-08-22

MiniMax Design: the H3 Video Model Gets an Agent Workbench

Three weeks after open-sourcing H3, MiniMax shipped the layer that tries to sell it: a natural-language creation agent workbench with a 3D director stage and ComfyUI hookup. A product, not an API — the model endpoints remain the programmatic path. Compact brief with sources.

Read →
Open weights Qwen3.8-2.4T-A95B 27B · Apache 2.0 Local-run
2026-08-21

Qwen3.8 Open Weights (2026): 2.4T Max Checkpoints + an Apache-2.0 27B for Your Laptop

Alibaba's first open-weight Max-class drop: Qwen3.8-2.4T-A95B (+FP8) on Hugging Face, plus a dense vision-language 27B under Apache 2.0 that runs from a 17GB file. Timeline, license fine print, $2/$6 API pricing, Simon Willison's local recipe, and a routing table.

Read →
Official API Kling 3.0 / Turbo / Omni Pricing $3B spin-off
2026-08-21

Kling 3.0 API (2026): Official Pricing, Turbo & Omni, and the $3B Spin-Off

Kuaishou's Kling 3.0 family on the official international API: list prices $0.084–$0.42/s (native 4K), Turbo bundling audio, Omni adding video input, JWT task-based API shape. Q2 earnings: RMB 850M+ revenue (+200% YoY), 100M+ users, July's ~$3B raise at ~$18B valuation. Routing table + 10-second cost math vs Seedance.

Read →
Official release deepseek-v4-flash-vision-exp Pricing Quickstart
2026-08-21

DeepSeek-V4-Flash-Vision-Exp (2026): Official Launch, Pricing, API Quickstart

Launched Aug 21 evening (UTC+8): experimental multimodal vision model at V4-Flash pricing ($0.22/1M input, images capped at 384 tokens), text parity with V4-Flash, multimodal agent close to Opus-4.8, new Files API, Harness v0.1.1-rc.1 support. Includes curl/Python quickstart, cost math vs Claude/Gemini/Grok vision, and which DeepSeek model to route.

Read →
Open source DeepSeek Harness rc.7 / rc.8 Multimodal
2026-08-20

DeepSeek Harness Adds Multimodal Image Input: rc.7 + rc.8

Two-step rollout: rc.7 (Aug 17) durable image attachments in MCP/ACP; rc.8 (Aug 19) native image requests for DeepSeek adapters plus image input for /goal and /plan. The harness can now carry images — but V4-Pro/V4-Flash stay text-only, and rc.8 rides the npm next tag. Turing Post calls the harness a "second DeepSeek moment."

Read →
Western watch Gemini 3.7 Flash Google
2026-08-19

Gemini 3.7 Flash: Google's Workhorse Model for Coding and Agents

Google announced Gemini 3.7 Flash on Aug 13 — 1M context, DeepSWE 65.3%, intro $0.75/$3.75 per 1M tokens through 2026. A Western frontier baseline note; Gemini is not routed by ChinaModelAPI.

Read →
Open source DeepSeek Harness v0.1 Agent harness · MIT
2026-08-14

DeepSeek Harness v0.1: DeepSeek Open-Sources Its Agent Harness

MIT-licensed, everything-is-a-plugin agent harness (dsh) announced Aug 13 evening — overseas press calls it an open Claude Code rival. Four work modes, session import from Claude Code, 288 plugins in 24h. Model-agnostic by design.

Read →
Official release GLM-5.3 Zhipu / Z.ai
2026-08-14 · Updated 2026-08-19

GLM-5.3 Officially Released: Post-Training Scaling, Emergent Cyber, Weights in 2 Weeks

Z.ai's official blog announced GLM-5.3 on Aug 14 — same base as GLM-5.2, Terminal-Bench 3.0 4.6→28.3, CyberGym 84.5%, 2,436 real vulnerabilities found. Updated Aug 19: API priced same as 5.2 ($1.40/$4.40 per 1M), independent AA index 60, weights due ~Aug 28; VentureBeat / The Decoder / Interconnects coverage in.

Read →
Official release DeepSeek-V4-Pro-0813 DeepSeek
2026-08-13

DeepSeek-V4-Pro-0813 Official: 正式版 GA With Stronger Agents

DeepSeek's official API docs updated deepseek-v4-pro to DeepSeek-V4-Pro-0813 on 2026-08-13 (evening). Same model ID, enhanced agent capability, minor pricing adjustment. What API builders need to do — almost nothing.

Read →
Western watch Grok 4.6 xAI
2026-08-13

Grok 4.6: xAI's New Recommended Flagship, Now in the API

xAI's official developer docs now list grok-4.6 as the latest flagship — 500k context, $2/$6 per 1M tokens, recommended for code. A Western frontier baseline note; Grok is not routed by ChinaModelAPI.

Read →
Weekly Model Watch Auto
2026-08-09

China Model Watch — Weekly Model Status

Weekly snapshot of Chinese AI model flagships tracked by ChinaModelAPI as of 2026-08-09, vs GPT-5.6 Sol / Claude Opus 5 baselines. Auto status brief for API builders.

Read →
Open weights MiniMax H3 Video
2026-08-08

MiniMax H3 Open Weights: Native A/V Video Model

H3 (not M3) is MiniMax’s multimodal video model with native stereo audio. HF weights + ComfyUI + hosted API. Not an official “open Kling” name — quality talk often vs Seedance.

Read →
Flagship LLM Qwen3.8-Max Alibaba
2026-08-08

Qwen3.8-Max: ~2.4T Flagship — Not a New Video Generator

August 2026 Qwen flagship for coding and long-horizon work. Clarifies rumors: not Qwen2.5-VL / Qwen2-VL-72B as a Seedance-class T2V release.

Read →
Market watch Seedance 2.0 Mini Pricing
2026-08-08

Seedance 2.0 Mini / Fast Discount Watch: Mini ~4折 Signals

August 2026 market note: Mini often advertised around 4折, Fast ~75折 on some platforms. Not ChinaModelAPI pricing — verify per console. Mini for iteration, 2.5 for finals.

Read →
Capabilities update Seedance 2.5 ByteDance Dreamina
2026-07-31 · Updated 2026-08-08

Seedance 2.5 Capabilities: 30s One-Take, 50 Refs, Live Platforms

Official launch plus August 2026 rollout: 30s single-pass, ~50 multimodal references, stronger editing. Where creators use it and what API builders should do.

Read →
Seedance 2.5 ByteDance Historical
2026-07-27 · superseded

Seedance 2.5 Status (Historical): July Slip Notes

Pre-launch status brief from late July. Superseded by the official 2026-07-31 launch post — kept for timeline context only.

Read →
DeepSeekAPIOpen Weights
2026-07-27

DeepSeek-V4-Pro for API Builders: What Replaced V3.2

DeepSeek-V4-Pro and V4-Flash are the 2026 flagship pair (1M context, open weights). How to migrate OpenAI-compatible clients from V3-era model IDs.

Read →
Kimi K3GLM-5.2Open Weights
2026-07-27

Kimi K3 + GLM-5.2: The 2026 Open Frontier Pairing

How Moonshot Kimi K3 and Zhipu GLM-5.2 complement each other for long-horizon coding, 1M context, and open-weight deployment — with OpenAI-compatible access notes.

Read →
Cadence: weekly Model Watch + urgent posts on major flagship drops. Automated drafts land in data/model-watch/; publish after source checks.