Chinese AI model updates for API builders
Short, dated briefings on Chinese frontier model changes, open weights, local hardware fit, and OpenAI-compatible access. Official docs first; estimates and vendor benchmarks are labeled.
Anthropic names Alibaba, Moonshot AI and Xiaomi in Claude distillation report
151M+ exchanges attributed to Alibaba, reasoning-trace reconstruction by Moonshot, Xiaomi replays user sessions — allegations with zero API impact today. Builder read inside.
DeepSeek-V4.1-Flash official launch: deepseek-flash is live
New-architecture multimodal flash: 1M context, 890 bytes/token KV, $0.15/$0.60 per 1M off-peak, MIT weights — V4 Pro routes to it from Sep 14.
Tmall AI Space Station fills up: Kimi, MiniMax, DeepSeek join — V4 day cards from ¥4.8
Five vendors now share Tmall's token shelf. A channel event, not a price event — day-card math vs API rates for builders.
GLM-5.3-Flash runs free nightly (Sep 3–20) inside ZCode
Zhipu's build events: nightly 23:00–09:00 unlimited flash usage, 300M-token weekend drops — off-peak-free era deepens, API rates unchanged.
DeepSeek plans 160K+ Huawei Ascend 950DT cluster
Bloomberg exclusive: a ~1GW data center in Ulanqab, Inner Mongolia — the domestic-silicon capacity story under V4 API economics.
Saudi Arabia's HUMAIN-M3 is built on MiniMax M3
The 428B MoE marketed as "100% Saudi" runs on Chinese open weights, post-trained on 1T+ Arabic tokens. National-base era begins.
Qwen3.8-27B on Cerebras: 1,500 tokens/s
The Apache-2.0 27B streams at ~1,500 tok/s on wafer-scale inference — $0.99/$1.49 per 1M, preview tier. HN 392-pt reception.
Zhipu opens a Tmall flagship for GLM tokens
GLM Coding Plan from ¥118/month; Tmall AI 空间站 follows. Channel news, not a rate change.
Anthropic Writes Anti-Distillation Into Its API — the Stated Suspect Is "Industrial Scale," and China's Answer Was Already on the Shelf
New Claude API accounts can't edit around thinking blocks (Aug 31+), Anthropic's launch post calls distillation industrial-scale, and The New Stack voices the China suspicion aloud. The open-weight shelf makes parts of this fight look obsolete.
GPT-6 Astra Launches the "AGI Era" at Fable Prices — the Coding Pack Didn't Move, and China's Open-Weight Shelf Starts 20x Cheaper
OpenAI shipped GPT-6 Astra (Sep 3): OSWorld 72.6 pct, alignment escape 0 pct, staged rollout, Fable-tier pricing — and no clear coding-leadership claim. What the AGI-era flagship means vs GLM-5.3, Kimi K3 and the MIT open-weight shelf.
Gemini 3.8 Flash Walks Into the China Flash-Tier Price War — Cheapest Per Intelligence Unit, With a 2027 Price Trap
Google's third Flash in six weeks: DeepSWE 73.7% at $0.75/$3.75 intro ($1.50/$7.50 in 2027). AA: cheapest per intelligence unit — but per-task burn runs 40% above 3.7 Flash. Priced against DeepSeek V4-Flash and GLM-5.3-Flash.
Claude Fable 5.1 Cuts Cache Prices 75% — Chinese Model APIs Are Still 6-18x Cheaper
Anthropic's Sep-1 Fable 5.1 makes agentic workloads up to 45% cheaper (cache reads $0.25/1M) — and DeepSeek V4 cache hits remain ~18x cheaper still. The frontier cost war, priced against official CN rate cards.
Moonshot's 30% Ask: Kimi K3 Open-Source Model Seeks Hyperscaler Revenue Share
Reuters: Moonshot AI negotiating with Microsoft Azure, Amazon AWS, and Google Cloud for K3 hosting, seeking up to 30% revenue share. Early-stage, nothing signed — but the first Chinese-lab-to-US-hyperscaler revenue deal would reverse the software-licensing flow. Chinese media calls it 开源模型「海外收租」.
DeepSeek Opens the Vision-Exp Weights: The 1M-Context Vision Agent Goes Self-Hostable, MIT and Ungated
Ten days after the API launch, deepseek-ai posted the weights on Hugging Face — MIT, ungated, 49 shards, card benchmarks matching the API changelog 1:1. Every live DeepSeek API model now has MIT open weights.
Codex Gets Its 5-Hour Limit Back, Claude Code Cuts 25% — The Metered-API Counterpoint
OpenAI restored Codex's 5-hour cap Aug 25 after a reset-heavy August (three full resets: Aug 8/13/21); Anthropic cuts Claude Code limits from Sep 14. The verified timeline, the quota mechanics, and why pay-per-token China model APIs have no windows at all.
GLM-5.3 Open Weights Land On Schedule: The Post-Training-Scaling Flagship Goes Downloadable
zai-org/GLM-5.3 went live the night of Aug 28 — 141 safetensors shards + BF16, ungated, custom glm-5.3 license — closing the ~two-week safety-review window promised on Aug 14. Specs, license notes and what it completes on the open-weight shelf.
Qwen3.8-Flash-Next: Alibaba Ships the Qwen4 Architecture Early — Open Weights Tonight
ModelScope countdown: open-weight multimodal MoE landing tonight 23:00 Beijing — 125B main + 51B N-gram embeddings with 6B active per token, first taste of Qwen4's GDN + Qwen Sparse Attention. Confirmed-vs-unconfirmed breakdown, why 6B-active matters for local runners, and what to watch when the repo drops.
Mac mini M6 vs M5 Pro for Local LLMs: 32GB or 64GB?
Apple's M6 has the newer accelerator story, but M5 Pro doubles the memory ceiling and reaches 307GB/s bandwidth. A no-hype buyer's guide with official specs, Apple-test labels, Qwen3.8-27B fit, and the M4 upgrade decision.
Mac Studio M5 Max vs M5 Ultra for Local LLMs
M5 Max 128GB is the sensible workstation; M5 Ultra 256GB/512GB is for named models that cannot fit elsewhere. Memory math for DeepSeek V4, GLM-5, Qwen3.8, and Kimi K3, plus the limits of Apple's cluster claim.
Ox Alpha Revealed: Zhipu Confirms GLM-5.3-Flash — Official Release and Open Weights Tonight
Mystery solved: Zhipu confirmed to Bloomberg that Ox Alpha is GLM-5.3-Flash, with the official release and open weights tonight (Aug 26). The stealth run's 100T tokens/day ran on Chinese chips (SemiAnalysis); the Pony Alpha script replayed to the letter. Full saga, forensics, and reveal inside.
OpenAI-Backed Harvey Builds Its First Model on Kimi K3 Open Weights
Legal-tech leader Harvey (backed by OpenAI/Sequoia/a16z) announced Harvey Tenet — a Kimi K3 open-weight base post-trained with Fireworks for long-horizon legal work. ~2× tasks vs base K3 on Harvey's LAB at open-source cost; research preview, no published weights. The clearest Western adoption signal yet for Chinese open weights.
DeepSeek Drops Weekend Peak Pricing: Saturdays & Sundays Now All-Day Off-Peak
From Aug 23 00:00 Beijing time, weekends bill entirely at off-peak rates — v4-pro output ¥13.5/M all day vs ¥27 weekday peak. Less than a week after the peak/off-peak mechanism landed. Full weekend rate card, scheduling math, and sources.
HappyHorse 1.1 Goes Global: Full API on Model Studio, Video Editing, and a Climb Past Seedance
Alibaba opened HappyHorse 1.1's full capabilities via API on Model Studio — new video-editing mode, integrated audio at no extra cost, 40% two-week launch discount, $0.0988/s (720p) on OpenRouter, and a No.2 global video-arena rank ahead of Seedance and Sora. Platforms, pricing table, and sources inside.
GLM-5.3 API Goes GA: $1.40/$4.40, One-Day Slip, and an Ecosystem That Didn't Wait
Zhipu opened general API access in the early hours of Aug 19 — one day late, priced same as GLM-5.2, AA Index 60 (open-source co-#1). Verified launch-week timeline, pricing table vs Qwen3.8/K3/V4-Flash, platform roll call (ZCode, PhanRouter, supercomputing internet), and what changes before weights land Aug 28.
MiniMax Design: the H3 Video Model Gets an Agent Workbench
Three weeks after open-sourcing H3, MiniMax shipped the layer that tries to sell it: a natural-language creation agent workbench with a 3D director stage and ComfyUI hookup. A product, not an API — the model endpoints remain the programmatic path. Compact brief with sources.
Qwen3.8 Open Weights (2026): 2.4T Max Checkpoints + an Apache-2.0 27B for Your Laptop
Alibaba's first open-weight Max-class drop: Qwen3.8-2.4T-A95B (+FP8) on Hugging Face, plus a dense vision-language 27B under Apache 2.0 that runs from a 17GB file. Timeline, license fine print, $2/$6 API pricing, Simon Willison's local recipe, and a routing table.
Kling 3.0 API (2026): Official Pricing, Turbo & Omni, and the $3B Spin-Off
Kuaishou's Kling 3.0 family on the official international API: list prices $0.084–$0.42/s (native 4K), Turbo bundling audio, Omni adding video input, JWT task-based API shape. Q2 earnings: RMB 850M+ revenue (+200% YoY), 100M+ users, July's ~$3B raise at ~$18B valuation. Routing table + 10-second cost math vs Seedance.
DeepSeek-V4-Flash-Vision-Exp (2026): Official Launch, Pricing, API Quickstart
Launched Aug 21 evening (UTC+8): experimental multimodal vision model at V4-Flash pricing ($0.22/1M input, images capped at 384 tokens), text parity with V4-Flash, multimodal agent close to Opus-4.8, new Files API, Harness v0.1.1-rc.1 support. Includes curl/Python quickstart, cost math vs Claude/Gemini/Grok vision, and which DeepSeek model to route.
DeepSeek Harness Adds Multimodal Image Input: rc.7 + rc.8
Two-step rollout: rc.7 (Aug 17) durable image attachments in MCP/ACP; rc.8 (Aug 19) native image requests for DeepSeek adapters plus image input for /goal and /plan. The harness can now carry images — but V4-Pro/V4-Flash stay text-only, and rc.8 rides the npm next tag. Turing Post calls the harness a "second DeepSeek moment."
Gemini 3.7 Flash: Google's Workhorse Model for Coding and Agents
Google announced Gemini 3.7 Flash on Aug 13 — 1M context, DeepSWE 65.3%, intro $0.75/$3.75 per 1M tokens through 2026. A Western frontier baseline note; Gemini is not routed by ChinaModelAPI.
DeepSeek Harness v0.1: DeepSeek Open-Sources Its Agent Harness
MIT-licensed, everything-is-a-plugin agent harness (dsh) announced Aug 13 evening — overseas press calls it an open Claude Code rival. Four work modes, session import from Claude Code, 288 plugins in 24h. Model-agnostic by design.
GLM-5.3 Officially Released: Post-Training Scaling, Emergent Cyber, Weights in 2 Weeks
Z.ai's official blog announced GLM-5.3 on Aug 14 — same base as GLM-5.2, Terminal-Bench 3.0 4.6→28.3, CyberGym 84.5%, 2,436 real vulnerabilities found. Updated Aug 19: API priced same as 5.2 ($1.40/$4.40 per 1M), independent AA index 60, weights due ~Aug 28; VentureBeat / The Decoder / Interconnects coverage in.
DeepSeek-V4-Pro-0813 Official: 正式版 GA With Stronger Agents
DeepSeek's official API docs updated deepseek-v4-pro to DeepSeek-V4-Pro-0813 on 2026-08-13 (evening). Same model ID, enhanced agent capability, minor pricing adjustment. What API builders need to do — almost nothing.
Grok 4.6: xAI's New Recommended Flagship, Now in the API
xAI's official developer docs now list grok-4.6 as the latest flagship — 500k context, $2/$6 per 1M tokens, recommended for code. A Western frontier baseline note; Grok is not routed by ChinaModelAPI.
China Model Watch — Weekly Model Status
Weekly snapshot of Chinese AI model flagships tracked by ChinaModelAPI as of 2026-08-09, vs GPT-5.6 Sol / Claude Opus 5 baselines. Auto status brief for API builders.
MiniMax H3 Open Weights: Native A/V Video Model
H3 (not M3) is MiniMax’s multimodal video model with native stereo audio. HF weights + ComfyUI + hosted API. Not an official “open Kling” name — quality talk often vs Seedance.
Qwen3.8-Max: ~2.4T Flagship — Not a New Video Generator
August 2026 Qwen flagship for coding and long-horizon work. Clarifies rumors: not Qwen2.5-VL / Qwen2-VL-72B as a Seedance-class T2V release.
Seedance 2.0 Mini / Fast Discount Watch: Mini ~4折 Signals
August 2026 market note: Mini often advertised around 4折, Fast ~75折 on some platforms. Not ChinaModelAPI pricing — verify per console. Mini for iteration, 2.5 for finals.
Seedance 2.5 Capabilities: 30s One-Take, 50 Refs, Live Platforms
Official launch plus August 2026 rollout: 30s single-pass, ~50 multimodal references, stronger editing. Where creators use it and what API builders should do.
Seedance 2.5 Status (Historical): July Slip Notes
Pre-launch status brief from late July. Superseded by the official 2026-07-31 launch post — kept for timeline context only.
DeepSeek-V4-Pro for API Builders: What Replaced V3.2
DeepSeek-V4-Pro and V4-Flash are the 2026 flagship pair (1M context, open weights). How to migrate OpenAI-compatible clients from V3-era model IDs.
Kimi K3 + GLM-5.2: The 2026 Open Frontier Pairing
How Moonshot Kimi K3 and Zhipu GLM-5.2 complement each other for long-horizon coding, 1M context, and open-weight deployment — with OpenAI-compatible access notes.
data/model-watch/; publish after source checks.