ChinaModelAPI

News / Model Watch · Qwen

Official release coding + agent upgrade $2 / $6 per 1M 1M context retained
2026-09-03 · Model Watch · QwenCloud model page + official @Alibaba_Qwen post, Sep 2

Qwen3.8-Max-0902: Alibaba Upgrades Its 2.4T Flagship for Coding and Agents — Same $2/$6 Slot, Cache at $0.25

Three weeks after open-sourcing the weights, Qwen shipped qwen3.8-max-0902 — an upgraded API snapshot (alias qwen3.8-max-2026-09-02) of the 2.4-trillion-parameter flagship. The official page leads with engineering-scale coding, multi-tool agent orchestration, and refined vision, keeping the 1M context and thinking mode. Priced at $2/$6 per 1M with implicit cache reads at $0.25, it lands squarely in frontier-coding price territory — a week when Gemini 3.8 is arriving and community threads are already ranking Kimi K3 and GLM-5.3 against it.

Direct answer

Qwen3.8-Max-0902 is a Sep-2 checkpoint upgrade of Qwen's flagship, live now on QwenCloud — model id qwen3.8-max-0902. Official highlights: coding gains on complex engineering-scale and long-horizon autonomous projects, significantly stronger collaborative-agent behavior (multi-tool orchestration, end-to-end delivery), and refined native vision (charts, documents, multimodal perception). It accepts text, image and video input, retains the 1M context window and thinking mode, and costs $2 input / $6 output per 1M, implicit cache $0.25, explicit cache creation $2.50.

What the official page confirms

  • Identity: qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02), described as "an upgraded snapshot of qwen3.8-max" — same 2.4T base, new post-training.
  • Coding: "breaks new ground, handling more complex engineering-scale projects and long-horizon autonomous development."
  • Agents: collaborative-agent performance "significantly enhanced, with greater composure in multi-tool orchestration and end-to-end task delivery" — the Cowork direction Qwen has been building all summer.
  • Vision: native vision refined across chart reasoning, document parsing, and multimodal perception; input modalities are text, image and video.
  • Continuity: 1M context window, thinking mode, and the full tool ecosystem all retained; a prefix-completion feature is listed for partial-mode calls.
  • Announcement: official @Alibaba_Qwen post (Sep 2) plus the QwenCloud model page; two HN threads picked it up within hours.

Pricing against the Chinese-model shelf

Model (per 1M tokens) Input Output Cache read
Qwen3.8-Max-0902$2.00$6.00$0.25 (implicit)
GLM-5.3$1.40$4.40$0.26 (cached)
DeepSeek V4-Pro (peak)$1.74$3.48$0.044 (hit)
DeepSeek V4-Flash (peak)$0.44$1.32$0.014 (hit)

Official rate cards as of Sep 3, 2026. DeepSeek peak/off-peak tiers: off-peak rates are roughly half; weekend off-peak billing applies. GLM-5.3 cache price from docs.z.ai.

The read: 0902 is not playing the budget game. Its $2/$6 slot sits above GLM-5.3 and DeepSeek V4-Pro on raw price, with implicit cache at $0.25 — nearly identical to GLM-5.3's $0.26 but ~6x DeepSeek's cache-hit economics. The bet is that coding-plus-agent capability clears the premium, the same wager Anthropic made with Fable 5.1's cache cuts two days earlier.

Community reception — and one claim to hold loosely

  • HN: two threads inside 24h — "Qwen3.8-Max Checkpoint 0902" (9 pts) and "Qwen3.8-Max just got upgraded" (6 pts) linking the official tweet and model page.
  • Single-source benchmark claim: CN tech media reports the 0902 coding gains surpassing Claude "Opus"-class models. No independent benchmark round-trip has landed yet — treat as vendor-adjacent until Artificial Analysis or equivalent publishes numbers. We label it, we don't bank on it.
  • Timing context: Google's Gemini 3.8 arrives this week, and an HN post already argues "Kimi K3 and GLM-5.3 are better than Gemini 3.8 Flash." The frontier-coding shelf is being restocked simultaneously on three continents — that is the frame for 0902, not a vacuum.
  • Weights nuance: the open 2.4T checkpoint from August 12 remains text-only without native 1M context; 0902's multimodal input is API-side. Self-hosters waiting for an upgraded open snapshot have no date yet.

Primary sources

FAQ (2026)

What is it?

A Sep-2 upgraded snapshot of Qwen's 2.4T flagship — id qwen3.8-max-0902, live on QwenCloud, focused on coding, agents and vision.

What improved?

Engineering-scale coding, multi-tool agent orchestration, end-to-end task delivery, chart/document vision — with 1M context and thinking mode retained.

Pricing?

$2/$6 per 1M in/out; implicit cache reads $0.25; explicit cache creation $2.50. Prefix completion supported.

vs the open weights?

Different artifacts: Aug-12 open checkpoint is text-only, no native 1M; 0902 is the multimodal API snapshot on the same base.

vs GLM-5.3 / DeepSeek?

Pricier than both on raw tokens (GLM $1.40/$4.40; V4-Flash peak $0.44/$1.32) — the pitch is frontier coding-agent capability, not budget.

On relays yet?

Live on official QwenCloud API; relay routing typically follows within days — join the waitlist for 0902 access updates.

Related guides