ChinaModelAPI

News / Model Watch · Coding-Tool Economics

Price war DeepSWE 73.7% $0.75 intro → $1.50 in 2027 3rd Flash in 6 weeks
2026-09-02 release · Model Watch · priced against official CN rate cards as of Sep 3

Gemini 3.8 Flash Walks Into the China Flash-Tier Price War — Cheapest Per Intelligence Unit, With a 2027 Price Trap

Google shipped its third Flash model in six weeks on Sep 2, and this one is aimed straight at the value tier Chinese labs have owned: DeepSWE v1.1 at 73.7% (second only to Claude Opus 5), 1M context, multimodal input, ~300 tok/s — at an introductory $0.75/$3.75 per 1M. Artificial Analysis calls it "the cheapest model at its level of intelligence." The catch: that rate doubles on January 1, and the model's own token appetite runs ~40% hotter than 3.7 Flash. Here's the full picture against the Chinese Flash tier.

Direct answer

Gemini 3.8 Flash (Sep 2, GA): $0.75/$3.75 per 1M through Dec 31 → $1.50/$7.50 regular from Jan 1, 2027; cache −90%; 1M context; DeepSWE v1.1 73.7% (No.2 overall, past GPT-5.6 Sol's 72.7%). AA: cheapest per intelligence unit measured ($0.58/task vs Fable 5.1's $3.76) — but per-task cost is ~40% above 3.7 Flash (more output tokens, more agent turns). China tier as of Sep 3: DeepSeek V4-Flash peak $0.44/$1.32 (off-peak half that) — still cheaper per token, and it won't double in January. GLM-5.3 and Kimi K3 (both AA intelligence 60) stay the open-weight, self-hostable picks.

What shipped (Sep 2)

  • Cadence as strategy. Third Flash in six weeks, fourth in as many months — Google iterating at Chinese-lab speed after the summer's credibility wobble (The Register's framing: 3.5 Flash was getting out-scored on intelligence by Chinese open-weight models).
  • Benchmarks. DeepSWE v1.1 73.7% — behind only Claude Opus 5 (74.0%), ahead of GPT-5.6 Sol (72.7%), Claude Sonnet 5 (53.8%), 3.7 Flash (65.3%). Also led Vals Finance Agent V2 and Harvey's Legal Agent. AA Intelligence Index: 4/4 units, ~300 tok/s output.
  • The token appetite. Google's own docs warn the model "works harder" — more reasoning steps, iterative tool calls. AA: per-task cost +40% vs 3.7 Flash on +30% output tokens. The Verge's practical note: stay on 3.7 Flash if you're minimizing tokens.
  • Cyber variant. Gemini 3.8 Flash Cyber hits 86.2% on CyberGym (vs GPT-5.5-Cyber 85.6%) and 47.2% Pass@1 on CWE-Bench automated patching — but it's restricted to trusted defenders via Google's Fairwind Program, not generally available.
  • Safety posture. Ships with CBRN and cyber-offense safeguards; output is text-only with text/image/video/speech input.

Priced against the Chinese Flash tier (per 1M tokens, Sep 3)

ModelInputOutputNotes
Gemini 3.8 Flash$0.75 → $1.50 (2027)$3.75 → $7.50intro to Dec 31; cache −90%; ~40% hotter token burn per task
DeepSeek V4-Flash (peak)$0.44$1.32off-peak ~$0.22/$0.66; no 2027 doubling; time-of-day pricing
GLM-5.3$1.40$4.40AA intelligence 60; weights on HF (MIT sibling Flash)
GLM-5.3-Flash~1/20 of GLM-5.3 pricing320B/A18B multimodal; GGUF/Ollama quants out
Claude Fable 5.1 (ref.)$10.00$50.00cache $0.25; $3.76 per AA task — 6x Gemini

Rates from official docs/rate cards as verified in earlier ChinaModelAPI briefings and re-checked Sep 3, 2026. GLM-5.3-Flash exact current card rate not re-verified today — shown as ratio per its launch terms.

The read: Google is pricing like a Chinese lab — mostly

Two things are genuinely new here. First, the intelligence-per-dollar frontier now runs through a US proprietary model: AA's $0.58 per task undercuts everything at its level, a direct shot at the economics that made DeepSeek and GLM-Flash the default budget picks. Second, Google has adopted the Chinese release cadence — four Flash models in four months is DeepSeek-style iteration, right down to the "introductory" pricing that everyone assumes will be superseded by a newer model before it expires. The honest caveats cut the other way: absolute token prices still favor the Chinese tier (V4-Flash remains ~41% cheaper on input, ~6x on output — and its price doesn't double in January), the 40% hotter token burn means the "cheap" model can bill like a mid-tier one on agentic workloads, and there are still no weights to self-host. For builders the decision tree is unchanged in shape, just tighter at the margins: meter tokens → China Flash tier; meter completed tasks and need frontier-grade quality → Gemini 3.8 Flash is now a real contender, at least until New Year's Day repricing.

Primary sources

FAQ (2026)

What launched?

Gemini 3.8 Flash, Sep 2 — third Flash in six weeks. 1M context, multimodal in, ~300 tok/s, agent/SWE positioning. A Cyber variant (86.2% CyberGym) is Fairwind-restricted.

Pricing?

$0.75/$3.75 per 1M intro through Dec 31 (cache −90%), then $1.50/$7.50 from Jan 1, 2027. Per-task cost runs ~40% above 3.7 Flash — more tokens, more turns.

Benchmarks?

DeepSWE v1.1 73.7% — No.2 behind Opus 5 (74.0%), past GPT-5.6 Sol (72.7%) and 3.7 Flash (65.3%). Led Vals Finance Agent V2 and Harvey Legal Agent.

vs Chinese Flash tier?

DeepSeek V4-Flash peak $0.44/$1.32 (off-peak half) stays cheaper per token and doesn't double in 2027. AA: Gemini is cheapest per intelligence unit ($0.58/task). GLM-5.3/K3 match it on intelligence (60) and are open-weight.

GLM-5.3 / Kimi K3 comparison?

The Register's nine-benchmark roundup: GLM-5.3 (max) 60, Kimi K3 (max) 60 — in the same frontier pack, self-hostable, with GLM-5.3-Flash at ~1/20 flagship pricing for the budget tier.

Switch or stay?

Token-metered budgets: China Flash tier still wins, esp. off-peak. Task-metered: Gemini 3.8 Flash is the new per-task price-performance leader — until Jan 1 repricing. Google itself: stay on 3.7 Flash to minimize tokens.

Related guides