News / Model Watch · Coding-Tool Economics
Gemini 3.8 Flash Walks Into the China Flash-Tier Price War — Cheapest Per Intelligence Unit, With a 2027 Price Trap
Google shipped its third Flash model in six weeks on Sep 2, and this one is aimed straight at the value tier Chinese labs have owned: DeepSWE v1.1 at 73.7% (second only to Claude Opus 5), 1M context, multimodal input, ~300 tok/s — at an introductory $0.75/$3.75 per 1M. Artificial Analysis calls it "the cheapest model at its level of intelligence." The catch: that rate doubles on January 1, and the model's own token appetite runs ~40% hotter than 3.7 Flash. Here's the full picture against the Chinese Flash tier.
Gemini 3.8 Flash (Sep 2, GA): $0.75/$3.75 per 1M through Dec 31 → $1.50/$7.50 regular from Jan 1, 2027; cache −90%; 1M context; DeepSWE v1.1 73.7% (No.2 overall, past GPT-5.6 Sol's 72.7%). AA: cheapest per intelligence unit measured ($0.58/task vs Fable 5.1's $3.76) — but per-task cost is ~40% above 3.7 Flash (more output tokens, more agent turns). China tier as of Sep 3: DeepSeek V4-Flash peak $0.44/$1.32 (off-peak half that) — still cheaper per token, and it won't double in January. GLM-5.3 and Kimi K3 (both AA intelligence 60) stay the open-weight, self-hostable picks.
What shipped (Sep 2)
- Cadence as strategy. Third Flash in six weeks, fourth in as many months — Google iterating at Chinese-lab speed after the summer's credibility wobble (The Register's framing: 3.5 Flash was getting out-scored on intelligence by Chinese open-weight models).
- Benchmarks. DeepSWE v1.1 73.7% — behind only Claude Opus 5 (74.0%), ahead of GPT-5.6 Sol (72.7%), Claude Sonnet 5 (53.8%), 3.7 Flash (65.3%). Also led Vals Finance Agent V2 and Harvey's Legal Agent. AA Intelligence Index: 4/4 units, ~300 tok/s output.
- The token appetite. Google's own docs warn the model "works harder" — more reasoning steps, iterative tool calls. AA: per-task cost +40% vs 3.7 Flash on +30% output tokens. The Verge's practical note: stay on 3.7 Flash if you're minimizing tokens.
- Cyber variant. Gemini 3.8 Flash Cyber hits 86.2% on CyberGym (vs GPT-5.5-Cyber 85.6%) and 47.2% Pass@1 on CWE-Bench automated patching — but it's restricted to trusted defenders via Google's Fairwind Program, not generally available.
- Safety posture. Ships with CBRN and cyber-offense safeguards; output is text-only with text/image/video/speech input.
Priced against the Chinese Flash tier (per 1M tokens, Sep 3)
| Model | Input | Output | Notes |
|---|---|---|---|
| Gemini 3.8 Flash | $0.75 → $1.50 (2027) | $3.75 → $7.50 | intro to Dec 31; cache −90%; ~40% hotter token burn per task |
| DeepSeek V4-Flash (peak) | $0.44 | $1.32 | off-peak ~$0.22/$0.66; no 2027 doubling; time-of-day pricing |
| GLM-5.3 | $1.40 | $4.40 | AA intelligence 60; weights on HF (MIT sibling Flash) |
| GLM-5.3-Flash | ~1/20 of GLM-5.3 pricing | 320B/A18B multimodal; GGUF/Ollama quants out | |
| Claude Fable 5.1 (ref.) | $10.00 | $50.00 | cache $0.25; $3.76 per AA task — 6x Gemini |
Rates from official docs/rate cards as verified in earlier ChinaModelAPI briefings and re-checked Sep 3, 2026. GLM-5.3-Flash exact current card rate not re-verified today — shown as ratio per its launch terms.
The read: Google is pricing like a Chinese lab — mostly
Two things are genuinely new here. First, the intelligence-per-dollar frontier now runs through a US proprietary model: AA's $0.58 per task undercuts everything at its level, a direct shot at the economics that made DeepSeek and GLM-Flash the default budget picks. Second, Google has adopted the Chinese release cadence — four Flash models in four months is DeepSeek-style iteration, right down to the "introductory" pricing that everyone assumes will be superseded by a newer model before it expires. The honest caveats cut the other way: absolute token prices still favor the Chinese tier (V4-Flash remains ~41% cheaper on input, ~6x on output — and its price doesn't double in January), the 40% hotter token burn means the "cheap" model can bill like a mid-tier one on agentic workloads, and there are still no weights to self-host. For builders the decision tree is unchanged in shape, just tighter at the margins: meter tokens → China Flash tier; meter completed tasks and need frontier-grade quality → Gemini 3.8 Flash is now a real contender, at least until New Year's Day repricing.
Primary sources
- The Verge — launch coverage, "works harder" token warning
- Artificial Analysis — $0.58/task, cheapest at its intelligence level; 300 tok/s; per-task cost +40% vs 3.7
- The Decoder — DeepSWE 73.7% breakdown, 2027 pricing, Cyber variant numbers
- Ars Technica — cadence analysis, DeepSWE leaderboard read
- The Register — intelligence roundup incl. GLM-5.3 60 / Kimi K3 60, cost-per-task comparison
- 站内基线:Gemini 3.7 Flash 发布文 · DeepSeek 谷时定价 · GLM-5.3-Flash 揭晓
FAQ (2026)
What launched?
Gemini 3.8 Flash, Sep 2 — third Flash in six weeks. 1M context, multimodal in, ~300 tok/s, agent/SWE positioning. A Cyber variant (86.2% CyberGym) is Fairwind-restricted.
Pricing?
$0.75/$3.75 per 1M intro through Dec 31 (cache −90%), then $1.50/$7.50 from Jan 1, 2027. Per-task cost runs ~40% above 3.7 Flash — more tokens, more turns.
Benchmarks?
DeepSWE v1.1 73.7% — No.2 behind Opus 5 (74.0%), past GPT-5.6 Sol (72.7%) and 3.7 Flash (65.3%). Led Vals Finance Agent V2 and Harvey Legal Agent.
vs Chinese Flash tier?
DeepSeek V4-Flash peak $0.44/$1.32 (off-peak half) stays cheaper per token and doesn't double in 2027. AA: Gemini is cheapest per intelligence unit ($0.58/task). GLM-5.3/K3 match it on intelligence (60) and are open-weight.
GLM-5.3 / Kimi K3 comparison?
The Register's nine-benchmark roundup: GLM-5.3 (max) 60, Kimi K3 (max) 60 — in the same frontier pack, self-hostable, with GLM-5.3-Flash at ~1/20 flagship pricing for the budget tier.
Switch or stay?
Token-metered budgets: China Flash tier still wins, esp. off-peak. Task-metered: Gemini 3.8 Flash is the new per-task price-performance leader — until Jan 1 repricing. Google itself: stay on 3.7 Flash to minimize tokens.