News / Model Watch · Western frontier
Gemini 3.7 Flash: Google's Workhorse Model for Coding and Agents
Google announced Gemini 3.7 Flash on August 13, 2026, calling it its "most intelligent workhorse model yet" for coding, software engineering, and agent workflows. We track Gemini as a Western frontier baseline against our Chinese open models (the same role GPT-5.6 Sol, Claude Opus 5, and Grok 4.6 play) — so here is a clean, sourced read on what Google actually shipped, and what it means for a China-model stack.
Gemini 3.7 Flash shipped August 13, 2026 — three weeks after Gemini 3.6 Flash.
Per Google's announcement and the official Gemini API docs, gemini-3.7-flash is a stable model with a
1,048,576-token input limit (~1M context) and 65,536-token output limit,
thinking levels low/medium/high, and agent features including function calling, code execution, search grounding, and computer use (preview).
Introductory pricing is $0.75 / $3.75 per 1M tokens (half of 3.6 Flash's launch price) through December 31, 2026.
Note: Gemini is a Western closed model, not routed by ChinaModelAPI — we cover it as a comparison baseline only.
What Google's announcement and docs say (verified)
We use the official announcement and API documentation as the source of truth. As of 2026-08-19:
| Spec | Gemini 3.7 Flash (per Google) |
|---|---|
| Model ID | gemini-3.7-flash · listed as stable in the Gemini API docs |
| Announced | August 13, 2026 (Google blog, Tulsee Doshi, Senior Director, Product Management) |
| Input token limit | 1,048,576 (~1M context) |
| Output token limit | 65,536 |
| Inputs → output | Text, image, video, audio, PDF → text |
| Thinking | low / medium / high (“minimal” returns an error) |
| Agent features | Function calling, code execution, search & Maps grounding, file search, URL context, structured outputs, caching, Batch API, Flex/Priority; computer use in preview. Audio/image generation and Live API not supported. |
| Where | Gemini API (AI Studio, Android Studio), Google Antigravity, Gemini Enterprise Agent Platform (Model Garden) & Enterprise app; Spark for AI Pro/Ultra subscribers (160+ countries; not EEA, Nigeria, Switzerland, UK) |
| Safety | Updated safeguards against CBRN and cyber-offense misuse |
Benchmarks vs Gemini 3.6 Flash
Google's announcement reports across-the-board gains over 3.6 Flash on coding, agent, and document-reasoning benchmarks:
| Benchmark | 3.7 Flash | 3.6 Flash |
|---|---|---|
| FrontierCode 1.1 Main | 43.6% | 34.4% |
| DeepSWE v1.1 | 65.3% | 49.0% |
| WebDev Arena (Arena.ai) Elo | 1588 | 1538 |
| GDP.pdf | 34.0% | 22.0% |
| AutomationBench (Zapier) | 30.4% | 17.0% |
Vendor-reported numbers from Google's announcement — not independently verified by us. Google also cites better debugging, higher first-pass code accuracy, improved multi-step planning and tool calls, and stronger design adherence in UI generation.
Pricing and a three-week release cadence
Gemini 3.6 Flash landed July 21, 2026; 3.7 Flash followed on August 13 — a cadence English media flagged prominently (Ars Technica: "just three weeks after previous release"). Google attributes the speed to developer feedback and algorithmic innovations it plans to reuse in future models. On price:
- Through Dec 31, 2026: $0.75 per 1M input tokens · $3.75 per 1M output tokens — half the per-token cost of 3.6 Flash's launch pricing.
- From Jan 1, 2027: standard pricing of $1.50 input / $7.50 output per 1M tokens.
- Framing: VentureBeat summarized the launch as Google "targeting coding and agents with a 50% introductory price cut."
How this fits a China-model builder
Gemini 3.7 Flash is not a model ChinaModelAPI routes — it is a Western closed model. We track it as a quality and cost ceiling reference, exactly like GPT-5.6 Sol, Claude Opus 5, and Grok 4.6 in our model-watch tables. The notable shift: its ~1M context and aggressive intro pricing close two gaps Chinese flagships used to hold clearly.
| Model | Origin / access | Best role in your stack |
|---|---|---|
| Gemini 3.7 Flash | Google · closed API | Western workhorse baseline (coding / agents) |
| DeepSeek-V4-Pro-0813 | DeepSeek · open weights | Frontier agents/reasoning, self-hostable, cheap |
| Qwen3.8-Max | Alibaba · open weights | Coding & long-horizon work |
| GLM-5.3 / Kimi K3 | Zhipu / Moonshot · open | 1M-context long-horizon & agents |
Full live comparison: Chinese AI model comparison.
Primary sources
- Google Blog — Gemini 3.7 Flash: our most intelligent workhorse model (official announcement, Aug 13, 2026)
- Gemini API docs — gemini-3.7-flash model page (1,048,576 input / 65,536 output, thinking levels, feature support)
- Gemini API docs — models catalog (gemini-3.7-flash listed as stable)
- 9to5Google — Gemini 3.7 Flash launch (Aug 13, 2026)
- VentureBeat — Google's Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
- Ars Technica — Google announces Gemini 3.7 Flash just three weeks after previous release
FAQ
Can I call Gemini 3.7 Flash via ChinaModelAPI?
No. Gemini is a Western closed model. ChinaModelAPI routes Chinese frontier models (DeepSeek, Qwen, GLM, Kimi, Seedance, HappyHorse). We list Gemini 3.7 Flash here only as a Western baseline for comparison.
1M context — parity with Chinese flagships?
On paper, yes: 1,048,576 input tokens matches the 1M class that DeepSeek-V4-Pro, Qwen3.8, GLM-5.2/5.3, and Kimi K3 advertise. Open-weight Chinese models remain self-hostable and typically cheaper per token.
Is the $0.75/$3.75 price permanent?
No — it is introductory pricing through December 31, 2026. From January 1, 2027, standard pricing is $1.50 input / $7.50 output per 1M tokens, per Google's announcement.
Are the benchmark numbers independent?
No. The 43.6% FrontierCode / 65.3% DeepSWE / 1588 WebDev Arena Elo figures are vendor-reported by Google comparing 3.7 Flash to 3.6 Flash. Treat them as directional until third-party evals publish.