News / Model Watch · Beta Watch
DeepSeek Quietly Opens a 48-Hour V4.1 Flash Beta: New Architecture, Native Multimodal — and an Official Survey Asking If It Can Replace V4-Pro
No launch event, no teaser art. On Sep 8 a notice appeared in DeepSeek's official developer group: V4.1 Flash, a middle-checkpoint build, is live for testing — set model='deepseek-v4.1-flash-expires-on-0910' and call it. The id is the deadline: the endpoint retires September 10. What's inside, per the official notice: a new model architecture and native multimodal support — with billing unchanged from V4-Flash, a 20-concurrency cap, and an attached survey that literally asks whether this build could replace V4-Pro.
Update (Sep 10 — final beta day): the story went mainstream overnight — the Hacker News thread "DeepSeek launching v4.1 flash cheaper and more capable than v4 pro" reached 391 points / 200+ comments, and Vercel's AI Gateway now lists deepseek-v4.1-flash-beta for third-party routing. The beta window closes today (09-10); official rate-card and changelog entries had not appeared as of 07:00 UTC+8.
DeepSeek V4.1 Flash beta is live (Sep 8-10): keep base_url, switch model to deepseek-v4.1-flash-expires-on-0910. Official claims: new architecture + native multimodal (text/image/audio in one base — no more Vision bolt-on) at V4-Flash pricing. Community timing: 300-507 tok/s. Limits: 20 concurrency/account (vs 2,500 production). The official survey explicitly tests "can it replace V4-Pro" — a positioning question, not a published result. Changelog and HF still show nothing: gray-stage, exactly DeepSeek's style. If the checkpoint validates, expect a formal release with MIT weights like the rest of the line.
The facts, as verified (Sep 9)
- How to call it. base_url unchanged; model id
deepseek-v4.1-flash-expires-on-0910. The suffix is literal — the endpoint expires Sep 10 and DeepSeek labels the build a "middle checkpoint" (中间版本), i.e. real traffic testing before a formal cut. - Billing and limits. Same rates as deepseek-v4-flash (RMB: off-peak ¥0.05 cached / ¥1.5 input / ¥4.5 output per 1M; peak ¥0.1 / ¥3 / ¥9). Concurrency capped at 20/account — against 2,500 on V4-Flash production. This is a validation sandbox, not an upgrade path.
- The architecture claim. Official notice says "new model structure" — architecture-level rebuild. No tech report, no parameter count, no benchmark. Chinese coverage (36Kr/Tencent Tech, ifeng, The Paper/Eastmoney, Tide News) is unanimous on the two headline claims and equally empty on specifics.
- Native multimodal, finally. The V4 line handled vision through the separate Vision-Exp model since Aug 21. V4.1 Flash integrates text, image and audio in the base model itself — same-direction as the whole industry's 2026 convergence, but notable because DeepSeek was the last major holdout shipping vision as a side model.
- Speed, from the field. Developer posts on X report 300+ tok/s sustained with peaks near 507 — on par with or above the Gemini 3.8 Flash-class speed tier, at a fraction of the price.
- The survey. The beta ships with an anonymous feedback form whose modules include an explicit "feasibility of replacing V4-Pro" assessment. V4-Pro launched less than a month ago; DeepSeek is already testing whether its own flagship's workloads can slide down to the cheap tier. That's the real announcement inside the announcement.
Per multiple outlets: as of Sep 9 the official changelog, API docs and Hugging Face list no V4.1 entry — as-of statements only; treat all capability claims as beta-stage vendor statements.
Context: five moves in forty days
The cadence is the second story. Jul 31: V4-Flash-0731 API. Aug 13: V4-Pro-0813 plus the open-source DeepSeek Harness. Aug 21: Vision-Exp on the API. Aug 31: MIT open weights. Sep 8: V4.1 Flash beta. Five releases in ~40 days, each one either expanding capability or opening supply. Set against this week's Western news — OpenAI declaring an "AGI era" behind a $50/1M output wall and Anthropic welding anti-distillation locks into its API — the Chinese frontier is running a different play entirely: ship fast, price flat, publish the weights. If V4.1 Flash's final cut really lands "near-Pro quality at Flash pricing with native multimodal," the budget-tier repricing we profiled in the Gemini 3.8 Flash analysis gets another leg down — from the side that isn't raising prices in January.
Primary sources
- DeepSeek 官方交流群通知(Sep 8)——模型 ID、计费、并发、问卷的一手来源(群通知截图经多媒体转述一致)
- Hacker News — DeepSeek launching v4.1 flash cheaper and more capable than v4 pro(391 分 / 202 评论,09-09 夜间)
- Vercel AI Gateway — deepseek-v4.1-flash-beta 模型页(第三方网关已上架内测模型)
- 36氪/腾讯科技 — V4.1 Flash 内测详情(含 changelog 未更新的灰度核实)
- 凤凰网科技 — 官方通知要点转述
- 澎湃新闻/东方财富 — 速度反馈与产品定位
- BigGo — 507 tok/s 实测、人民币计价明细、问卷「V4-Pro 替代可行性」模块
- 站内对照:Vision-Exp API 文 · MIT 权重开源文 · 峰谷定价文
FAQ (2026)
How do I try it?
base_url unchanged, model='deepseek-v4.1-flash-expires-on-0910'. Billing = V4-Flash. But the endpoint dies Sep 10 — this is a test id, don't build on it.
What's actually new?
Per official notice: new architecture + native multimodal (text/image/audio in one base). No tech report or benchmarks published — the "how" is still the open question.
Limits?
20 concurrency/account (vs 2,500 on production V4-Flash). It's a validation sandbox: functional checks and benchmarks yes, production traffic no.
Can it replace V4-Pro?
Unknown — but DeepSeek's own survey explicitly tests that question. If near-Pro quality lands at Flash pricing, Pro-tier traffic migrating down is the obvious commercial play.
Official release?
Not yet — changelog/docs/HF show no V4.1 entry as of Sep 9. Gray-stage checkpoint. Given the line's MIT pattern, a formal release with weights is the likely follow-up if it validates.
Market impact?
If final cut = near-Pro at Flash pricing with native multimodal: budget-tier repricing pressure on GLM-5.3-Flash, Qwen Flash, and Gemini 3.8 Flash's intro rate — from the side not raising prices in 2027.