News / Model Watch · DeepSeek ecosystem pricing
Ollama Cuts DeepSeek-V4 Prices 50% Off-Peak: The Convention DeepSeek Invented Is Going Industry-Standard
Ollama — the platform that put local LLMs on every laptop, now a hosted API for frontier open weights — introduced off-peak token rates for DeepSeek-V4: peak pricing applies 12:00-18:00 UTC Monday to Friday; everything else, including the entire weekend, is half price. deepseek-v4-flash drops to $0.22/$0.66 per 1M off-peak — and the peak rates are digit-for-digit DeepSeek's own card. The peak/off-peak convention DeepSeek launched on August 16 just became a western-industry standard.
Ollama now charges 50% less for DeepSeek-V4 outside 12:00-18:00 UTC on weekdays, and all weekend, with more models getting off-peak rates soon. Rates per the official pricing page: v4-flash $0.44/$1.32 peak → $0.22/$0.66 off-peak (cached $0.014 → $0.007); v4-pro $1.32/$3.96 → $0.66/$1.98. DeepSeek models are hosted in the US and Europe with zero data retention — a compliance-friendly alternative for workloads that can't touch China-hosted endpoints. GLM-5.3 rides along at z.ai parity ($1.40/$4.40).
What the official sources confirm
- The announcement: Ollama's official X account and LinkedIn announced "off-peak token rates — DeepSeek V4" with 50% savings; the HN discussion linked the pricing page directly.
- The window: peak = 12:00-18:00 UTC, Mon-Fri. Off-peak = weekday early mornings/evenings + the entire weekend.
- The rates: v4-flash peak $0.44 input / $1.32 output / $0.014 cached — identical to DeepSeek's own peak card — halving to $0.22/$0.66/$0.007 off-peak. v4-pro: $1.32/$3.96/$0.044 peak → $0.66/$1.98/$0.022 off-peak.
- The compliance angle: US & Europe hosting with ZDR per Ollama's own post.
- The shelf: the same page lists GLM-5.3 at exactly z.ai's rates, plus GLM-5.3-Flash, Qwen3.5-397B, Gemma4, Nemotron-3-Ultra — western-hosted open-weight frontier under one API.
- The roadmap: "off-peak pricing will be available soon for more models."
Why this matters
- The convention is escaping its origin. When DeepSeek introduced peak/off-peak + weekend billing in August, it read as a capacity-management quirk. A major western provider adopting the same clock turns it into a de-facto pricing standard for frontier open weights.
- Batch economics just doubled down. Scheduling non-interactive workloads into evenings/weekends now halves costs on both DeepSeek's own API and Ollama's US/EU cloud — same clock, same discount structure.
- ZDR hosting changes the buyer set. Regulated teams that couldn't use China-hosted endpoints can now run V4 at (peak-)parity rates — expanding DeepSeek's addressable market without DeepSeek operating it.
- For relay users: complements, not substitutes. Off-peak scheduling optimizes unit cost for batch; an OpenAI-compatible relay optimizes routing breadth (Qwen/GLM/Kimi/video models, one key, global billing) for production traffic. The cheapest stack uses both.
Primary sources
FAQ (2026)
What launched?
Ollama off-peak rates: 50% off outside 12:00-18:00 UTC weekdays + all weekend, for DeepSeek-V4 (more models soon).
V4-Flash rates?
Peak $0.44/$1.32 (cached $0.014) — same as DeepSeek's own card; off-peak $0.22/$0.66 (cached $0.007).
vs DeepSeek's API?
Same convention (since Aug 16) and same flash peak numbers; Ollama sets its own pro rates and hosts in US/EU with ZDR.
GLM on Ollama?
Yes — GLM-5.3 at z.ai parity ($1.40/$4.40), plus GLM-5.3-Flash, Qwen3.5-397B, Gemma4, Nemotron-3-Ultra.
Compliance angle?
US/Europe hosting with zero data retention — usable where China-hosted endpoints are off-limits.
Switch production to it?
Batch jobs: schedule off-peak on either provider. Interactive multi-model traffic: a relay stays the simpler spine. Cheapest stacks use both.