ChinaModelAPI

News / Model Watch · Qwen ecosystem

Ecosystem launch ~1,500 tok/s $0.99 / $1.49 per 1M Apache 2.0 · 262K ctx
2026-09-04 · Model Watch · Cerebras inference docs + official X post, Sep 3

Qwen3.8-27B Hits ~1,500 Tokens/s on Cerebras: The Open-Weights Speed Play, Priced at $0.99/$1.49

Alibaba's Apache-2.0 Qwen3.8-27B — the laptop-class daily driver of the Qwen3.8 generation — now runs on Cerebras' wafer-scale inference at roughly 1,500 output tokens per second, listed on the official Cerebras inference docs and pricing page since September 3. The Hacker News thread hit 392 points overnight. Here's what shipped, what it costs, and where it fits for builders routing the Qwen family.

Direct answer

Cerebras now hosts Qwen3.8-27B at ~1,500 tokens/s output (expectations cited above 2,000), at $0.99 input / $1.49 output per 1M tokens, as a preview model for evaluation. The model itself is Alibaba's 27B dense, natively multimodal, Apache-2.0 open-weights release with 262K context — Cerebras' own launch post credits it with beating Qwen3.7-Plus overall and excelling at real-world coding. For speed-critical workloads on one open model, this is the fastest shelf in town; for production routing across the whole Qwen family (Max-0902, Flash tiers, vision), an OpenAI-compatible relay remains the practical spine.

What's confirmed from official sources

  • Hosting: Qwen3.8-27B listed on Cerebras' official model overview and pricing pages since Sep 3; preview status, intended for evaluation.
  • Speed: ~1,500 output tokens/s quoted at launch, with expectations of exceeding 2,000 — territory that would set a new serving-speed record for this class.
  • Pricing: $0.99 per 1M input, $1.49 per 1M output.
  • Model credentials: Cerebras' official post congratulates the Qwen team and describes 27B as a native multimodal dense model that outperforms Qwen3.7-Plus overall, strong on real-world coding — consistent with the Apache-2.0 release (Aug 14) carrying 262K native context.
  • Reception: HN front-page thread at 392 points / 120+ comments within hours — the biggest China-model infra story this week alongside Max-0902.

What 1,500 tok/s actually changes

  • Interaction model flips. At GPU-typical speeds, a 4,000-token answer streams for a minute-plus; at 1,500 tok/s it lands in under three seconds. Agentic loops that re-generate whole files stop punishing users with wait states.
  • 262K context + speed = live document Q&A. The pairing that made Cerebras' earlier Qwen3-235B launch notable (1-2 minutes → 0.6 seconds) now arrives on a model anyone can also self-host — the weights are Apache 2.0.
  • The economics are sane, not loss-leader crazy. $0.99/$1.49 sits below Qwen's own flagship API rates ($2/$6 on Max-0902) while delivering single-model speed GPU pools can't match — this is the wafer-scale bet paying out on a free model someone else trained.
  • Preview caveat: listed for evaluation; rate limits, availability and pricing can shift once it graduates. Treat as the fast lane, not the production spine, until then.

Primary sources

FAQ (2026)

What launched?

Cerebras hosting of Qwen3.8-27B since Sep 3 — ~1,500 tok/s output, preview tier, on official docs and pricing pages.

Pricing?

$0.99/$1.49 per 1M in/out — below Qwen flagship API rates, for the fastest single-model serving available.

Which model is it?

The Aug-14 Apache-2.0 release: 27B dense, natively multimodal, 262K context — beats Qwen3.7-Plus per Cerebras' launch post.

Why does speed matter?

4,000-token answers land in ~3s instead of a minute — flips agentic regeneration and big-context Q&A into interactive flows.

vs a Qwen relay?

Complementary: Cerebras = one open model at max speed; a relay = whole-family routing (Max-0902, Flash, vision) through one OpenAI-compatible endpoint.

Production-ready?

Preview status, listed for evaluation — expect limits and possible pricing shifts at graduation. Fast lane, not spine.

Related guides