Key Terms
- Wafer-Scale Engine (WSE-3) - Cerebras's chip architecture delivering up to ~3,000 tokens/sec on GPT OSS 120B, the core differentiator versus GPU inference. Source: Inference-Docs – Overview
- Dedicated Endpoints - enterprise reserved capacity with 30+ additional model families, production SLAs, custom weights, and fine-tuning; the only way to run more than the two shared-API models. Source: Inference-Docs – Overview
- REAP - Router-weighted Expert Activation Pruning, Cerebras's MoE compression research; pruned models are published on HuggingFace for research only and explicitly not served on the public API. Source: Inference-Docs – Overview
Latest Changes
Changes since the 2026-07 report.
July watch items verified:
- GLM 4.7 deprecation (Aug 17): confirmed. The shared model catalog now lists exactly two models: GPT OSS 120B and Gemma 4 31B. No replacement was added. Source: Inference-Docs – Overview
- Cerebras Code subscription reopening: failed (4th consecutive month). The pricing page no longer lists the Cerebras Code Pro/Max subscription product at all; only Inference API tiers (Free Trial $5 credits, Developer, Enterprise) are shown. Source: Cerebras – Pricing
- AMD partnership availability (H2 2026): still-pending. No update found on the July-announced AMD Helios + WSE disaggregated inference pairing. Source: Cerebras – Partners
New August changes:
- Infrastructure: Cerebras and Compute Nordic Finland announced a 165 MW AI data centre in Mikkeli, Finland (announced late August, site banner). Source: Investors – Cerebras And Compute Nordic Finland Announce New 165 Mw Ai Data
- Transparency publication: the model catalog now documents compression state per model (no pruned models on public endpoints; selective weight-only quantization in storage with full-precision activations and KV cache). Source: Inference-Docs – Overview
Tracking status: HLD paused Cerebras tracking as of July. August confirms the pause conditions persist: the shared catalog shrank to two models and the coding subscription has been unavailable for four consecutive months. Recommend keeping tracking paused; revisit when subscriptions reopen or a new shared coding-class model is added.
Plans
| Tier | Price | Notes |
|---|---|---|
| Free Trial | $5 free credits | All shared models, Discord support |
| Developer | Pay-as-you-go from $10 top-up | 10x free-tier rate limits, higher priority |
| Enterprise | Custom | Dedicated endpoints, custom weights, fine-tuning, SLAs |
Cerebras Code Pro ($50/mo) and Max ($200/mo): still not offered. Source: Cerebras – Pricing
API Pricing
Shared-endpoint per-token rates: GPT OSS 120B and Gemma 4 31B at their published rates (unchanged; see developer-tier pricing table). Cached input tokens remain billed at the same rate as fresh input (no cache discount), with a dual-bucket rate-limit system giving 3x TPM headroom for cached tokens. Source: Cerebras – Pricing , Inference-Docs – Pricing
Model Performance / Benchmarks
No new benchmarks published. Standing speeds: ~3,000 tok/s (GPT OSS 120B), ~1,850 tok/s (Gemma 4 31B). Source: Inference-Docs – Overview
Latest News
The Finland data centre (165 MW, Mikkeli) is the month's only substantive announcement, extending capacity for the inference business. Source: Investors – Cerebras And Compute Nordic Finland Announce New 165 Mw Ai Data
Community Signals
No significant HN/Reddit threads on Cerebras found in August. Community attention to Cerebras-hosted models flows through third-party references (e.g., SWE-1.7's 1000 TPS serving claim on Cerebras hardware from July's Cognition launch).
Enterprise Readiness
| Feature | Available? | Details |
|---|---|---|
| SSO (SAML) | No | Not documented on self-serve tiers. |
| SCIM | No | Not documented. |
| Audit logs | No | Not documented. |
| IP indemnity | No | Not documented. |
| Data residency | Partial | Finland data centre adds EU capacity; residency commitments not published yet. Source: Investors |
| HIPAA | No | Not advertised. |
| Air-gapped / on-prem | No | Dedicated endpoints are cloud-reserved capacity. |
| SLA | Yes | Enterprise dedicated endpoints with "guaranteed uptime" and response-time guarantees. Source: Cerebras – Pricing |
| Admin controls (RBAC) | No | Not documented. |
Transparency Gaps
| Gap | Details | Severity |
|---|---|---|
| Subscription roadmap | Four months with no Cerebras Code availability and no reopening timeline. | High |
| Shared catalog shrinkage | Two models remain; whether new coding-class models will be added to shared endpoints is undisclosed. | Medium |
| Cache pricing | Cached tokens still billed at full input rate. | Medium |