Cerebras

Executive Summary

What it is: Cerebras is an inference-only provider running open-weight models on wafer-scale chips (~3,000 tok/s on GPT OSS 120B), with a $5-credit free tier, pay-as-you-go Developer access, and enterprise dedicated endpoints; the Cerebras Code subscription ($50/$200) has been unavailable for four months. Source: https://www.cerebras.ai/pricing

What to watch out for: The shared API shrank to two models (GPT OSS 120B, Gemma 4 31B) after GLM 4.7's scheduled Aug 17 deprecation, cached tokens still bill at full input rates, and there is no timeline for subscription reopenings. Source: https://inference-docs.cerebras.ai/models/overview

Bottom line: Tracking remains paused under HLD rules and August confirmed the pause: shrinking shared catalog, no subscriptions, and only infrastructure news (a 165 MW Finland data centre); revisit when a coding subscription reopens. Source: https://investors.cerebras.ai/news-releases/news-release-details/cerebras-and-compute-nordic-finland-announce-new-165-mw-ai-data

Key Terms

  • Wafer-Scale Engine (WSE-3) - Cerebras's chip architecture delivering up to ~3,000 tokens/sec on GPT OSS 120B, the core differentiator versus GPU inference. Source: Inference-Docs – Overview
  • Dedicated Endpoints - enterprise reserved capacity with 30+ additional model families, production SLAs, custom weights, and fine-tuning; the only way to run more than the two shared-API models. Source: Inference-Docs – Overview
  • REAP - Router-weighted Expert Activation Pruning, Cerebras's MoE compression research; pruned models are published on HuggingFace for research only and explicitly not served on the public API. Source: Inference-Docs – Overview

Latest Changes

Changes since the 2026-07 report.

July watch items verified:

  • GLM 4.7 deprecation (Aug 17): confirmed. The shared model catalog now lists exactly two models: GPT OSS 120B and Gemma 4 31B. No replacement was added. Source: Inference-Docs – Overview
  • Cerebras Code subscription reopening: failed (4th consecutive month). The pricing page no longer lists the Cerebras Code Pro/Max subscription product at all; only Inference API tiers (Free Trial $5 credits, Developer, Enterprise) are shown. Source: Cerebras – Pricing
  • AMD partnership availability (H2 2026): still-pending. No update found on the July-announced AMD Helios + WSE disaggregated inference pairing. Source: Cerebras – Partners

New August changes:

Tracking status: HLD paused Cerebras tracking as of July. August confirms the pause conditions persist: the shared catalog shrank to two models and the coding subscription has been unavailable for four consecutive months. Recommend keeping tracking paused; revisit when subscriptions reopen or a new shared coding-class model is added.

Plans

Tier Price Notes
Free Trial $5 free credits All shared models, Discord support
Developer Pay-as-you-go from $10 top-up 10x free-tier rate limits, higher priority
Enterprise Custom Dedicated endpoints, custom weights, fine-tuning, SLAs

Cerebras Code Pro ($50/mo) and Max ($200/mo): still not offered. Source: Cerebras – Pricing

API Pricing

Shared-endpoint per-token rates: GPT OSS 120B and Gemma 4 31B at their published rates (unchanged; see developer-tier pricing table). Cached input tokens remain billed at the same rate as fresh input (no cache discount), with a dual-bucket rate-limit system giving 3x TPM headroom for cached tokens. Source: Cerebras – Pricing , Inference-Docs – Pricing

Model Performance / Benchmarks

No new benchmarks published. Standing speeds: ~3,000 tok/s (GPT OSS 120B), ~1,850 tok/s (Gemma 4 31B). Source: Inference-Docs – Overview

Latest News

The Finland data centre (165 MW, Mikkeli) is the month's only substantive announcement, extending capacity for the inference business. Source: Investors – Cerebras And Compute Nordic Finland Announce New 165 Mw Ai Data

Community Signals

No significant HN/Reddit threads on Cerebras found in August. Community attention to Cerebras-hosted models flows through third-party references (e.g., SWE-1.7's 1000 TPS serving claim on Cerebras hardware from July's Cognition launch).

Enterprise Readiness

Feature Available? Details
SSO (SAML) No Not documented on self-serve tiers.
SCIM No Not documented.
Audit logs No Not documented.
IP indemnity No Not documented.
Data residency Partial Finland data centre adds EU capacity; residency commitments not published yet. Source: Investors
HIPAA No Not advertised.
Air-gapped / on-prem No Dedicated endpoints are cloud-reserved capacity.
SLA Yes Enterprise dedicated endpoints with "guaranteed uptime" and response-time guarantees. Source: Cerebras – Pricing
Admin controls (RBAC) No Not documented.

Transparency Gaps

Gap Details Severity
Subscription roadmap Four months with no Cerebras Code availability and no reopening timeline. High
Shared catalog shrinkage Two models remain; whether new coding-class models will be added to shared endpoints is undisclosed. Medium
Cache pricing Cached tokens still billed at full input rate. Medium