Skip to content

Cerebras Inference Release notes from Cerebras

45 releases · 3 this month

Official notes RSS
DateRelease
Sep 22Temporary increase to Qwen 3.8 27B total token rate limit The Developer tier total token rate limit for...unverified
Update
Sep 3Gemma 4 31B availability changes on public endpoints Starting September 3, 2026, gemma-4-31b is no longer...unverified
Update
Sep 3Qwen 3.8 27B available on public endpoints qwen-3.8-27b is now available on Cerebras public endpointsunverified
Update
Aug 28Flowise integration moved to legacy status Flowise has archived its repositories and ends official support on...unverified
Update
Jul 27GLM 4.7 reasoning_logprobs default will not ship before deprecation zai-glm-4.7 is scheduled for deprecation...unverified
Deprecation
Jul 26Image dimension limit for image inputs Image inputs now enforce a maximum width and height of 15,000 pixels...unverified
Update
Jul 24Predicted Outputs limited to dedicated endpoints Predicted Outputs is available by request on dedicated...unverified
Preview
Jul 23Image limit increased for Developer and Enterprise accounts Developer and Enterprise accounts using...unverified
Update
Jul 22API Version 2 Default Rollout As of July 22, 2026, API version 2 is the defaultunverified
Update
Jul 16Free Trial credits for new accounts New accounts now receive \$5 in free credits after adding a verified...unverified
Pricing
Jul 16Uncached Rate Limits Cerebras now uses a dual-bucket rate limiting modelunverified
Update
Jul 14Gemma 4 31B available across coding integrations gemma-4-31b is now a selectable model in our Cline, Kilo...unverified
Update
Jun 29Gemma 4 31B and Image Input Support - Gemma 4 31B (gemma-4-31b) is now available in previewunverified
Preview
Jun 1New dedicated models: StepFun Step 3.5 Flash and Step 3.7 Flash StepFun's Step 3.5 Flash and Step 3.7 Flash...unverified
Update
May 1Projects is now generally available Projects is out of private preview and available to all organizationsunverified
GA
Apr 27New dedicated models: GLM 5, GLM 5.1, and Kimi K2.6 Z.AI's GLM 5 and GLM 5.1 and Moonshot AI's Kimi K2.6 are...unverified
Update
Apr 24Validation Errors Now Return 400 Instead of 422 API requests that fail validation now return HTTP 400 Bad...unverified
Update
Apr 22New parameter: prompt_cache_key /v1/chat/completions and /v1/completions now accept an optional...unverified
Update
Apr 10Payload Optimization: msgpack and gzip /v1/chat/completions and /v1/completions now accept request bodies...unverified
Update
Mar 31New Sampling Parameters The following parameters are now available on all models: - frequency_penalty –...unverified
Update
Mar 30Streaming: Server-Sent Events (SSE) Streaming responses for Chat Completions (/v1/chat/completions) and...unverified
Update
Mar 27Rolling out selective 4-bit weight-only quantization for supported models, targeting non-sensitive layers...unverified
Update
Mar 26developer Message Role (gpt-oss-120b) The Chat Completions API now supports the developer message role for...unverified
Update
Mar 26Temperature Parameter: Increased Maximum The maximum allowed value for the temperature parameter for Chat...unverified
Update
Mar 25Logprob Consistency Logprobs are now computed from the model's raw output, before applying temperature scalingunverified
Update
Mar 24Deprecating disable_reasoning parameter for zai-glm-4.7 The disable_reasoning parameter is deprecated and...unverified
Deprecation
Feb 27reasoning_effort="none" support for GLM 4.7 zai-glm-4.7 now accepts the reasoning_effort parameterunverified
Update
Jan 22Monitor your dedicated inference endpoints with the new Metrics API, providing Prometheus-compatible metrics...unverified
Update
Jan 21API Version 2 Available for Testing API version 2 is now available for testing via the...unverified
Update
Jan 14Prioritize requests to the API to balance latency sensitivity and resource allocation across your workloads...unverified
Update
Jan 9Constrained Decoding Generally Available Constrained decoding has moved to General Availability (GA) with...unverified
GA
Jan 6Added preview support for Z.ai GLM 4.7: zai-glm-4.7unverified
Preview
Dec 18Process large-scale inference workloads asynchronously with the new Batch API and Files APIunverified
Update
Dec 17Added support for parallel tool calling, allowing models to request multiple tool invocations simultaneously...unverified
Update
Dec 16Streaming Support Streaming is now supported for all models, including with structured output requests...unverified
Update
Dec 10Added support for prompt caching, which automatically stores and reuses prompt prefixes to reduce latency and...unverified
Update
Nov 24We now support Predicted Outputs for faster generation when large portions of the output are known in advanceunverified
Update
Nov 14Deprecated qwen-3-235b-a22b-thinking-2507 We recommend migrating to GPT OSS 120Bunverified
Deprecation
Nov 7Added preview support for Z.ai GLM 4.6: zai-glm-4.6unverified
Preview
Nov 5Deprecated qwen-3-coder-480bunverified
Deprecation
Nov 3Deprecated llama-4-scout-17b-16e-instruct See Deprecations for more detailsunverified
Deprecation
Oct 22Added support for retrieving log probabilities via the logprobs parameter for the Completions endpointunverified
Update
Oct 15Deprecated llama-4-maverick-17b-128e-instruct See Deprecations for more detailsunverified
Deprecation
Oct 6We're gradually introducing multi-token streaming across our modelsunverified
Update
Oct 2Updated support for OpenAI GPT-OSS (gpt-oss-120b) - Tool calling with strict: true (constrained decoding) -...unverified
Update
About and activity

Fast inference on wafer-scale chips.

Releases per month
4Apr
1May
2Jun
8Jul
1Aug
3Sep

Sources: each vendor's own release notes, changelogs and GitHub releases, read daily to monthly by how often it posts. Logos via logo.dev; trademarks belong to their owners.

New releases by email

Tuesday mornings: the week's data, AI and developer-tools releases, only in weeks when something shipped.

Double opt-in. Unsubscribe any time.