Skip to content

Cerebras Release notes and new features

Public · Sunnyvale, US · 45 releases · 3 this month

Website RSS

Also onTechConf(1 conference)

Cerebras releases

September 2026 digest
DateRelease
Sep 22Temporary increase to Qwen 3.8 27B total token rate limit The Developer tier total token rate limit for...unverified
Sep 3Gemma 4 31B availability changes on public endpoints Starting September 3, 2026, gemma-4-31b is no longer...unverified
Sep 3Qwen 3.8 27B available on public endpoints qwen-3.8-27b is now available on Cerebras public endpointsunverified
Aug 28Flowise integration moved to legacy status Flowise has archived its repositories and ends official support on...unverified
Jul 27GLM 4.7 reasoning_logprobs default will not ship before deprecation zai-glm-4.7 is scheduled for deprecation...unverified
Jul 26Image dimension limit for image inputs Image inputs now enforce a maximum width and height of 15,000 pixels...unverified
Jul 24Predicted Outputs limited to dedicated endpoints Predicted Outputs is available by request on dedicated...unverified
Jul 23Image limit increased for Developer and Enterprise accounts Developer and Enterprise accounts using...unverified
Jul 22API Version 2 Default Rollout As of July 22, 2026, API version 2 is the defaultunverified
Jul 16Free Trial credits for new accounts New accounts now receive \$5 in free credits after adding a verified...unverified
Jul 16Uncached Rate Limits Cerebras now uses a dual-bucket rate limiting modelunverified
Jul 14Gemma 4 31B available across coding integrations gemma-4-31b is now a selectable model in our Cline, Kilo...unverified
Jun 29Gemma 4 31B and Image Input Support - Gemma 4 31B (gemma-4-31b) is now available in previewunverified
Jun 1New dedicated models: StepFun Step 3.5 Flash and Step 3.7 Flash StepFun's Step 3.5 Flash and Step 3.7 Flash...unverified
May 1Projects is now generally available Projects is out of private preview and available to all organizationsunverified
Apr 27New dedicated models: GLM 5, GLM 5.1, and Kimi K2.6 Z.AI's GLM 5 and GLM 5.1 and Moonshot AI's Kimi K2.6 are...unverified
Apr 24Validation Errors Now Return 400 Instead of 422 API requests that fail validation now return HTTP 400 Bad...unverified
Apr 22New parameter: prompt_cache_key /v1/chat/completions and /v1/completions now accept an optional...unverified
Apr 10Payload Optimization: msgpack and gzip /v1/chat/completions and /v1/completions now accept request bodies...unverified
Mar 31New Sampling Parameters The following parameters are now available on all models: - frequency_penalty –...unverified
Mar 30Streaming: Server-Sent Events (SSE) Streaming responses for Chat Completions (/v1/chat/completions) and...unverified
Mar 27Rolling out selective 4-bit weight-only quantization for supported models, targeting non-sensitive layers...unverified
Mar 26developer Message Role (gpt-oss-120b) The Chat Completions API now supports the developer message role for...unverified
Mar 26Temperature Parameter: Increased Maximum The maximum allowed value for the temperature parameter for Chat...unverified
Mar 25Logprob Consistency Logprobs are now computed from the model's raw output, before applying temperature scalingunverified
Mar 24Deprecating disable_reasoning parameter for zai-glm-4.7 The disable_reasoning parameter is deprecated and...unverified
Feb 27reasoning_effort="none" support for GLM 4.7 zai-glm-4.7 now accepts the reasoning_effort parameterunverified
Jan 22Monitor your dedicated inference endpoints with the new Metrics API, providing Prometheus-compatible metrics...unverified
Jan 21API Version 2 Available for Testing API version 2 is now available for testing via the...unverified
Jan 14Prioritize requests to the API to balance latency sensitivity and resource allocation across your workloads...unverified
Jan 9Constrained Decoding Generally Available Constrained decoding has moved to General Availability (GA) with...unverified
Jan 6Added preview support for Z.ai GLM 4.7: zai-glm-4.7unverified
Dec 18Process large-scale inference workloads asynchronously with the new Batch API and Files APIunverified
Dec 17Added support for parallel tool calling, allowing models to request multiple tool invocations simultaneously...unverified
Dec 16Streaming Support Streaming is now supported for all models, including with structured output requests...unverified
Dec 10Added support for prompt caching, which automatically stores and reuses prompt prefixes to reduce latency and...unverified
Nov 24We now support Predicted Outputs for faster generation when large portions of the output are known in advanceunverified
Nov 14Deprecated qwen-3-235b-a22b-thinking-2507 We recommend migrating to GPT OSS 120Bunverified
Nov 7Added preview support for Z.ai GLM 4.6: zai-glm-4.6unverified
Nov 5Deprecated qwen-3-coder-480bunverified
Nov 3Deprecated llama-4-scout-17b-16e-instruct See Deprecations for more detailsunverified
Oct 22Added support for retrieving log probabilities via the logprobs parameter for the Completions endpointunverified
Oct 15Deprecated llama-4-maverick-17b-128e-instruct See Deprecations for more detailsunverified
Oct 6We're gradually introducing multi-token streaming across our modelsunverified
Oct 2Updated support for OpenAI GPT-OSS (gpt-oss-120b) - Tool calling with strict: true (constrained decoding) -...unverified
About and activity

Cerebras Systems Inc., headquartered in Sunnyvale, California, develops semiconductors, supercomputers, and related software to power artificial intelligence deep-learning applications such as inference engines.

Releases per month

Sources: each vendor's own release notes, changelogs and GitHub releases, read daily to monthly by how often it posts. Logos via logo.dev; trademarks belong to their owners.

New releases by email

Tuesday mornings: the week's data, AI and developer-tools releases, only in weeks when something shipped.

Double opt-in. Unsubscribe any time.