vLLM 0.31.0
Update5 days agoUnverifiedAdded Oct 10, 2026
v0.31.0 Highlights This release features 717 commits from 307 contributors (96 new)! - DeepSeek-V4.1-Flash performance: FlashMLA mega attention with the V4.1 NVFP4 compressed KV cache is now the SM100 default ; DeepGEMM sparse MQA logits for the indexer and Mega-Gate fusing the gate GEMM with expert selection ; decoder boundaries fuse the TP all-reduce, mHC...
Topics: Developer tools, Performance
More vLLM releases
Every vLLM releasevLLM 0.30.0
v0.30.0 Highlights This release features 762 commits from 315 contributors (104 new)! New models : DeepSeek-V4.1-Flash ( #56214 , #56228 , #56208 ) with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100 ( #56893 ), DeepGEMM Mega-mHC ( #56962 ), and async Engram prefetch with Engram DP sharding ( #56512 ); DeepSeek-V4-Flash-Vision-Exp (...
vLLM proto-v0.3.0
Release vllm-proto 0.3.0
proto-v0.2.0: vllm-proto 0.2.0
Validated by PR #56538 CI at fa2a26f .
vLLM proto-v0.1.0
vllm-proto 0.1.0
vLLM 0.29.0
v0.29.0 Highlights This release features 594 commits from 277 contributors (91 new)! Model Runner V2 is now the default for all models ( #53183 ), completing the rollout that began with pooling models ( #48290 ).
Also shipped on Oct 5, 2026
vLLM in October 2026Scheduled Search and Reporting is now generally available
New Relic is excited to share that Scheduled Search and Reporting is now generally available. Instead of manually re-running the same NRQL query to check on your data, Scheduled Search runs it for you, on whatever cadence you choose, and delivers the results straight to your inbox. What is Scheduled Search and Reporting?
Starting on July 1, 2026, Identity Service for GKE is deprecated in GKE version 1.36 and earlier
Starting on July 1, 2026, Identity Service for GKE is deprecated in GKE version 1.36 and earlier. This feature is also unavailable in organizations that were created on or after July 1, 2025. GKE version 1.37 and later don't support Identity Service for GKE.
GKE support for using the c4-standard-* machine types (up to 192 vCPUs) as Confidential GKE Nodes with Intel...
GKE support for using the c4-standard-* machine types (up to 192 vCPUs) as Confidential GKE Nodes with Intel TDX is generally available. For more information, see the following pages: To use this feature with GKE , see Encrypt workload data in-use with Confidential GKE Nodes .
AI Deep Scan with linked repositories | GitHub Cloud | Essentials and above, AI Deep Scan access, Beta
AI Deep Scan with linked repositories | You can now give AI Deep Scan context from linked repositories to analyze your primary repository more thoroughly. Source from dependencies, such as a shared authentication library, helps CodeRabbit investigate and verify security findings in the primary repository. Learn more in the linked repositories documentation .
Lyria 3 music models for your app's AI features
Your app's AI features can now generate music with two Google models, from a text prompt or an image, with vocals or as an instrumental: Lyria 3 Clip Preview , for clips of about 30 seconds. Lyria 3 Pro Preview , for full tracks up to 184 seconds. Ask Lovable to use either model in your app, or describe what you want and let Lovable pick.
iOS 27.2 beta 3 (24B5099f)
- View downloads View release notes