vLLM 0.31.0
v0.31.0 Highlights This release features 717 commits from 307 contributors (96 new)! - DeepSeek-V4.1-Flash performance: FlashMLA mega attention with the V4.1 NVFP4 compressed KV cache is now the SM100 default ; DeepGEMM sparse MQA logits for the indexer and Mega-Gate fusing the gate GEMM with expert selection ; decoder boundaries fuse the TP all-reduce, mHC...