Development·EASYHUB JOURNAL
vLLM 0.31.0 adds GPU-resident weight caching for fast restarts, initialized-engine snapshots and a new wave of large-scale serving optimizations

What changed
vLLM released 0.31.0 on October 5 with 717 commits from 307 contributors. The new vllm preload CLI starts a weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts, while experimental vllm snapshot create/restore uses CRIU to restore a fully initialized TP1 engine. The release also expands Model Runner V2 speculative decoding, large-scale serving through MoonEP/DeepEPv2, scheduling and KV-offload controls, with major work for DeepSeek-V4.1-Flash, GLM-5.3-Flash and Kimi K3. The project also documents several breaking changes, so quantization and multimodal request settings should be checked before upgrading.
- Original title
- v0.31.0
- Source
- vLLM · github.com
- Topic
- Development
- Source month
- 2026-10
This is a concise EasyHub summary of the linked source, not the full report or original reporting. Availability and preview conditions are described in the summary and original.
Summary page published · Updated · Editorial information