EasyHubExplore
Explore
EN

Site appearance

Your color. Your style.

Accent colorRose
Visual styleSame content, fresh look

Soft gradients, dimensional icons

Applied: Rose · Studio. Saved in this browser.

Development·EASYHUB JOURNAL

vLLM 0.31.0 adds GPU-resident weight caching for fast restarts, initialized-engine snapshots and a new wave of large-scale serving optimizations

vLLMSource published
Editorial illustration: vLLM 0.31.0 adds GPU-resident weight caching for fast restarts, initialized-engine snapshots and a new wave of large-scale serving optimizations
EasyHub editorial illustration · Not a source photograph

What changed

vLLM released 0.31.0 on October 5 with 717 commits from 307 contributors. The new vllm preload CLI starts a weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts, while experimental vllm snapshot create/restore uses CRIU to restore a fully initialized TP1 engine. The release also expands Model Runner V2 speculative decoding, large-scale serving through MoonEP/DeepEPv2, scheduling and KV-offload controls, with major work for DeepSeek-V4.1-Flash, GLM-5.3-Flash and Kimi K3. The project also documents several breaking changes, so quantization and multimodal request settings should be checked before upgrading.

Read original source
Original title
v0.31.0
Source
vLLM · github.com
Topic
Development
Source month
2026-10

This is a concise EasyHub summary of the linked source, not the full report or original reporting. Availability and preview conditions are described in the summary and original.

Summary page published · Updated · Editorial information

Back to the news timeline