vLLM
vLLM keeps model weights in GPU memory across engine restarts
vLLM added a preload command that keeps weights in GPU memory through restarts, gated multimodal request settings behind a flag and removed several options.

vLLM, the open-source server for running large language models, added a command that keeps model weights in GPU memory across engine restarts. The same release blocks per-request multimodal settings unless an operator opts in. It also removes or renames several options that existing launch scripts may use.
The vLLM project published the update on GitHub, and its v0.31.0 release notes are dated Oct. 6, 2026. The notes list 717 commits from 307 contributors.
The vllm preload command holds weights through a restart
The new vllm preload command keeps post-quantized weights, the weights after they are compressed to lower precision, in GPU memory across engine restarts. Experimental vllm snapshot commands go further. They use CRIU, a Linux checkpoint and restore tool, to bring back an engine that was already initialized.
Multimodal request settings now need a server flag
vLLM now returns an error for per-request multimodal processor kwargs unless the server starts with --trust-request-mm-kwargs. The release also tags prefix-cache keys by their source to prevent collisions.
Removed and renamed options can break an existing setup
The release breaks existing launch scripts in several places, according to the notes.
tokenizer_mode="slow"is removed.--enable-mamba-fine-grained-prefix-cacheis renamed--enable-mamba-shared-prefix-checkpoint.quantization="fp8"now redirects to thefp8_per_tensorshorthand, and Quark silent online quantization is removed.- The AllSpark INT8 W8A16 backend is removed.
--enforce-eagernow also disables JIT kernel warmup unless fault tolerance is enabled.- XPU graphs are on by default, and
VLLM_XPU_ENABLE_XPU_GRAPHis removed. VLLM_PLE_CPU_OFFLOADis removed in favor of--engram-config.
DeepSeek-V4.1-Flash gets a new default attention path on Blackwell
FlashMLA mega attention with an NVFP4 compressed KV cache is now the default for DeepSeek-V4.1-Flash on SM100 GPUs, Nvidia’s Blackwell-class chips. The release adds several fused kernels and Engram sharding for the model.
Scheduling and large-scale serving get new controls
Model Runner V2 now supports draft-model speculative decoding and custom logits processors. It also adds new drafters and async scheduling for DFlash. Scheduling gets a new --max-num-active-seqs option, an adaptive long-prefill threshold and a reworked waiting queue.
For large deployments, vLLM added a MoonEP all-to-all backend, prefill context parallelism combined with data parallelism, and NCCL M2N weight transfer for reinforcement learning.
New model support covers Gemma, MiMo, GLM and Granite
The release adds DiffusionGemma structured generation and MiMo V2 MXFP4 mixture-of-experts support. Cohere2MoE now exposes auxiliary hidden states for EAGLE3 and DFlash drafters. GLM-5.2-MXFP4 and GLM-5.3-Flash Quark MXFP4 checkpoints are supported, along with an AMD-Quark mixed-precision DeepSeek-V4.1-Flash-MXFP4 checkpoint. Granite 4.2 gets a built-in granite_thinking_parser. The notes also list optimizations for Qwen3.8-Flash-Next, Kimi-K3 and MiniMax-M3.
Analysis
We think operators should read the breaking changes before the feature list. A removed option or a new error can stop a server from starting after an upgrade. A faster restart only helps once the server is running. Anyone who passes multimodal kwargs in individual requests will need to add the new flag and decide whether they trust those requests.
