# vLLM v0.31.0 cuts ROCm decode latency 13x and adds fast-restart weight caching

From The Forward Pass daily issue, October 6, 2026 (https://theforwardpass.net/archive/daily/2026-10-06). Source: https://theforwardpass.net/archive/daily/2026-10-06/vllm-v0-31-0-cuts-rocm-decode-latency-13x-and-adds-fast-restart-weight-caching

> This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

**Top Repo** · 6 Stars today

vLLM v0.31.0 adds a preload CLI that keeps post-quantized weights in GPU memory across engine restarts. It also broadens serving across hardware in one large community release (717 commits from 307 contributors).

Here's what changed:
- The preload CLI runs a weight-cache daemon and supports data parallelism, MTP draft models, health checks and readiness waits.
- Serving improves across NVIDIA CUDA, AMD ROCm, Intel XPU and CPU, with platform-specific wheels and Docker images.
- vLLM reports a ROCm Hy4 path with 13x lower decode latency, fused KimiViT QK RoPE up to 29x faster, and Arm CPU paged attention up to 25% faster.
- Per-request multimodal processor and media I/O kwargs are now rejected by default unless you opt in with the trust flag.

One catch: slow tokenizer mode is gone and the online FP8 quantization shorthand changed, so check your launch scripts before upgrading.

Try it: `pip install vllm` for CUDA 13.0, or `docker pull vllm/vllm-openai:v0.31.0`.

Sources: [GitHub](https://github.com/vllm-project/vllm/releases/tag/v0.31.0)
