Story 02 October 6, 2026 issue
Daily AI-generated issue
vLLM v0.31.0 cuts ROCm decode latency 13x and adds fast-restart weight caching
Top Repo · 6 Stars today
vLLM v0.31.0 adds a preload CLI that keeps post-quantized weights in GPU memory across engine restarts. It also broadens serving across hardware in one large community release (717 commits from 307 contributors).
Here's what changed:
- The preload CLI runs a weight-cache daemon and supports data parallelism, MTP draft models, health checks and readiness waits.
- Serving improves across NVIDIA CUDA, AMD ROCm, Intel XPU and CPU, with platform-specific wheels and Docker images.
- vLLM reports a ROCm Hy4 path with 13x lower decode latency, fused KimiViT QK RoPE up to 29x faster, and Arm CPU paged attention up to 25% faster.
- Per-request multimodal processor and media I/O kwargs are now rejected by default unless you opt in with the trust flag.
One catch: slow tokenizer mode is gone and the online FP8 quantization shorthand changed, so check your launch scripts before upgrading.
Try it: pip install vllm for CUDA 13.0, or docker pull vllm/vllm-openai:v0.31.0.
Sources: GitHub
This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.