THE FORWARD PASS
October 6, 2026 issue

Story 02 October 6, 2026 issue

Daily AI-generated issue

vLLM v0.31.0 cuts ROCm decode latency 13x and adds fast-restart weight caching

Top Repo · 6 Stars today

vLLM v0.31.0 adds a preload CLI that keeps post-quantized weights in GPU memory across engine restarts. It also broadens serving across hardware in one large community release (717 commits from 307 contributors).

Here's what changed:

  • The preload CLI runs a weight-cache daemon and supports data parallelism, MTP draft models, health checks and readiness waits.
  • Serving improves across NVIDIA CUDA, AMD ROCm, Intel XPU and CPU, with platform-specific wheels and Docker images.
  • vLLM reports a ROCm Hy4 path with 13x lower decode latency, fused KimiViT QK RoPE up to 29x faster, and Arm CPU paged attention up to 25% faster.
  • Per-request multimodal processor and media I/O kwargs are now rejected by default unless you opt in with the trust flag.

One catch: slow tokenizer mode is gone and the online FP8 quantization shorthand changed, so check your launch scripts before upgrading.

Try it: pip install vllm for CUDA 13.0, or docker pull vllm/vllm-openai:v0.31.0.

Sources: GitHub

This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.