October 9, 2026 issue

Story 03 October 9, 2026 issue

Daily AI-generated issue

vLLM cuts DeepSeek-V4.1-Flash time to first token by nearly 70%

Image: vllm.ai

Top News

If you serve DeepSeek-V4.1-Flash, vLLM and Inferact have sped up the path from prompt to output. Their combined optimizations cut time to first token by nearly 70% at around 100K throughput.

Here's what changed:

  • AgentX throughput improved about 5x versus vLLM's day-0 implementation. Low-latency performance improved 1.9x.
  • Decoder CUDA graphs plus bounded replay cut prefill computation time by 30-40%.
  • DeepSeek's open-source MegaAttention kernel uses an NVFP4 KV format reported as 45% smaller than the prior FP8 KV cache.
  • Bounded replay is enabled by default for DeepSeek-V4.1 and replays only the last 128 tokens.

One catch: Bounded replay trades exactness for less computation, and without CUDA graphs it can slow short prompts. vLLM found no meaningful accuracy difference on GSM8K and GPQA.

Sources: vllm.ai

This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.