---
title: "vLLM cuts DeepSeek-V4.1-Flash time to first token by nearly 70%"
description: "From the The Forward Pass daily issue, October 9, 2026."
canonical: https://theforwardpass.net/archive/daily/2026-10-09/vllm-cuts-deepseek-v4-1-flash-time-to-first-token-by-nearly-70
updated: 2026-10-09
---

# vLLM cuts DeepSeek-V4.1-Flash time to first token by nearly 70%

From The Forward Pass daily issue, October 9, 2026 (https://theforwardpass.net/archive/daily/2026-10-09). Source: https://theforwardpass.net/archive/daily/2026-10-09/vllm-cuts-deepseek-v4-1-flash-time-to-first-token-by-nearly-70

> This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

**Top News**

If you serve DeepSeek-V4.1-Flash, vLLM and Inferact have sped up the path from prompt to output. Their combined optimizations cut time to first token by nearly 70% at around 100K throughput.

Here's what changed:
- AgentX throughput improved about 5x versus vLLM's day-0 implementation. Low-latency performance improved 1.9x.
- Decoder CUDA graphs plus bounded replay cut prefill computation time by 30-40%.
- DeepSeek's open-source MegaAttention kernel uses an NVFP4 KV format reported as 45% smaller than the prior FP8 KV cache.
- Bounded replay is enabled by default for DeepSeek-V4.1 and replays only the last 128 tokens.

One catch: Bounded replay trades exactness for less computation, and without CUDA graphs it can slow short prompts. vLLM found no meaningful accuracy difference on GSM8K and GPQA.

Sources: [vllm.ai](https://vllm.ai/blog/2026-10-07-deepseek-v41-flash)
