THE FORWARD PASS
Archive

Daily edition October 6, 2026

Daily AI-generated issue

AI engineering, October 6, 2026

Reflection introduces 501B open-weight Beam, vLLM v0.31.0 cuts ROCm decode latency 13x, Reka unveils 19B omni model Rho-1, two-token cue lifts Olmo-3-7B to 78% on MATH-500

This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

  • Agents
  • APIs
  • Audio
  • Benchmarks
  • Business
Edition
Daily
Published
October 6, 2026
Read time
4 min
Stories
3

One 19B network handles text, images, video, reasoning and actions. Rho-1's distilled variant generated a 5.3-second video in about a second in Reka's internal tests, and Reka trained it from scratch on 320 H100s in three months.

Meanwhile, Reflection's 501B open-weight Beam scores 90.5 on GPQA Diamond, close to GLM 5.2 at 91.2, with Apache 2.0 weights planned. And vLLM v0.31.0 is out on PyPI for CUDA 13.0, with wheels and Docker images for ROCm, XPU and CPU.

If you read one thing today: the vLLM v0.31.0 release. You can install it now, and its breaking changes (multimodal kwargs rejected by default, slow tokenizer mode removed) can break your deployment if you upgrade blind.

Key takeaways

  1. 01On GPQA Diamond, Beam scores 90.5, against 91.2 for GLM 5.2, 91.7 for GLM 5.3, 93.5 for Kimi K3 and 90.9 for DeepSeek V4.1 Flash.
  2. 02On SWE-Bench Verified, it hits 80.9, ahead of Inkling (77.6) and Nemotron 3 Ultra (70.7).
  3. 03Reflection plans to release the weights under Apache 2.0, with tools for running, evaluating and fine-tuning.
  4. 04Reflection pitches inference efficiency over raw capability, with 23B parameters active.

Top News · 383 HN points

Reflection positions Beam as a Western entry in the open-weight frontier, and it is the company's first open-weight model. It is a sparse MoE with 501B total and 23B active parameters, aimed at coding, reasoning and agentic work.

Here's what changed:

  • On GPQA Diamond, Beam scores 90.5, against 91.2 for GLM 5.2, 91.7 for GLM 5.3, 93.5 for Kimi K3 and 90.9 for DeepSeek V4.1 Flash.
  • On SWE-Bench Verified, it hits 80.9, ahead of Inkling (77.6) and Nemotron 3 Ultra (70.7).
  • Reflection plans to release the weights under Apache 2.0, with tools for running, evaluating and fine-tuning.
  • Reflection pitches inference efficiency over raw capability, with 23B parameters active.

One catch: Beam is text-only, and there are no API details or pricing yet.

Why care? A planned permissive license and a small active footprint make it one to watch for self-hosted coding work.

Try it: join Reflection's early access waitlist.

Sources: reflection.ai

Top Repo · 6 Stars today

vLLM v0.31.0 adds a preload CLI that keeps post-quantized weights in GPU memory across engine restarts. It also broadens serving across hardware in one large community release (717 commits from 307 contributors).

Here's what changed:

  • The preload CLI runs a weight-cache daemon and supports data parallelism, MTP draft models, health checks and readiness waits.
  • Serving improves across NVIDIA CUDA, AMD ROCm, Intel XPU and CPU, with platform-specific wheels and Docker images.
  • vLLM reports a ROCm Hy4 path with 13x lower decode latency, fused KimiViT QK RoPE up to 29x faster, and Arm CPU paged attention up to 25% faster.
  • Per-request multimodal processor and media I/O kwargs are now rejected by default unless you opt in with the trust flag.

One catch: slow tokenizer mode is gone and the online FP8 quantization shorthand changed, so check your launch scripts before upgrading.

Try it: pip install vllm for CUDA 13.0, or docker pull vllm/vllm-openai:v0.31.0.

Sources: GitHub

Top News

Reka presents Rho-1 as an alternative to multi-model agentic pipelines. It is one 19B network, trained from scratch, that handles text, images, video, reasoning and actions. Reka calls it a proof of concept.

Here's what makes this credible:

  • A distilled variant cuts denoising from 99 steps to 8 and generated a 5.3-second video in about a second in Reka's internal tests.
  • The first clip arrives in 7.0 seconds, against 13.8 seconds for an illustrative multi-agent pipeline.
  • Action channels are native to the model, not added through a robotics wrapper.
  • Understanding and generation expert streams share attention and a KV cache, trained with next-token prediction and flow matching.
  • Training took 320 H100 GPUs for three months.

One catch: it is a research preview with structural drift in long rollouts, unreliable object grounding and native video capped at 672×384.

Why care? Reka pitches Rho-1 as a unified alternative for real-time simulation and robotics.

Sources: reka.ai · reka.ai

Also in this issue

Signals

Keep reading