---
title: "AI engineering, October 9, 2026"
description: "The Forward Pass daily issue, October 9, 2026."
canonical: https://theforwardpass.net/archive/daily/2026-10-09
updated: 2026-10-09
---

# AI engineering, October 9, 2026

The Forward Pass daily issue, October 9, 2026. Source: https://theforwardpass.net/archive/daily/2026-10-09

> This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

_GPT-6 brings Intelligent UI to ChatGPT, Anthropic cuts Haiku costs, vLLM cuts DeepSeek time to first token, SambaCloud discounts MiniMax M3 cached tokens_

Haiku 4.5 scored 0.0% on Terminal-Bench 4.0. Haiku 5.5 now scores 39.2% on the same test.

Meanwhile, GPT-6 is launching globally in ChatGPT with Intelligent UI. DeepSeek has open-sourced DeepSeek-V4.1-Flash and its kernels.

If you read one thing today: Claude Haiku 5.5. The 75% average cost reduction gives you a cheaper option for high-volume summaries and classification.

## GPT-6 launches globally in ChatGPT with faster responses and Intelligent UI

**Top News** · 751 HN points

GPT-6 is rolling out globally in ChatGPT with faster responses and Intelligent UI. The release pairs the new model with visuals and interactive experiences that you can explore and use directly. The launch brings the model and the interface together inside ChatGPT, where users can interact directly with those experiences.

Here's what changed:
- GPT-6 is rolling out globally in ChatGPT, alongside the introduction of Intelligent UI.
- Faster responses are part of the launch, alongside changes to how users interact with results.
- Intelligent UI includes visuals and interactive experiences that users can explore and use directly.

Why care? If you work in ChatGPT, this is a model launch and an interface change to assess together.

Sources: [openai.com](https://openai.com/index/gpt-6-for-everyone)

## Anthropic launches Claude Haiku 5.5 at 75% lower average cost

**Top News** · 1,040 HN points

High-volume summaries, classification, and subagent coding need a model you can afford to call repeatedly. Anthropic launches Claude Haiku 5.5 at a stated 75% lower average cost, with adjustable effort to trade cost against capability.

Here's what changed:
- Prompts up to 100,000 tokens cost 90% less, at $0.10 per million input tokens and $0.50 per million output tokens.
- Prompts above 100,000 tokens cost 50% less, at $0.50 for input and $2.50 for output per million tokens.
- Anthropic reports 39.2% on Terminal-Bench 4.0, versus Haiku 4.5's 0.0% and GPT-6 Luna's 16.4%.
- Adjustable effort is new to the Haiku class, letting you choose the cost-capability tradeoff.

One catch: The changed tokenizer consumes slightly more tokens per task.

Try it: Use `claude-haiku-5-5` on the Claude Platform, Amazon Web Services, Google Cloud or Microsoft Azure. Also available via Vercel, Amazon Web Services, Databricks, LiteLLM

Sources: [anthropic.com](https://www.anthropic.com/claude-haiku-5-5)

## vLLM cuts DeepSeek-V4.1-Flash time to first token by nearly 70%

**Top News**

If you serve DeepSeek-V4.1-Flash, vLLM and Inferact have sped up the path from prompt to output. Their combined optimizations cut time to first token by nearly 70% at around 100K throughput.

Here's what changed:
- AgentX throughput improved about 5x versus vLLM's day-0 implementation. Low-latency performance improved 1.9x.
- Decoder CUDA graphs plus bounded replay cut prefill computation time by 30-40%.
- DeepSeek's open-source MegaAttention kernel uses an NVFP4 KV format reported as 45% smaller than the prior FP8 KV cache.
- Bounded replay is enabled by default for DeepSeek-V4.1 and replays only the last 128 tokens.

One catch: Bounded replay trades exactness for less computation, and without CUDA graphs it can slow short prompts. vLLM found no meaningful accuracy difference on GSM8K and GPQA.

Sources: [vllm.ai](https://vllm.ai/blog/2026-10-07-deepseek-v41-flash)

## Signals

1. [JetBrains releases Mellum2.1 coding model with 2.5B active parameters for local agents](https://blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/)
2. [SambaCloud adds MiniMax M3 prompt caching with a 90% discount on cached tokens](https://sambanova.ai/blog/introducing-prompt-caching-for-minimax-m3-on-sambacloud)
3. [TokenRouter runs token-level LLM routing with 2.01-64.15x higher decoding throughput than existing systems](https://huggingface.co/papers/2610.12242) · 97 Upvotes
4. [reViT-B/16 matches DeiT III vision accuracy with about 70% fewer stored parameters](https://huggingface.co/papers/2610.12448) · 2 Upvotes
5. [Docker Agent runs AI agent teams from YAML configs without writing code](https://github.com/docker/docker-agent) · 301 HN points
6. [Underdog Saluki cuts Qwen3.8-27B file size from 54 GB to 7.89 GB](https://huggingface.co/ConwayResearch/Underdog-Saluki-27B-1.0) · 15,274 Downloads
