# AI engineering, October 1, 2026

The Forward Pass daily issue, October 1, 2026. Source: https://theforwardpass.net/archive/daily/2026-10-01

> This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

_Google discounts Gemini 4 Argon cached input tokens by 95%, GPT-6.1 Sol costs one-fifth of Astra rates, Cloudflare starts 100 sandboxes 6.2x faster, Ollama adds decision-model support_

Gemini 4 Argon’s output-token limit is 1 million, up from 64,000. That gives long coding jobs more room in a single response.

Meanwhile, GPT-6.1 Sol costs one-fifth of Astra’s standard API prices while offering near-Astra intelligence for coding, computer use, and professional work. Cloudflare cut median startup for 100 concurrent sandboxes from 4.049 seconds to 648 milliseconds.

If you read one thing today: Cloudflare’s `durable_object` policy lets your Durable Object choose a sandbox image and instance at runtime instead of requiring a separate deployed app for each combination. That’s a practical change if your agents need different environments.

## Google discounts Gemini 4 Argon cached input tokens by 95%

**Top News** · 1,184 HN points

Google is opening Gemini 4 Argon in stages for long-horizon coding, enterprise work, multimodal tasks, and cybersecurity defense. Its announced introductory API rates are $2 per million input tokens and $10 per million output tokens.

Here’s what changed:
- Cached input tokens cost 95% less, which can lower bills when a request reuses prior context.
- Argon’s output-token limit is 1 million, up from 64,000, leaving more room for long responses.
- Google reports 77.9% on DeepSWE v1.1 and 68% on CWE-bench v1, tied for first.
- Google says the model can find, validate, and patch critical software vulnerabilities.

One catch: Access is currently limited to trusted cyber defenders through the Fairwind Program and a cohort of trusted testers. Google plans broader access starting with paid API customers and Google AI Ultra subscribers.

Sources: [blog.google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/) · [deepmind.google](https://deepmind.google/blog/gemini-4-argon-our-next-era-of-frontier-intelligence/)

## GPT-6.1 Sol costs one-fifth of Astra’s standard API prices

**Top News** · 1,052 HN points

GPT-6.1 Sol targets teams that want near-Astra intelligence for coding, computer use, and professional work without paying Astra’s standard API rates. Its standard input and output token prices are each one-fifth of Astra’s.

Here’s what changed:
- Standard input and output rates both fall to one-fifth of Astra’s prices, so the reduction applies to either side of a request.
- Sol is described as offering near-Astra intelligence for coding, giving engineers a cheaper option to evaluate for coding workloads.
- Computer use and professional work are also listed as areas where Sol offers near-Astra intelligence.
- GPT-6.1 Sol is listed on Databricks, Vercel, Amazon Web Services, and LiteLLM.

Try it: Select GPT-6.1 Sol on a supported platform. Also available via Databricks, Vercel, Amazon Web Services, LiteLLM.

Sources: [openai.com](https://openai.com/index/introducing-gpt-6-1-sol)

## Cloudflare starts 100 concurrent sandboxes 6.2x faster at the median

**Top News**

Cloudflare Containers are Linux workspaces for agents, and the latest changes make them quicker to start and easier to configure. In ComputeSDK’s Burst TTI Benchmark, which launched 100 sandboxes concurrently, median startup fell from 4.049 seconds to 648 milliseconds, a 6.2x improvement.

Here’s what changed:
- The same test measured 910 milliseconds at p95 and 1,129 milliseconds at p99.
- The `durable_object` scheduling policy is in public beta for all and lets a Durable Object select an image and instance type at runtime, rather than requiring a separate deployed application for every pairing.
- Filesystem snapshots can save and restore a workspace, including files and setup, or provide a shared starting point for isolated sandboxes.
- The prebuilt `cloudflare/debian-trixie` image includes Debian Trixie Slim and Node.js 24.20.0 LTS, and can start without first building and pushing a Docker image.

One catch: The new capabilities are available only through the native `ctx.container` API. Legacy Container and Sandbox classes will be maintained through December 31, 2026, but will not receive updates after that date.

Try it: In your Durable Object, declare the images it can use, then call `this.ctx.container.start({ image, instance, ... })`. For a ready-made base, use `cloudflare/debian-trixie`.

Sources: [blog.cloudflare.com](https://blog.cloudflare.com/faster-agent-sandboxes/)

## Signals

1. [Ollama adds decision-model support that returns choices, probabilities, and scores instead of text](https://github.com/ollama/ollama/releases/tag/v0.35.0) · 6 Stars today
2. [Magnitude ships an inference engine that runs open models up to 2x faster than llama.cpp](https://github.com/magnitudedev/magnitude) · 142 HN points
3. [Google's TabFM delivers zero-shot predictions and ranks first among default models on 51 TabArena datasets](https://huggingface.co/papers/2609.37959) · 8 Upvotes
4. [EngiWorld benchmarks engineering agents across 1,301 tasks, best model scores 44.3](https://huggingface.co/papers/2609.37686) · 52 Upvotes
5. [Agent Error Dataset releases 50,228 diagnosis pairs, corrections lift verifier pass rates from 18.4% to 51.1%](https://huggingface.co/papers/2609.40111) · 15 Upvotes
6. [JEV-like decision models compress ordinal choices, using 67-76% of gold support across 36 datasets](https://huggingface.co/papers/2609.38827) · 26 Upvotes
