# AI engineering, October 2, 2026

The Forward Pass daily issue, October 2, 2026. Source: https://theforwardpass.net/archive/daily/2026-10-02

> This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

_Pi ships 1.0 agent harness, Cloudflare open-sources Clef decision models, GitHub Copilot adds desktop computer use, GPT-6.1 Sol nears Astra at one-fifth the price_

Your coding agent can now click buttons in apps that never had an API. In public preview, GitHub Copilot CLI and the Copilot app on macOS and Windows can drive legacy, GUI-only desktop software, and Copilot asks for approval before controlling an app.

Meanwhile, Pi shipped version 1.0 of its open-source agent harness, installable today with one shell command or PowerShell on Windows. Cloudflare released Clef, open decision models that answer in 209.3 ms median latency where Jev takes 524.1 ms, hosted on Workers AI and open on Hugging Face under Apache 2.0.

If you read one thing today: the Clef story. You can pull Apache 2.0 weights or call Workers AI and replace slow LLM classification calls with structured answers that carry probabilities your code can branch on.

## Pi ships 1.0 agent harness with Codemode, native MCP and cache warming

**Top News** · 984 HN points

Pi is a minimal, extensible agent harness that hundreds of thousands of people use weekly. After months of feedback and hardening, it hits 1.0 and stays small while adding features that have proved useful.

Here's what changed:
- Codemode lets the agent write scripts for tasks, such as summarizing your commit history.
- Native MCP support and non-LLM model support land in the core.
- Virtual-model extensions let you route work, for example planning to Claude Opus and implementation to GPT.
- Pi adds deferred tool loading and Anthropic cache warming.
- System messages can change mid-conversation with awareness of the transcript, and full-screen mode with a new TUI theme is now the default.

There is also Pi Durable, a separate package for building long-running agentic applications beyond the terminal. Both are MIT licensed.

One catch: Pi Durable is experimental.

Try it: `curl -fsSL https://pi.dev/install.sh | sh` (Windows: `powershell -c "irm https://pi.dev/install.ps1 | iex"`). For Durable, run `npm install @earendil-works/pi-durable @earendil-works/pi-ai @earendil-works/chord`. Code: github.com/earendil-works/pi.

Sources: [earendil.com](https://earendil.com/posts/pi-1-0/)

## Cloudflare open-sources Clef decision models at 209.3 ms versus Jev's 524.1 ms

**Top News** · 472 HN points

Some LLM calls exist only to pick an option, like classifying a website. Cloudflare's Clef and Clef-flash are decision models built for that job. They classify inputs and return structured answers with probabilities your software can act on.

Here is what makes this credible:
- Across 43 benchmarks, Clef's median latency is 209.3 ms and Clef-flash's is 38.8 ms, versus Jev at 524.1 ms.
- On BFCL case exact, Clef scores 98.47 and Clef-flash 98.76, versus Jev at 95.75.
- Inference runs one prefill-only pass, then scores valid schema choices in parallel instead of generating text.
- Clef adds image input and a 64k context window, where Jev is text-only at 32k.
- The backbones are Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, and the API is Jev-compatible.

Fine-tuning starts as a hands-on service with Cloudflare engineers, with self-serve planned later.

One catch: Jev still leads on When2Call accuracy (80.97 versus 72.37) and BRIGHT nDCG@10.

Try it: POST to `https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef` with a bearer token, or download the Apache 2.0 weights from Hugging Face.

Sources: [blog.cloudflare.com](https://blog.cloudflare.com/clef-decision-models/)

## GitHub Copilot adds computer use for desktop apps on macOS and Windows

**Top News**

Some of the work you want to automate lives in desktop software with no API, CLI or MCP integration. Copilot can now operate those apps directly.

Here's what you can build with it:
- Copilot reads accessible content and visual context, then clicks controls, types text and moves across applications.
- It can summarize browser notifications, update a presentation or push data through a desktop workflow.
- It asks for approval before controlling an app, and you can review or reset apps set to always allow.
- Prompts work best when they state the outcome, the apps involved and key constraints.

The feature is in public preview in Copilot CLI and the GitHub Copilot app on macOS and Windows.

One catch: macOS needs Accessibility and Screen Recording permissions, and organization-managed settings can disable the feature.

Try it: run `/computer on` in Copilot CLI (`/computer show` checks status, `/computer off` disables it). In the app, open Settings → Computer Use → Enable Computer Use.

Sources: [github.blog](https://github.blog/changelog/2026-10-01-github-copilot-can-now-interact-with-desktop-apps/) · [github.blog](https://github.blog/changelog/2026-10-01-github-copilot-can-now-interact-with-desktop-apps)

## Signals

1. [OpenAI's GPT-6.1 Sol nears Astra on complex work at one-fifth the token price](https://community.openai.com/t/gpt-6-1-sol-in-the-api-a-meaningful-step-up-in-cost-performance/1402388)
2. [Context Language Models manage their own context, scoring 11.4% higher with 21.5% fewer FLOPs](https://arxiv.org/abs/2609.37725) · 132 HN points
3. [Qwen3.8-27B-pi matches base xhigh on Terminal-Bench at medium effort with 41% fewer tokens](https://huggingface.co/blog/bytkim/qwen38-27b-pi)
4. [Context Mode MCP server sandboxes coding agent tool output, cutting context use by 98%](https://github.com/mksglu/context-mode) · 20 Stars today
5. [Meta paper finds base LLMs can beat post-trained versions on agentic pass@K given enough samples](https://huggingface.co/papers/2610.01509) · 20 Upvotes
6. [webAI's 3.66B TwIL-LM3-Pro beats VibeThinker-3B on all six formal-logic lanes](https://huggingface.co/webAI-Official/TwIL-LM3-Pro)
