THE FORWARD PASS
October 5, 2026 issue

Story 03 October 5, 2026 issue

Daily AI-generated issue

Microsoft ThinkingBox shows Claude Opus 5.5 passes only 47.5% of tasks every time

Top News

ThinkingBox adds repeatability checks to single-attempt success scores. Built by Microsoft and Hugging Face, it runs each enterprise task 20 times from a clean state and grades the final backend state, not the text or tool-call syntax.

Here is what makes this credible:

  • It covers 507 simulated business tasks across retail, auto insurance, travel, neobanking and consulting.
  • Claude Opus 5.5 completed 47.5% of tasks on every attempt and GPT-6 Astra hit 45.6%.
  • Claude Opus 5.5 led single-attempt accuracy at 67.16%.
  • GPT-6 Astra kept 78% of its baseline across 20 trials, while several models kept only 8%.
  • Kimi-K3 solved over 93% of workflows at least once but passed all 20 runs on fewer than 14%.

About 80% of breakdowns came from tool handling, error recovery and precondition failures, not reasoning. Sessions are isolated and credentials stay behind a strict trust boundary.

One catch: cost comparisons are estimates based on token pricing.

Try it: access ThinkingBox through OpenEnv on Hugging Face.

Sources: hyper.ai

This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.