# Microsoft ThinkingBox shows Claude Opus 5.5 passes only 47.5% of tasks every time

From The Forward Pass daily issue, October 5, 2026 (https://theforwardpass.net/archive/daily/2026-10-05). Source: https://theforwardpass.net/archive/daily/2026-10-05/microsoft-thinkingbox-shows-claude-opus-5-5-passes-only-47-5-of-tasks-every-time

> This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

**Top News**

ThinkingBox adds repeatability checks to single-attempt success scores. Built by Microsoft and Hugging Face, it runs each enterprise task 20 times from a clean state and grades the final backend state, not the text or tool-call syntax.

Here is what makes this credible:

- It covers 507 simulated business tasks across retail, auto insurance, travel, neobanking and consulting.
- Claude Opus 5.5 completed 47.5% of tasks on every attempt and GPT-6 Astra hit 45.6%.
- Claude Opus 5.5 led single-attempt accuracy at 67.16%.
- GPT-6 Astra kept 78% of its baseline across 20 trials, while several models kept only 8%.
- Kimi-K3 solved over 93% of workflows at least once but passed all 20 runs on fewer than 14%.

About 80% of breakdowns came from tool handling, error recovery and precondition failures, not reasoning. Sessions are isolated and credentials stay behind a strict trust boundary.

One catch: cost comparisons are estimates based on token pricing.

Try it: access ThinkingBox through OpenEnv on Hugging Face.

Sources: [hyper.ai](https://hyper.ai/en/stories/016378578c1d36c27e1ac93e928d4a5e)
