Story 03 October 5, 2026 issue
Daily AI-generated issue
Microsoft ThinkingBox shows Claude Opus 5.5 passes only 47.5% of tasks every time
Top News
ThinkingBox adds repeatability checks to single-attempt success scores. Built by Microsoft and Hugging Face, it runs each enterprise task 20 times from a clean state and grades the final backend state, not the text or tool-call syntax.
Here is what makes this credible:
- It covers 507 simulated business tasks across retail, auto insurance, travel, neobanking and consulting.
- Claude Opus 5.5 completed 47.5% of tasks on every attempt and GPT-6 Astra hit 45.6%.
- Claude Opus 5.5 led single-attempt accuracy at 67.16%.
- GPT-6 Astra kept 78% of its baseline across 20 trials, while several models kept only 8%.
- Kimi-K3 solved over 93% of workflows at least once but passed all 20 runs on fewer than 14%.
About 80% of breakdowns came from tool handling, error recovery and precondition failures, not reasoning. Sessions are isolated and credentials stay behind a strict trust boundary.
One catch: cost comparisons are estimates based on token pricing.
Try it: access ThinkingBox through OpenEnv on Hugging Face.
Sources: hyper.ai
This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.