Microsoft's ThinkingBox: Kimi-K3 tops discovery at 93.89% but only 13.41% solve tasks 20/20
tuhin_k · reddit · 2026-10-09
- Microsoft researchers released Thinkingbox-Bench: 507 policy-conditioned business workflows across 5 domains, each run 20 times from an identical clean backend, graded on terminal database state (10,140 trials per model).
- Key finding: discovery and repeatability rank models in nearly reversed orders — Kimi-K3 solves 93.89% of tasks at least once but only 13.41% on all 20 attempts, while Claude Opus 5 covers fewer (79.09%) but repeats far more (47.53%). Qwen3.8-27B: 89.35% vs 7.50%.
- In an ablation over 121,680 trials, 67.24% of failures terminated cleanly with no tool error — a completion-style proxy would have scored them as done; 77.61% involved wrong field values, 43.30% unintended side effects.
- Paper, code, and dataset are public; tasks runnable via HF OpenEnv.
More from coding & agent
- WorldBox spatial memory boosts Opus 5.5 Minecraft progress 133% and GPT-6 Astra 50% — Lianhuiq · 2026-10-09
- After GPT-6 Made Generative UI Mainstream, This Dev Argues the Next Layer Is Generative Workflow — Over_Accountant_2311 · 2026-10-09
- Next Token podcast ep.5: the Personal Agent wars, software replication, and small-hardware opportunities — op7418 · 2026-10-09
- mattyp's build a bot takes voice feedback via Grok, feeding each report to an agent for fixes — mattyp · 2026-10-09
- Reverse-engineering a 3D pixel-emoji animation style into a Claude artifact in 10 minutes — justin_hart · 2026-10-09
- How Delphi's Predict Anything works: LLM-generated odds seed an LMSR market — benfielding · 2026-10-09