Microsoft's ThinkingBox: grade agents by database changes, not their words
SergioPaniego · x · 2026-10-05
Microsoft's ThinkingBox is now available as an OpenEnv environment. It's a sandbox for testing agents on business workflows: it simulates a customer, gives the agent MCP tools over a real database, then checks what actually changed in the database rather than trusting the agent's final message.
The accompanying ThinkingBox-Bench covers 507 tasks across retail, insurance, travel, banking and consulting. Each task runs as an episode in an isolated backend and returns a pass/fail reward from those backend-state checks.
The HF article gives an example: an agent makes nine well-formed tool calls and closes the ticket, yet the carrier exception remains open and the customer never got a real answer — exactly the gap this benchmark is built to catch.
More from coding & agent
- HuggingFace tackles harness overfitting with multi-harness RL across Claude Code, Codex and more — huggingface · 2026-10-05
- Garry Tan: Lab-built harnesses burn tokens, and that's why startup harnesses like Grep matter — garrytan · 2026-10-05
- Dev builds a color-bar planning board to visualize long-running agent sessions — kevinkern · 2026-10-05
- Security-One: open-weight 27B model outputs probabilities for agent security decisions — huggingface · 2026-10-05
- 3D game scene built in under an hour of prompting with GPT Sol 6.1 — aitrendz_xyz · 2026-10-05
- A week of GPT-6.1 Sol + Three.js yields a playable browser game — aitrendz_xyz · 2026-10-05