TypeSafe's Workflow Evals: decomposing policies into workflows beats prompts on accuracy, cost and speed
multiply_matrix · x · 2026-09-30
Two weeks after launch, eval platform TypeSafe is making its released datasets easier to work with — while doubling down on being anti-public-benchmark by deprecating all datasets it evaluates on.
Method: instead of asking a model to solve a whole task in one prompt, TypeSafe decomposes policies into programmatic rules plus three question types (Noul yes/no, Choice, Score), then assembles outputs in code. Using expense-approval as a toy example, every policy sentence becomes a typed question or a code rule.
Result: averaged over four example tasks, every model is more accurate, cheaper and faster running the same policy as a workflow than as a prompt. Scatter plots of accuracy vs cost and time show no frontier model is simultaneously cheaper-and-more-accurate or faster-and-more-accurate — each trades off differently. The project assumes the harness is correct and focuses on measuring models, not debating labels.
More from coding & agent
- Pi Puts MCP at Its Core After Once Publicly Refusing It — omarsar0 · 2026-09-30
- One prompt before bed: agent delivers a stack of verified performance PRs via Alchemy — samgoodwin89 · 2026-09-30
- Developer Explores What AI Agents Should Actually Remember in Long-Term Memory Design — Bh0nu0077 · 2026-09-30
- Why production agent builders say OpenAI's Agents API won't kill agent frameworks — BhAAI777 · 2026-09-30
- LlamaIndex's Jerry Liu and Snorkel AI on why evals and RL environments remain unsolved — ajratner · 2026-09-30
- Tibo gets his Codex reset live on stream, viewers call it a 'canonic event' — flavioAd · 2026-09-30