TypeSafe's Workflow Evals: decomposing policies into workflows beats prompts on accuracy, cost and speed

multiply_matrix · x · 2026-09-30

Two weeks after launch, eval platform TypeSafe is making its released datasets easier to work with — while doubling down on being anti-public-benchmark by deprecating all datasets it evaluates on.

Method: instead of asking a model to solve a whole task in one prompt, TypeSafe decomposes policies into programmatic rules plus three question types (Noul yes/no, Choice, Score), then assembles outputs in code. Using expense-approval as a toy example, every policy sentence becomes a typed question or a code rule.

Result: averaged over four example tasks, every model is more accurate, cheaper and faster running the same policy as a workflow than as a prompt. Scatter plots of accuracy vs cost and time show no frontier model is simultaneously cheaper-and-more-accurate or faster-and-more-accurate — each trades off differently. The project assumes the harness is correct and focuses on measuring models, not debating labels.

Original post →

More from coding & agent

coding & agent channel →