WorkflowEvals: typesafe open collection for evaluating agents on real workflows

_lewtun · x · 2026-09-30

typesafe released WorkflowEvals, a typesafe collection of agent workflow evaluation datasets with a reproducible GitHub repo and an "anti-benchmaxxing" blog post.

The collection covers multiple real-world scenarios:

The release argues for evaluating agents on realistic workflows rather than gaming leaderboards, shipping datasets, eval tooling, and reproduction code together.

Original post →

More from coding & agent

coding & agent channel →