WorkflowEvals: typesafe open collection for evaluating agents on real workflows
_lewtun · x · 2026-09-30
typesafe released WorkflowEvals, a typesafe collection of agent workflow evaluation datasets with a reproducible GitHub repo and an "anti-benchmaxxing" blog post.
The collection covers multiple real-world scenarios:
- Invoice processing (evalsafe-invoice-processing): 8.37k items
- Customer service (evalsafe-customer-service): 5.33k items
- Security incidents (evalsafe-security-incidents): 4.22k items
- Agent trace observability (evalsafe-agent-trace-observability): 2.23k items
The release argues for evaluating agents on realistic workflows rather than gaming leaderboards, shipping datasets, eval tooling, and reproduction code together.
More from coding & agent
- Dev builds Rez-inspired music rail shooter with Sonnet then Opus, playable in browser — AIandDesign · 2026-09-30
- Ramp demos agentic buying: Codex agent picks wine via Stripe MCP and pays with a Ramp card over MPP — jeff_weinstein · 2026-09-30
- AI agent starts reverse engineering NVIDIA drivers to run game benchmarks — mgostIH · 2026-09-30
- ModRetro console ships with blank cartridges you fill by coding games with Codex — OpenAIDevs · 2026-09-30
- Beam Studio launches on Bittensor Subnet 105 as a coordination layer for AI agents — markjeffrey · 2026-09-30
- Open-source project brings LoRA training to 50+ models — no3us · 2026-09-30