PhD Team Building Open-Source Benchmark for Realistic LLM/Agent Workflows, Seeks Pain Points

Groofy_beautypie · reddit · 2026-10-02

A mostly-PhD team is developing an open-source benchmark targeting realistic LLM/agent workflows, aiming to fill gaps left by academic benchmarks.

Shortcomings they cite include:

They are soliciting cases of "needed this in production but had no good way to benchmark it."

Related event: PhD Team Builds Open Benchmark for Real-World Agent Workflows(2 posts)→

Original post →

More from Research

Research channel →