PhD Team Building Open-Source Benchmark for Realistic LLM/Agent Workflows, Seeks Pain Points

Groofy_beautypie · reddit · 2026-10-02

Duplicate post of the same announcement: a mostly-PhD team is developing an open-source benchmark targeting realistic LLM/agent workflows and soliciting cases where existing benchmarks fall short.

Related event: PhD Team Builds Open Benchmark for Real-World Agent Workflows(2 posts)→

Original post →

More from Research

Research channel →