PhD Team Building Open-Source Benchmark for Realistic LLM/Agent Workflows, Seeks Pain Points
Groofy_beautypie · reddit · 2026-10-02
Duplicate post of the same announcement: a mostly-PhD team is developing an open-source benchmark targeting realistic LLM/agent workflows and soliciting cases where existing benchmarks fall short.
Related event: PhD Team Builds Open Benchmark for Real-World Agent Workflows(2 posts)→
More from Research
- Clarification as Supervision lands NeurIPS Oral: denser training signal via model interaction — iatitov · 2026-10-02
- Ego-Exo4D-HM: SMPL-H reconstructions for 523 hours of egocentric-exocentric video, open-sourced — geopavlakos · 2026-10-02
- Cohere Labs Talk: Making AI Math Reasoning Machine-Checkable with Lean — Cohere_Labs · 2026-10-02
- New paper dissects only task-relevant network parts, making mechanisms inspectable and editable at far lower cost — leedsharkey · 2026-10-02
- Transformers can't hide reasoning but may cryptographically obfuscate their CoTs — gsarti_ · 2026-10-02
- CloudAnalyzer: Browser-based LiDAR SLAM map correction and Lanelet2 editing in Rust + WASM — rsasaki0109 · 2026-10-02