PatronusAI evals panel covers SpeedrunBench, world models, and spreadsheet agent evals
DynamicWebPaige · x · 2026-09-24
DynamicWebPaige previews an evals panel session with PatronusAI and gracegong, covering SpeedrunBench (how fast agents finish a game), benchmarks for world models, and evals for agent actions in spreadsheets with embedded formulas and charts.
More from coding & agent
- Why AI-made dev tools beat one-shot game generation: freedom of process — eschadiol · 2026-09-24
- RealSense VP on AgenticROS: Letting AI Agents Directly Control Physical Robots — chrismatthieu · 2026-09-24
- AWS API Gateway's Hard 10MB Upload Limit and the Presigned URL Fix — _jaydeepkarale · 2026-09-24
- Open-source Pragma gives coding agents a terminal-first workspace with Git worktrees — tech_w0rld · 2026-09-24
- Dev builds dense task annotation system with GPT-6 Astra, ships it as an LLM skill — chris_j_paxton · 2026-09-24
- Jev-as-a-Judge: hybrid agent eval flow escalates low-confidence calls to frontier models — omarsar0 · 2026-09-24