Harvey AI team shares its agent evals research agenda
sarahcat21 · x · 2026-09-02
The post highlights a talk by Spencer from the @harvey AI team, covering its research agenda: the right evals framework unlocks frontier research on topics ranging from knowledge compression to harness engineering.
More from coding & agent
- Fable 5.1 Tops Bug Hunt Bench, Outperforming GPT-5.6 — PawelHuryn · 2026-09-02
- Replit MCP Launches: Control Powerful Agents from Anywhere — amasad · 2026-09-02
- Claude Code 2.1.258 Released with macOS 12 Fixes — ClaudeCodeLog · 2026-09-02
- Claude Code 2.1.258 Released, Fixes macOS 12 Launch Bug — ClaudeCodeLog · 2026-09-02
- Automating Boring Business Tasks with 5 Skydive Agents — nima_owji · 2026-09-02
- Hands-on: Claude Code Feels Smarter with Better Writing — ivan_bezdomny · 2026-09-02