Build benchmarks: Logan says spend 25% of your time, plus Terminal-Bench-Science
BenBlaiszik · x · 2026-09-22
- @OfficialLoganK advises companies building with AI to spend >25% of their time creating benchmarks and getting model labs to care about them—"easiest path to accelerate your progress as a company."
- @StevenDillmann adds that the same applies to academic scientists, hence Terminal-Bench-Science: an open hub where researchers contribute problems they actually care about.
More from coding & agent
- Josh Rosen says the job is to build System 1.5: connect fast models to frontier reasoning via software — iamrobotbear · 2026-09-22
- Agent Substrate roadmap: sub-second suspend/resume runtime for dense agent deployments — rakyll · 2026-09-22
- Hot take: Fable beats Astra at coding, but Astra wins at computer use — PratikKadam_ · 2026-09-22
- Building a (deliberately unsafe) restricted shell MCP server for local coding agents — ag789 · 2026-09-22
- Rumor: OpenAI, Anthropic, and Cognition to launch personal agent platforms within a month — altryne · 2026-09-22
- One test decides if you need an AI agent: if you can write the steps down, you don't — Virtual_Hair_1987 · 2026-09-22