OpenAI's Logan Kilpatrick: AI product teams should spend >25% of time on benchmarks
OfficialLoganK · x · 2026-09-22
Logan Kilpatrick of OpenAI advises teams building AI products to spend over 25% of their time creating benchmarks and lobbying model labs to care about them. Building custom evals and getting labs to optimize for them, he argues, is the easiest way to accelerate a company's progress.
More from coding & agent
- GitHub Copilot app adds Sentry canvas: from crash report to fix PR in one app — mariorod1 · 2026-09-22
- Giving an autonomous agent homeostatic sleep: fatigue from real token cost, plus REM dreaming — Dzikula · 2026-09-22
- Dev on AI coding's biggest pain point: models are still too stupid, slow and expensive — remilouf · 2026-09-22
- MCP Is Speedrunning the Web 2.0 Story: Platforms Will Tighten Integration Gates — dbreunig · 2026-09-22
- Firebase Tutorial: Build Apps That Talk, Laugh and Whisper with Gemini TTS — jggomezt · 2026-09-22
- Debate: RL Environment Synthesis Is Valuable but Not Novel — stochasticchasm · 2026-09-22