KateBench details: 86% automation rate reveals consistency issues
every · x · 2026-08-28
The Every team shared insights on KateBench, their internal AI copy-editing tool. It automates the editor-in-chief's skills by leaving tracked suggestions in Google Docs. Recent runs show an 85-90% acceptance rate. However, the maintaining engineer discovered that the tool stops filing suggestions after 40 and produces non-deterministic edits across reads. This highlights engineering challenges regarding consistency and stability in AI products.
More from coding & agent
- AI agents make sim-to-real training more accessible — chris_j_paxton · 2026-08-28
- Eval anti-cheat idea: serve models a stale HF cache from before grader fixes — willcb · 2026-08-28
- Agent saves hours by automating video migration via browser control — evielync · 2026-08-28
- Local Coding on 5090: Copilot Auto Beats Qwen3 in Task — Efficient_Raisin7645 · 2026-08-28
- Airbnb to share its prototype-to-production LangChain/LangGraph agent stack at Interrupt NYC — LangChain · 2026-08-28
- Study of 400k Claude Code sessions: Domain expertise beats coding skills — alex_verem · 2026-08-28