Benchmarking AI assistants: measure time to a checked result, not time to an answer

OriginalHospital · reddit · 2026-09-10

The author proposes defining "finished" before comparing AI assistants: a response appearing on screen is one event, a usable result another.

A fictional example: assistant A drafts in 1 minute but needs 14 minutes of checking; assistant B takes 4 to generate and 3 to review — A looks faster if the stopwatch stops at generation, but B reaches the defined finish sooner. For source-based summaries, "finished" means every claim verified, caveats kept, output fit for the audience.

Tips for fair small trials: same sources and acceptance criteria, count failed attempts, don't give one assistant extra context or retries, blind model names, and rotate review order since learning the source on the first answer biases the second. Report generation and review time plus residual errors separately — don't turn a tiny comparison into a universal ranking.

Original post →

More from coding & agent

coding & agent channel →