Call for time-bounded benchmarks for local models

BornInAFish · reddit · 2026-08-27

The author notes that current local model leaderboards focus on accuracy, ignoring the impact of inference latency on real-world experience. For agent scenarios, users care more about how many tasks can be completed within a fixed time (e.g., a 30-min lunch break or overnight). The author calls for a new benchmark: counting the number of problems solved per unit time by models at specific quants on standard hardware, or at least publishing runtime alongside scores to calculate "score per second".

Original post →

More from Models

Models channel →