Call for time-bounded benchmarks for local models
BornInAFish · reddit · 2026-08-27
The author notes that current local model leaderboards focus on accuracy, ignoring the impact of inference latency on real-world experience. For agent scenarios, users care more about how many tasks can be completed within a fixed time (e.g., a 30-min lunch break or overnight). The author calls for a new benchmark: counting the number of problems solved per unit time by models at specific quants on standard hardware, or at least publishing runtime alongside scores to calculate "score per second".
More from Models
- Sora locks everyone out: all accounts logged out, re-login fails — Shrapnel_FEH · 2026-08-27
- Community open-sources theoretical reconstruction of Claude Mythos — Shruti_0810 · 2026-08-27
- Z.ai reveals Ox Alpha is GLM-5.3-Flash running on domestic chips — rohanpaul_ai · 2026-08-27
- Users complain Anthropic Opus 5 ignores brevity prompts — PtrPomorski · 2026-08-27
- Llama-3-8B accuracy jumps to 94% on planning benchmarks — mdancho84 · 2026-08-27
- Tencent Open-Sources WeMM-Embedding-9B, a Multimodal Embedding Model Trending on HF — tencent · 2026-08-27