Dev calls for speed/productivity benchmarks: frontier LLMs still ship with no sense of time
gandamu_ml · x · 2026-09-17
gandamuml argues it's wild that frontier LLMs still ship with no sense of time. Instead of task-completion metrics and pelicans-on-bikes style benchmarks, he suggests scoring models on speed and productivity.
More from Models
- LAB legal benchmark flaw: case docs leak planted issues directly to models — andersonbcdefg · 2026-09-17
- MiMo eval chart shows judge and probe disagree 60% of the time, sparking reward-hacking concerns — andrew_n_carr · 2026-09-17
- Fans mourn the end of Codex lead's regular daily reset cadence — kimmonismus · 2026-09-17
- Users report Google Astra feels noticeably degraded over past two days — pwlot · 2026-09-17
- OpenAI burns 20% as much compute on monitoring as the model itself, SemiAnalysis says — kevinnbass · 2026-09-17
- Google releases Gemma 3n: 2GB RAM multimodal model, first sub-10B to top 1300 on LMArena — joemeno · 2026-09-17