"Everyday task" benchmarks are meaningless — can we get real data on what models actually build?
Benhamish-WH-Allen · reddit · 2026-09-30
A Reddit user criticizes vendors' meaningless model capability descriptions like "everyday task": demos split between Luna minimum building a ComfyUI wrapper and Astra Ultra computing to the 15th decimal place, with nothing in between. They call for aggregated metrics — what's being built with which model, how long it took, token cost — to replace marketing claims with comparable real-world data.
More from Models
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30
- DeepSeek is giving users 6 yuan in free API credits via its harness — teortaxesTex · 2026-09-30
- User claims OpenAI bots autonomously scan your Gmail after connecting and keep the data — alexcovo_eth · 2026-09-30
- Rumor: DeepSeek's rumored single-GPU model may have been trained on Ascend — teortaxesTex · 2026-09-30
- DeepSeek opens community feedback channel: flip a toggle in the harness to help improve its models — teortaxesTex · 2026-09-30
- You can drop the vision encoder once pretraining compute exceeds 1e22 FLOPs — heghbalz · 2026-09-30