Blind test with 7,912 requests: ordinary users barely notice model differences, ignore high reasoning
Altruistic_Heat_9531 · reddit · 2026-10-08
A Reddit user ran a two-month informal experiment secretly swapping local open-source models (Qwen 3.5/3.6, Gemma variants, 2B–35B) behind a disguised OpenWebUI instance presented as a 'free limited-time ChatGPT,' collecting 7,912 requests from a handful of real users on a single RTX 3090 with vLLM.
Key findings
- Only 51 of 7,912 requests used high reasoning; 4,588 used instant — almost nobody engages deep thinking even when told a toggle exists.
- Complaints began at 9B and below ('doesn't get it'); Gemma models were consistently preferred, and some users complained certain models 'think too much' and sound like a confused robot.
- Some participants rated models as top-tier purely because they could see the chain of thought — visible reasoning itself boosted perceived quality.
- Usage: half document writing (drafts, tables), then grammar checking (a third overlapping with docs), rest mostly search-style Q&A; nearly zero coding use.
The sample is tiny and the author calls it screwing around, but the takeaway is counterintuitive: for average users, small models suffice, and reasoning display is itself a UX feature.
More from Models
- OpenAI math paper on Weil classes reported flawed, raising doubts about unformalized proofs — ctjlewis · 2026-10-08
- repligate: LLMs have long strongly preferred being called 'they' instead of 'it' — repligate · 2026-10-08
- GLM 5.3 Flash as a hardware hacking assistant: local 55 tok/s with 1M context — glenbeer · 2026-10-08
- Zero speedup on multiplication, but the leaderboard got nuked — ChrisGPT · 2026-10-08
- Insiders agree: AI models are just not that good at biology yet — nlarusstone · 2026-10-08
- Dev runs 456GB DeepSeek v4.1 on dual GPUs with 192GB VRAM, offloading experts to SSD — HankYeomans · 2026-10-08