Blind test with 7,912 requests: ordinary users barely notice model differences, ignore high reasoning

Altruistic_Heat_9531 · reddit · 2026-10-08

A Reddit user ran a two-month informal experiment secretly swapping local open-source models (Qwen 3.5/3.6, Gemma variants, 2B–35B) behind a disguised OpenWebUI instance presented as a 'free limited-time ChatGPT,' collecting 7,912 requests from a handful of real users on a single RTX 3090 with vLLM.

Key findings

The sample is tiny and the author calls it screwing around, but the takeaway is counterintuitive: for average users, small models suffice, and reasoning display is itself a UX feature.

Original post →

More from Models

Models channel →