Comparing 33 Qwen Models: Over 1,100 One-Shot Outputs Analyzed
kms_dev · reddit · 2026-08-03
A developer spent the weekend testing the cheapest models on OpenRouter, conducting a comprehensive comparison across 33 different Qwen models.
- Test Scale: Used 35 distinct prompts for one-shot generation, successfully collecting 1,109 valid outputs.
- Covered Versions: Includes everything from Qwen 2.5 up to the newest Qwen 3.7, covering various sub-versions like Coder and VL (Vision-Language), as well as different parameter sizes.
- Accessibility: All results are aggregated on the OneshotLM website, allowing users to visually compare the actual generation quality of different parameter sizes and model versions under identical prompts.
More from Models
- Users Complain Claude Opus 5 is Wordy and Judgmental vs Opus 4.6 — JeremyNguyenPhD · 2026-08-03
- OpenAI Models Consume 10x Tokens for the Same Prompt, Dev Reports — michael_g_williams · 2026-08-03
- Deep Dive into Kimi K3: Architecture and Training of the 2.78T Model — imrancoder · 2026-08-03
- DeepSeek V4 Flash Performance Varies Wildly Across Coding Agents — PMinervini · 2026-08-03
- Rumor: Zhipu's GLM 5.5 and DeepSeek Pro Set for August Release — bindureddy · 2026-08-03
- MiniMax H3 Multimodal Video Model Coming to ComfyUI — NerdyRodent · 2026-08-03