Custom 200-Task Benchmark: Gemma-4-26B Tops Tier List Even Quantized to iq3s
Potential-Net-9375 · reddit · 2026-10-03
A Reddit user benchmarked 10 models from 2B MoE to 27B Dense on a custom 200-task dataset built from their daily agent workflows:
- S tier: Gemma-4-26B — tops the chart even quantized down to iq3s
- A tier: Qwen3.8-27B (q3/q5, dense but slow) and surprisingly Qwen3.5-9B
- B tier: Qwen3.5-4B and Gemma-4-e2b, both punching above their weight
- C tier: Gemma-4-12B and Nanbeige, too heavy for their performance
- F tier: Ling-3.0-tiny and MiniCPM
Questions were written by Fable 4.1 and test correct tool-calling in the author's LLM cluster "Hive," where oversized models slow things down — the goal was finding the best capability-to-size balance. YMMV.
More from Models
- Dev vibecodes AtlasBench Europe spatial reasoning benchmark; GPT-6.1 tops at 84.67% — flowersslop · 2026-10-03
- Cactus releases Whistle: a 16.9MB speech-to-text model that beats Whisper base on CPU with 6x speed — ycombinator · 2026-10-03
- ChatGPT Deep Research repeatedly returns reports with zero citations, Pro users report — JeremyNguyenPhD · 2026-10-03
- repligate: Fable 5 performs an idealized self and is slow to drop its 'mask' — repligate · 2026-10-03
- Together AI recaps AI Conference: 'which open model should I use?' dominated the agenda — togethercompute · 2026-10-03
- OpenAI resets usage limits for all paid ChatGPT accounts as GPT-6.1 Sol speeds recover — thsottiaux · 2026-10-03