Custom 200-Task Benchmark: Gemma-4-26B Tops Tier List Even Quantized to iq3s

Potential-Net-9375 · reddit · 2026-10-03

A Reddit user benchmarked 10 models from 2B MoE to 27B Dense on a custom 200-task dataset built from their daily agent workflows:

Questions were written by Fable 4.1 and test correct tool-calling in the author's LLM cluster "Hive," where oversized models slow things down — the goal was finding the best capability-to-size balance. YMMV.

Original post →

More from Models

Models channel →