Lobehub bench shows high reasoning beats max, blogger slams vendor marketing spin

karminski3 · x · 2026-09-15

Citing Lobehub's internal benchmark (his own tests agree, and even medium beat some quantized small models), karminski3 argues that the vendor marketed max as strongest, then reframed high as the right setting when max underperformed—dressing up weak post-training as "best practices" (use high, Linux bash, patches, watch the CoT). He quips Starbucks wouldn't dare claim a medium has 250ml more than a large.

Original post →

More from Models

Models channel →