Blind multi-model review: Opus 5 at Medium effort beats Opus 4.8 High
OHOLshoukanjuu · reddit · 2026-08-22
Amid widespread criticism of Opus 5, a non-coder user ran their Cross-Instance Review pipeline: multiple models independently draft from the same brief, then peer-review and vote via STAR (Score Then Automatic Runoff) to pick the base draft, used here to build v2 of their Research Synthesis skill.
Scores (max 30): Opus 5 Medium took first with a unanimous 30, beating Opus 4.8 High (25) 6–1 in the runoff; Fable 5 Medium (23) beat Fable 5 High (20); Opus 4.6 High came last at 8, with multiple seats citing concrete missing requirements.
Two counterintuitive findings: Fable's effort profile runs inverse (Medium beats High, a third independent observation — no case for paying High on Fable for build work); and Opus 4.6's weak precision instruction-following argues for dropping it from build seats. The author also found Sonnet 5 at Low effort beats Opus 5 for synthesizing multiple research reports, thanks to lower variance and cost.
More from Models
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Qwen 27B 3.8 low quantization tested: Q3 XXS works well locally — jeremyckahn · 2026-08-24
- Users notice significant quality shift in GPT-5.6 output — haider1 · 2026-08-24
- Ramp Stats: Anthropic Opus 4.8 and Sonnet 4.6 Lead Usage — vista8 · 2026-08-24
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24