Custom eval harness ranks Fable 5 and Opus 4.6 above 10 models
rudrank · x · 2026-07-24
- The author argues that each person should build their own evals and use their own harness to compare models under identical conditions.
- They ran an “article taste” eval across 10 models.
- Fable 5 and Opus 4.6 ranked at the top, with Kimi K3 close behind.
- The post frames eval setup as painful but highly iterative and fun once it is in place.
More from Models
- Kimi K3 tops Frontend Arena and appears to use memorized Unsplash image IDs — BlackHC · 2026-07-24
- Zhipu shows GLM 5.2 on Agent Platform, with an 8×H200 deployment estimate near $84,187 a month — dejanseo · 2026-07-24
- Users Praise OpenAI's New Models: GPT-5.6 Series is Excellent, Sol is a Defining Moment — brianmichel · 2026-07-24
- A user says GPT-4.5 was the best AI writing model and wants it back — niloofar_mire · 2026-07-24
- OpenAI and Anthropic’s live voice models may need Cerebras-level inference — downingARK · 2026-07-24
- Claude Opus 5 appears to be rolling out as “Opus 4.8” for some users — legit_api · 2026-07-24