Personalized benchmark noahbench re-released, crowning Fable 5.1, Grok 4.7 and Opus 5.5
sarahcat21 · x · 2026-09-29
- Developer Noah re-released noahbench™, a deliberately personalized benchmark that grades model output quality on real-world tasks he solves daily.
- His top 3 models: Fable 5.1, Grok 4.7, and Opus 5.5.
- Quoting it, Sarah Cat argues the industry is entering an era of customization — per-customer and personalized models — which will require new modeling approaches plus tools to scale benchmarks/evaluation and to monitor, update, and orchestrate dozens to thousands of models.
More from Models
- Community finds ~17% of MiMo-V2.6-Pro's ~1T params can be pruned with no KLD change — QuixiAI · 2026-09-29
- Image JevBench v0.1.4 live: Imajev-4B holds #1, Wity-1 debuts at #2 with 74.36 — airesearch12 · 2026-09-29
- Opus 5.5 Generates a Gorgeous Scrollable 3D Explainer on Chipmaking — threepointone · 2026-09-29
- Why steering vectors work on complex models: a new post offers three intuitions — gleech · 2026-09-29
- OpenAI Pushes Pre-Event Update, Wiping Users' Custom Sprite Sheets — whurley · 2026-09-29
- UsageBench launches to track AI subscription usage cuts, starting with Claude Max 20x and ChatGPT Pro — alejandroll10 · 2026-09-29