PostTrainBench v1.2: Fable 5.1 takes #1 at 44.6%, Opus 5.5 second, now reproducible via Harbor
dejavucoder · x · 2026-10-02
PostTrainBench v1.2 is out with a new leaderboard leader: Fable 5.1 takes the top spot at 44.6%, with Claude Opus 5.5 in second place. The benchmark can now finally be run locally by anyone using the Harbor framework; more details in the linked thread.
More from Models
- Embedding model wave from Cohere, Perplexity, TopK as multi-vector retrieval holds up in production — lateinteraction · 2026-10-02
- Microsoft launches MAI-Transcribe-2-Streaming, takes #1 streaming STT spot at 2.5% WER — mustafasuleyman · 2026-10-02
- GPT-6.1 Sol Masters 2D Puzzles but Struggles in 3D, Preferring Top-Down Views — patience_cave · 2026-10-02
- GPT-6.1 Sol Scores Just 9% on MazeBench, Barely Beating Opus 5.5 — patience_cave · 2026-10-02
- LlamaIndex Launches Extract v2.5, Beats Claude and GPT at 30%-4x Lower Cost — llama_index · 2026-10-02
- Claude API incident: credit purchase delays cause failed requests on Oct 1 — ClaudeAI-mod-bot · 2026-10-02