PostTrainBench v1.2 ranks Fable 5.1 > Opus 5.5 > GPT-6 Astra, adds Harbor support
maksym_andr · x · 2026-10-02
PostTrainBench v1.2 is out, ranking Fable 5.1 > Opus 5.5 > GPT-6 Astra > Fable 5. The update adds Harbor support for easier runs, removes BFCL, fixes scoring bugs, and now runs the reward-hacking judge and final evaluation multiple times to reduce variance.
Related event: PostTrainBench v1.2 Released: Fable 5.1 Takes Top Spot(2 posts)→
More from Models
- Local AI roundup: 27B reasoning in 5.9GB, phone-class 35B, and dozens more — vramkickedin · 2026-10-02
- Users: new Gemini shines on open-ended overnight research tasks, Flash covers daytime — JMateosGarcia · 2026-10-02
- Rumor: Google has an internal model more powerful than Argon (unverified) — Independent-Wind4462 · 2026-10-02
- Chollet: base LLMs have ~0 fluid intelligence, LRMs saturated ARC 1 in 2025 — fchollet · 2026-10-02
- Users find ChatGPT chat usage secretly drains Codex credits despite docs saying otherwise — jonistaken · 2026-10-02
- Grok 4.7 Reportedly Rolling Out on Grok Web and Mobile as Base Model Across All Modes — testingcatalog · 2026-10-02