PostTrainBench v1.2: Fable 5.1 tops leaderboard at 44.6%, cloud GPU support
karinanguyen · x · 2026-10-02
PostTrainBench v1.2 adds cloud GPU runs via a new Harbor + Modal adapter, refreshes the leaderboard (Fable 5.1 #1 at 44.6%, Opus 5.5 at 43.8%, GPT-6 (Astra) at 41.9%), and fixes evaluation: BFCL removed, HumanEval and remote-code scoring fixed, multi-seed averaging, and majority-vote contamination checks.
Related event: PostTrainBench v1.2 Released: Fable 5.1 Tops Leaderboard(3 posts)→
More from Models
- Nat Eliason suspects @bot now runs on 4.7, says he reaches for Claude and Codex less — nateliason · 2026-10-02
- Users report queries being routed to Fable 5.5, with SVG tests said to far beat 5.1 — kimmonismus · 2026-10-02
- 180B Qwen model runs on one DGX Spark: 2.39-bit quant keeps 95.5% of BF16 scores — TheZachMueller · 2026-10-02
- Grok 4.7 reportedly live across all modes in the Grok app — XFreeze · 2026-10-02
- Claude's cloud sessions don't cost extra — bonus credits are cloud-only tokens — stablequan · 2026-10-02
- One 'please continue' Prompt Burned a 5-Hour Usage Cap in 6.5 Minutes — thawingfrog · 2026-10-02