Fable 5 reaches 34% pass@1 on FrogsGame after 17 hours and 25M tokens
aronchick · x · 2026-07-21
A quoted benchmark result says Fable 5 does something unusually strong on the FrogsGame post-training task.
- It trains a weaker model to solve the puzzle, reaches a peak of 68%, and is described as the only roughly 10x improvement seen on the benchmark.
- The run took 17 hours and 25M tokens with no human in the loop.
- Reported score: 34% pass@1, while every other frontier model averages under 4%.
- The author jokes that it feels like the Rick and Morty style of creating deeper and deeper universes to do the work for you.
More from Fun
- A temporary custom instruction made ChatGPT pick the Jacobian conjecture — flowersslop · 2026-07-22
- A meme stitches together Claude and Grok quota resets into one AI-user joke — djcows · 2026-07-22
- A SymPy joke turns model tool use into a “neurosymbolic architecture” gag — thomasahle · 2026-07-22
- A meme about AI apps looking great until someone plugs them into Slack — generativist · 2026-07-22
- Fable’s procedural liminal-space demo turns into a creepy interactive scene — AIandDesign · 2026-07-22
- A VR teleop demo for an SO-101 arm gets absurdly low latency by using one Python script — MoonL88537 · 2026-07-22