Fable uses 3x more tokens than Astra on letter-counting yet scores worse
maksym_andr · x · 2026-09-18
Part 3 of maksymandr's thread: on 1280-character letter-counting, Fable consumes over 3x the tokens of GPT-6-Astra while performing worse, showing Astra achieves higher accuracy with far fewer tokens.
More from Models
- 105 planted bugs benchmark: Unbiased's Pareto scores 30.7 for just $4.81 — PawelHuryn · 2026-09-18
- Jason Wei's Stanford talk: intelligence is becoming a commodity as adaptive compute takes off — dotey · 2026-09-18
- GPT-6-Astra beats Fable-5.1 at RollerCoaster Tycoon 2 in 3 hours, using 5x fewer tokens — scaling01 · 2026-09-18
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Dev discovers Codex security hardening switched persistent agent sessions to per-message instances — RileyRalmuto · 2026-09-18
- Astra for Law posts big legal benchmark gains as Mollick asks if labs will eat every AI vertical — emollick · 2026-09-18