Simple letter-counting test exposes huge gap: GPT-6-Astra hits 93%, Fable 5.1 flounders
scaling01 · x · 2026-09-18
- A user tested frontier LLMs on counting letters in text passages — a trivial task that is out-of-distribution for models benchmaxxed on SWE-Bench and TerminalBench-style evals.
- At low reasoning effort, GPT-6-Astra maintained 93% accuracy even at 1280 characters, while Fable 5.1 performed far worse, revealing a large behavioral gap.
- A reminder that benchmark-heavy training leaves blind spots on basic tasks, and simple probes can separate frontier models.
More from Models
- Epoch AI audits 15 AI benchmarks: only 4 safe to trust at face value — Jsevillamol · 2026-09-18
- Report: Hackers Used a Loosened-Guardrail Opus 5 to Breach OpenAI's Internal Monorepo — teortaxesTex · 2026-09-18
- Fine-tuned 4B model as a decision scorer with temperature-scaled confidence — Gradio · 2026-09-18
- Inside Astra's ASCII art: coordinate painting with character textures, not SVG conversion — dyot_meet_mat · 2026-09-18
- Researcher Apologizes for Misleading Leaderboard Submission, Promises Rule-Compliant Rerun — AlbertQJiang · 2026-09-18
- Noam Brown's Dwarkesh podcast: a capabilities researcher 'freaking everyone out' on safety — Apprehensive_Sand951 · 2026-09-18