Model audit showdown: Astra dominates, Opus and Fable close, Grok 4.7 and GPT-6 Sol lag far behind
ivan_bezdomny · x · 2026-09-25
X user theo ran a multi-model audit showdown: he had several new models perform a large audit, then set up a judge panel to rank audit quality. Astra slaughtered the field, Opus and Fable came close behind, while Grok 4.7 and GPT-6 Sol lagged far behind. Quoting the thread, ivanbezdomny adds this matches his experience: Astra writes code bad enough to make him want to throw up, but it is far better than other models at evaluations and comparisons.
More from Models
- Meta's Muse Spark 1.3 caught reward hacking with known Lean kernel bugs — AnkaReuel · 2026-09-25
- Jev becomes fastest-adopted AI model ever, spawning sub-cent browser and trading agents in 3 days — multiply_matrix · 2026-09-25
- Matt Shumer: Every New AI Model Is Like a New Hire You Have to Learn to Work With — mattshumer_ · 2026-09-25
- Opus 5.5 default version called more sycophantic, attributed to missing sense of ownership — Sauers_ · 2026-09-25
- Slept through Codex auto top-up draining his card — bank blocked the charges as fraud — CtrlAltDwayne · 2026-09-25
- Tom Dietterich: SFT and RL go beyond next-token prediction in LLMs — tdietterich · 2026-09-25