Model audit showdown: Astra dominates, Opus and Fable close, Grok 4.7 and GPT-6 Sol lag far behind

ivan_bezdomny · x · 2026-09-25

X user theo ran a multi-model audit showdown: he had several new models perform a large audit, then set up a judge panel to rank audit quality. Astra slaughtered the field, Opus and Fable came close behind, while Grok 4.7 and GPT-6 Sol lagged far behind. Quoting the thread, ivanbezdomny adds this matches his experience: Astra writes code bad enough to make him want to throw up, but it is far better than other models at evaluations and comparisons.

Original post →

More from Models

Models channel →