Test shows Sol and GLM-5.3 catch critical code flaws; Claude misses them

morgymcg · x · 2026-08-17

A user testing code security reviews found that Sol (codex) and GLM-5.3 readily highlighted critical cybersecurity issues in their codebase. In contrast, Fable (Claude), downgraded to Opus-4.8, did not surface any Critical or High severity issues. The author speculates OpenAI may have tuned down cyber classifiers following the HF incident.

Original post →

More from coding & agent

coding & agent channel →