GLM-5.3 nearly matches Claude at exploiting bugs; $20 in tokens found a Chrome flaw

DeepLearningAI · x · 2026-10-04

In this week's letter, Andrew Ng analyzes Anthropic's evaluation of open-weight GLM-5.3's cyber capabilities: on a subset of ExploitBench exploit tasks, GLM-5.3 solved 12% versus closed-weight Claude Mythos's 14% at comparable token cost (GLM's own full-benchmark numbers show a wider gap: 54.4% vs 78.0%).

Related event: Open-source GLM-5.3 nearly matches Claude in exploit capability(2 posts)→

Original post →

More from Models

Models channel →