GLM 5.2 Ties with Opus in Enterprise Code Eval

rohanpaul_ai · x · 2026-07-11

Databricks conducted evaluations using real internal PRs, tests, and million-line codebases rather than relying solely on public benchmarks.

Results show:

Related event: Databricks' Internal Coding Benchmark: Harness Design Matters More Than Model Price(16 posts)→

Original post →

More from coding & agent

coding & agent channel →