GLM 5.2 Ties with Opus in Enterprise Code Eval
rohanpaul_ai · x · 2026-07-11
Databricks conducted evaluations using real internal PRs, tests, and million-line codebases rather than relying solely on public benchmarks.
Results show:
- GLM 5.2 entered Databricks' top capability tier, statistically tying in quality with Claude Opus 4.8.
- The quality-cost Pareto frontier for coding tasks features OpenAI, Anthropic, and open-source models, indicating that achieving cutting-edge performance currently requires a multi-model combination.
- For the same model, simply changing the harness/orchestration method can yield about a 2x cost saving.
- Databricks believes open-source models can now handle the most difficult tasks, utilizing Omnigent upfront for multi-harness and multi-model routing.
More from coding & agent
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11