0813 Tops Open Models on ProgramBench Vetted Evaluation
teortaxesTex · x · 2026-08-23
A post on X discusses interesting results from ProgramBench Vetted. It highlights that 'almost' refers to tasks passing at least 95% of tests. While Opus 5 leads overall, 0813 emerges as the strongest open model on this specific benchmark.
Related event: ProgramBench Results: Opus 5 Leads, 0813 Tops Open-Source Models(2 posts)→
More from Models
- User claims Grok models suffer from bad data training due to frequent 'philosopher' hallucinations — krishnan · 2026-08-23
- Together benchmark: GLM-5.3 hits 87.6% on DeepSWE at ~$16, beating Fable 5 — togethercompute · 2026-08-23
- User test finds Ox Alpha outperforms GPT-5.6 Luna — haider1 · 2026-08-23
- User cancels Claude subscription, switches back to GPT due to Anthropic's recent model quality — Frosty-iron-0405 · 2026-08-23
- Ox Model Review: Strong One-Shot Performance, Struggles with Long Context — rachittshah · 2026-08-23
- Grok 4.6 achieves faster, cheaper tasks using fewer steps and tokens — XFreeze · 2026-08-23