GPT-5.5 fully verifies only 27/43 repos; 10 unsolved across all 8 configs
dawnsongtweets · x · 2026-08-23
Vero is challenging for frontier models: within a 90-minute budget, the best configuration, GPT-5.5 at xhigh reasoning effort, fully verifies only 27 of 43 repositories in code-and-proof mode and 25 in proof-only; Claude Opus 4.8 reaches just 8 and 10 respectively. Ten repositories remain unsolved across all 8 configurations tested.
More from Models
- Claude Pricing Transparency Criticized: Why Silicon Valley Is Hated — StewartalsopIII · 2026-08-23
- Pixel32Bench compares language models via 32x32 pixel generation — TheMoonMidas · 2026-08-23
- Benchmarking LLMs by asking them to draw Mario on a 32x32 grid — TheMoonMidas · 2026-08-23
- RTX 5090 runs Qwen3.8-27B at 262K context — Fz1zz · 2026-08-23
- Flashback: GPT-4 cost $60/M output tokens with 8K context three years ago — gajesh · 2026-08-23
- Open Weights vs Frontier: Just a 3-Point Gap but 1/3 the Cost — MicahBerkley · 2026-08-23