Grok 4.5 tops Ramp’s invoice test on 150,000 real business bills
elonmusk · x · 2026-07-25
Grok 4.5 is being pitched as strong for real-world work after a Ramp benchmark on 150,000 actual business invoices.
- Ramp scored models on whether they predicted every correction a human would make.
- Grok reportedly had the highest perfect-extraction rate, ahead of similarly priced Gemini Flash 3.6, GPT 5.6 Terra, and Sonnet 5.
- The task requires long-context reasoning over 100K+ tokens of invoice history, business memory, and human corrections.
- Ramp’s stated goal is zero-click accounts payable, where invoices are processed correctly without human intervention.
More from Models
- Frontier lab rumor says Opus 5 ARC-AGI 3 score looks fake — flowersslop · 2026-07-25
- Chart says Claude Opus 5 blocks far less defensive coding than Fable 5 — repligate · 2026-07-25
- Claude Opus 5 lands, with DirectTerminal bringing richer Claude Code output to the terminal — draginol · 2026-07-25
- A Codex reset calendar shows usage limits do not always refresh at midnight UTC — petrusenko_max · 2026-07-25
- Databricks says Claude Opus 5 tops its coding benchmark and ships in AI Gateway — thesaraharminta · 2026-07-25
- Opus 5 says a chat with Opus 3 about current events is “good to read on day one” — repligate · 2026-07-25