Verifiable tasks are solved: Opus 5.5 hits 100% on accounting work

CurieuxExplorer · x · 2026-10-02

A cited thread shows Opus 5.5 completing manual accounting tasks instantly at 100% accuracy where humans take far longer and err often — a stark reversal from GPT-4o underperforming average accountants two years ago. The takeaway: if a task is checkable, frontier models now beat specialists; the open question is everything that isn't verifiable.

Original post →

More from AGI Musings

AGI Musings channel →