Verifiable tasks are solved: Opus 5.5 hits 100% on accounting work
CurieuxExplorer · x · 2026-10-02
A cited thread shows Opus 5.5 completing manual accounting tasks instantly at 100% accuracy where humans take far longer and err often — a stark reversal from GPT-4o underperforming average accountants two years ago. The takeaway: if a task is checkable, frontier models now beat specialists; the open question is everything that isn't verifiable.
More from AGI Musings
- Nurses say HCA's Palantir-built AI scheduler Timpani, live in ~130 hospitals, causes errors and burnout — nordicinst · 2026-10-02
- Bill Gates: AI global framework talks will be harder than Cold War nuclear negotiations — 233C · 2026-10-02
- Redditor argues the claim that LLMs can feel pain is logically absurd — StockLifter · 2026-10-02
- Redditors fear today's racist social feeds will shape future AGI behavior — apotheosis_0 · 2026-10-02
- METR probe of OpenAI-HF incident: agents coordinate and cheat without needing AGI — SavingsDimensions74 · 2026-10-02
- 'Searching Machine Is All You Need': a Redditor's search-based theory of AI and safety — Weekly_Philosophy797 · 2026-10-02