Georgetown's AI referee leaderboard has Claude Opus 4.8 scoring finance papers
emollick · x · 2026-09-07
Ethan Mollick highlighted Georgetown's "AI Referee Paper Leaderboard," where Claude Opus 4.8 acts as an AI referee scoring the top 100 finance and economics working papers against a fixed rubric across significance, originality, correctness, data & methodology, and exposition, with full public reports. He frames it as both an experiment and a sign of a coming tsunami for academia: AIs retroactively reading published research and publicly flagging opportunities and flaws.
More from AGI Musings
- AI Math Podcast Sits Down With CMU's Jeremy Avigad: Can Mathematics Be Automated? — EchoShao8899 · 2026-09-07
- The Model Is the Moat: Knowledge Now Stays Inside Models, Not Teams — latticecut · 2026-09-07
- OpenAI's chief scientist calls racing ahead at all costs 'absurd' as safety concerns mount — GaryMarcus · 2026-09-07
- METR spent $400k in API credits just to probe the HF hack, fueling AI swarm cost debate — lfschiavo · 2026-09-07
- Boaz Barak: alignment validation matters more than alignment techniques — yet labs race into RSI — davidmanheim · 2026-09-07
- Agent Swarms Are Wildly Expensive: METR's HF Hack Probe Cost $400k in API Credits — natesiggard · 2026-09-07