Chartography Benchmark: Zoom Tool Boosts Claude's Chart Accuracy to 44%
echen · x · 2026-07-28
SurgeAI introduced the Chartography benchmark, designed to evaluate how well models read and interpret professional charts.
In a test of 100 questions based on dense real-world charts, both Fable 5 and Claude 3.5 Sonnet showed significant improvement when equipped with a zoom tool. Fable 5's accuracy jumped from 29% to 73%, while Sonnet's increased from 13% to 44%.
Related event: Claude's Zoom Tool Significantly Boosts Chart Recognition Accuracy(3 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23