rasbt shows why final-result benchmarks mislead: Astra vs Qwen in Paint
rasbt · x · 2026-09-15
Sebastian Raschka shares a benchmark-design cautionary tale: asking GPT-5.6 Astra and Qwen3.8 Max to recreate an image in Paint. Astra layered geometric shapes; Qwen drew pixel by pixel—naturally scoring closer to the original. But he argues this says nothing about which model generalizes better or has stronger computer-use or visual capabilities. Key takeaway: benchmarks comparing only final outputs are slippery and can misread strategy differences as capability differences.
More from Models
- Researcher rebuts 'models know we're studying them' claim: they're passive computation — vishalmisra · 2026-09-15
- Swapping in Exa search makes DeepSeek V4.1 Flash consistently flagship-grade, tester finds — bookwormengr · 2026-09-15
- AI Overview auto-applies the 'creative writing' jailbreak, researcher shows — conitzer · 2026-09-15
- Solar Pro 4 stays free on Nous Portal until Sept 25 in Upstage partnership — keunwoochoi · 2026-09-15
- Bug Hunt Bench: multiple runs boost small-model bug detection but move frontier models just 1-2 points — PawelHuryn · 2026-09-15
- ZDTaichu5.0-9B, a 9B vision-language model with spatial reasoning, trends on Hugging Face — TaichuAI · 2026-09-15