rasbt shows why final-result benchmarks mislead: Astra vs Qwen in Paint

rasbt · x · 2026-09-15

Sebastian Raschka shares a benchmark-design cautionary tale: asking GPT-5.6 Astra and Qwen3.8 Max to recreate an image in Paint. Astra layered geometric shapes; Qwen drew pixel by pixel—naturally scoring closer to the original. But he argues this says nothing about which model generalizes better or has stronger computer-use or visual capabilities. Key takeaway: benchmarks comparing only final outputs are slippery and can misread strategy differences as capability differences.

Original post →

More from Models

Models channel →