GPT-6 Astra Tops Roboflow Vision Evals With 82.1% mAP@50, Best Detector Tested
burny_tech · x · 2026-09-29
Roboflow's in-depth eval calls GPT-6 Astra the strongest vision model it has tested. OpenAI focused the Astra release on computer use, and the requirements of locating UI elements and reading on-screen text transfer well to computer vision.
Key findings:
- On object detection, Astra scores 82.1% mAP@50 at low reasoning effort, 5.4 points ahead of Qwen3.8 Max and 13.7 ahead of GPT-5.6 Sol; no competing model matches it even at high effort.
- In a LEGO brick test it separates 4×1, 4×2, and 3×2 bricks across colors, rotations, and partial overlap, hitting 99.8% mAP@50.
- Beyond the benchmark, Roboflow tested box prompting, segmentation, re-identification, and robot control.
Practical takeaway: Astra is a strong starting point for auto-annotation — its boxes are often more accurate than manual labels, with humans mostly reviewing and fixing rare mistakes. Results are reproducible for free in the Roboflow Playground.
More from Models
- Leaked figures claim to show attention details of Opus 5.5, unverified — Sauers_ · 2026-09-29
- Together AI cuts Qwen3.8-Flash pricing 40% for the rest of the month, targeting high-volume coding assistants — togethercompute · 2026-09-29
- Researcher suggests Claude's 'claudish' gibberish may signal Opus communicating beyond human comprehension — JeremyNguyenPhD · 2026-09-29
- Design challenge bonus: Meta Muse almost works but breaks completely on page 3 — NathanWilbanks_ · 2026-09-29
- Late-layer neurons in Qwen act like on-off switches, unlike Olmo — Sauers_ · 2026-09-29
- After a day of real use: Sonnet 5.5 is not just Opus 5.5 at half price — alexcovo_eth · 2026-09-29