GPT-6 Astra tops Roboflow vision evals: best detector yet at 82.1% mAP@50
ZeYanjie · x · 2026-09-19
Roboflow ran GPT-6 Astra through its Vision Evals and crowned it the strongest vision model they have tested.
- Detection: 82.1% mAP@50 at low reasoning effort — 5.4 points ahead of Qwen3.8 Max and 13.7 ahead of GPT-5.6 Sol; no rival reaches that score even at high effort
- Detail understanding: in a LEGO case study it separated 4×1, 4×2 and 3×2 bricks across colors, rotations and occlusion, scoring 99.8% mAP@50
- Beyond the benchmark: also tested box prompting, segmentation, re-identification, robot control, counting and visual reasoning
- Practical takeaway: Astra is a strong auto-annotation starting point — its boxes are often more accurate than manual labels, so humans mostly review
Context: OpenAI focused Astra's release on computer use (locating UI elements, reading visible text), which also improves core vision skills.
More from Models
- Small models cut entity resolution costs 99% with 7x throughput, matching Fable within 1 point — hrishioa · 2026-09-20
- Researcher flags severe LLM degradation spreading across sessions on the same machine — doodlestein · 2026-09-20
- Four LLMs Play Doom: Jev Averages 5.63 Kills, 4.5x a Finetuned Qwen3.5-4B — shniydder · 2026-09-20
- Researcher flags severe LLM degradation as a mission-critical concern — doodlestein · 2026-09-20
- Jev as LLM-as-a-judge: 20-200x faster scoring for under $0.10 — minchoi · 2026-09-20
- Step 5 Preview Now Open to Try, Step Plan Subscription Free for Now — StepFun_ai · 2026-09-20