GPT-6 Astra (high) hits 92.9% on WeirdML, matching Fable 5.1 (max)
zainhas · x · 2026-09-05
A WeirdML benchmark run shows GPT-6 Astra at the high setting scoring 92.9%, tying Fable 5.1 (max) for the top score and setting new individual highs on 6 of 17 tasks. Astra also writes notably less code than Sol and recent GPT models, though still more than Claude. Commenters suspect higher compute tiers (xhigh/max) won't move the needle much further, with fuller analysis to come.
More from Models
- Qwen-Drive-1.0-4B: Alibaba's compact 4B vision-language model for autonomous driving — solyarisoftware · 2026-09-06
- Shader eval puts Astra below GLM 5.3 flash on 3D spatial reasoning — teortaxesTex · 2026-09-06
- Claimed GPT-6 Astra + H3 Max combo builds an insanely fast realtime Blender renderer — jfischoff · 2026-09-06
- Users report ChatGPT has become 'handicapped' lately, forgetting context and demanding files — lonelyroom-eklaghor · 2026-09-06
- Early Astra hands-on: code quality and data analysis feel incremental, says developer — xeophon · 2026-09-06
- Providers wage price war over serving DeepSeek-v4-flash, user burns massive tokens for a few dollars — MaziyarPanahi · 2026-09-06