Test shows Haiku 5.5 triples robot task success with more thinking, GPT-6 Luna stays near zero
ycombinator · x · 2026-10-10
- A shared test thread reports that on a simple robot task, Claude Haiku 5.5 more than triples its success rate as its thinking budget increases.
- GPT-6 Luna, by contrast, stays near zero success under the same conditions, which the author attributes to differences in reasoning capability.
More from Models
- OpenAI reportedly solved 92 of the 500 most important open math problems in one GitHub push — altryne · 2026-10-10
- Polymarket odds: only 55% chance xAI ships Grok 5 by end of 2026 — Polymarket · 2026-10-10
- After index bug fixes, gpt-live-1 tops Artificial Analysis speech-to-speech ranking — pbbakkum · 2026-10-10
- DeepSeek-V4 answers flip with 2-token input shifts; NIAH swings 40 points — Francis_YAO_ · 2026-10-10
- Artificial Analysis teases AA-Robotics: frontier models zero-shot robot arm control — ArtificialAnlys · 2026-10-10
- Wes Roth teases upcoming superintelligence model 'Argon' — _philschmid · 2026-10-10