Rails benchmark: Grok 4.6 hits 84% accuracy near Opus at 60% lower cost
AccBalanced · x · 2026-08-19
dhh shared the Rails coding benchmark results, marveling at how fast Grok is closing on the frontier and asking "what happened to Google?"
Of the 4 new models benchmarked, Grok 4.6 performed best: 4th in accuracy at 84% (vs. current best Claude Opus 5 at 92%) while costing 60% less; it also scored 33.3% recall on Rails APIs, third on the leaderboard. Claude Opus 4.8 ranked second for speed at 3m36s, just 7 seconds behind the fastest (Luna). Gemini Flash 3.7 performed mid-to-low across the board but remains one of the cheaper options.
More from Models
- Researcher: GPT 5.6 Sol Ultra Beats Pro for Long-Horizon Hard Problems — arankomatsuzaki · 2026-08-24
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Qwen 27B 3.8 low quantization tested: Q3 XXS works well locally — jeremyckahn · 2026-08-24
- Users notice significant quality shift in GPT-5.6 output — haider1 · 2026-08-24
- Ramp Stats: Anthropic Opus 4.8 and Sonnet 4.6 Lead Usage — vista8 · 2026-08-24
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24