Rails agent evals: OpenAI stays ahead, Luna Max the dark horse at 18% completion for just $11
npew · x · 2026-09-23
DHH shared the latest Rails agent evals: OpenAI retains a solid lead despite Opus 5.5 making a good jump. The dark horse is Luna Max, completing 18% of tasks at just $11 — not far off GPT-6 Sol. npew added that Astra is crushing it and called Luna "a little beast." The benchmark measures real task completion on an actual engineering framework, with cost efficiency as a differentiator.
More from coding & agent
- Scoring context with small models: a cheaper alternative to one-shot summarization — marlene_zw · 2026-09-23
- Building a Custom Agent Harness with Pi and Jev, With Interactive Playground — dair_ai · 2026-09-23
- PSAISuite: a PowerShell abstraction layer unifying 15+ GenAI providers — dfinke · 2026-09-23
- Why AI-Edited UI Keeps Breaking, and the Visual Feedback Loop Fix — nikola_mr64990 · 2026-09-23
- Long AI sessions degrade suddenly because context accumulates garbage, not because models got worse — ClickOk5811 · 2026-09-23
- Same model, same prompt, worse API extraction: apps do invisible pre-processing you're skipping — TangeloOk9486 · 2026-09-23