Databricks evals: Opus 5.5 cuts task costs 20%, GPT-6 Luna is 20x cheaper
pwendell · x · 2026-09-29
Databricks' Peter Wendell shared findings from online workload analysis of 2,400 engineers plus offline evals on recent model releases:
- Opus 5.5 and GPT-6 Luna meaningfully expand the cost/quality Pareto frontier.
- Opus 5.5 is now the highest-quality mid-tier model, beating all prior Opus versions as well as GPT-6 Sol and GPT-5.6 Sol.
- Against Opus 4.8 (the prior cheapest Opus; 5.0 was a dud), Opus 5.5 consistently reduces same-task costs by 20%, and Databricks is now encouraging it as the everyday default for coding.
- GPT-6 Luna is extremely cheap: at least 20x cheaper per task than Opus 5.5 in every offline benchmark and observed online use.
- Luna is surprisingly capable — on one of Databricks' hardest eval suites it roughly matches Opus 4.6 while costing 99.3% less per task (preliminary finding).
Their production setup routes workloads via Unity Gateway, with harnesses including Omingent, Claude Code, Codex, and Cursor.
Related event: Databricks Tests Show GPT-6 Luna Costs 20x Less Than Opus 5.5(2 posts)→
More from coding & agent
- New GPU Prices API Tracks Real-Time H100/B200 Rental Rates via REST or MCP — virattt · 2026-09-29
- FastMCP Podcast: LangChain on the New MCP Spec, Agent Evals, and Decision Models — Hacubu · 2026-09-29
- Tested 6 models: WebMCP is optional for agents but consistently cuts steps and speeds them up — rseroter · 2026-09-29
- AgentPaySec tested a payment-capable AI agent: 5 of 16 adversarial security tests exploited — Miserable_Gas_1527 · 2026-09-29
- Two reusable prompts: p5.js story trailers and news-meme YouTube explainers — dotey · 2026-09-29
- A single JavaScript prompt generates a typography-driven minimalist video — dotey · 2026-09-29