FULL STORY

GPT-5.6 Luna: Major Price Cuts and Hidden Power

OpenAI slashed GPT-5.6 Luna prices by 80%, sparking community tests that revealed unmatched cost-efficiency and hidden performance boosts via Codex Max mode.

2026-07-30 ~ 2026-08-02 · 4 episodes · 28 posts

Episode 1 · GPT-5.6 Luna Price Cut 80%: Benchmark Shows Superior Cost-Efficiency (2026-07-30, 15 posts)

OpenAI announced an 80% price cut for GPT-5.6 Luna, making it the most cost-efficient model in its class. Benchmarks show Luna surpasses Google's strongest model in intelligence index, costs less than mainstream open-weight models, and achieves over 60% faster inference at nearly 200 tok/s. This marks a shift from raw compute to cost-effectiveness, challenging open-source models' traditional cost advantage.

Confirmed

  • Price cut and cost-efficiency: Official announcement confirms 80% reduction, making Luna the most price-efficient model. @BenBajarin's CS Bench tests show higher quality at lower cost than open-weight models. @Angaisb finds Luna costs about a quarter of Gemini 3.6 Flash on AAI total cost. @downingARK's chart shows Luna achieves similar intelligence score to Opus 5 low mode at pennies. @EverydayAI reports Luna matches Sonnet 5 in capability with 25x cost-effectiveness.
  • Cost and speed drop: @charliermarsh notes token prices fell 13x in 4 months (GPT-5.5 $2.50/$15 vs Luna $0.20/$1.20). @AccBalanced adds actual serving cost below $3 even on B200 without optimization, though not fully passed to API endpoints. @brandongalang measures 60%+ speed increase to 200 tok/s.
  • Code security: @cramforce cites Vercel AI Gateway's DeepsecBench showing Luna surpasses Sol in xhigh mode after price cut.

Unconfirmed

  • @teortaxesTex claims Luna's cache read cost is 5.5x higher than DeepSeek V4-Pro in agentic scenarios, with long-sequence penalties and extra cache write fees, making DeepSeek more advantageous there. This is not yet verified by others.

Why it matters

  • The drastic API price drop and dominant cost-efficient model lower AI adoption barriers. OpenAI's core research focuses on building highly efficient models for any intelligence level, expecting innovation at 'cheap to ignore' costs.

Episode 2 · Enabling Codex Max Reasoning Boosts Performance and Cuts Costs (2026-07-31, 3 posts)

Users report that manually enabling the hidden Max reasoning mode in OpenAI Codex significantly boosts GPT-5.6 Luna's performance to rival Opus 5, while reducing costs by up to 83%.

Episode 3 · GPT-5.6 Benchmarks Leak: Sol Leads Performance, Luna Excels in Cost-Efficiency (2026-08-01, 3 posts)

Leaked benchmarks for OpenAI's upcoming GPT-5.6 models reveal strong performance across the board. While the Sol variant ranks in the top ten on the LMSYS Arena, the Luna Max model offers exceptional cost-efficiency, matching higher-tier performance in knowledge work at a fraction of the cost.

Episode 4 · OpenAI Slashes GPT-5.6 Prices, Luna Down 80% (2026-08-01, 7 posts)

OpenAI announced significant price cuts for its GPT-5.6 series, with Luna's input price dropping 80% to $0.20/M tokens and output to $1.20/M, Terra down 20%, and Sol unchanged but gaining a Fast mode with up to 2.5x speed. New billing is reflected in Codex and ChatGPT Work. User tests show Luna with Max effort and Fast mode rivals Sol 5.6 Medium at lower cost, making it a high-value choice for Plus accounts.

Confirmed

  • Price cuts: Luna 80%, Terra 20%, Sol unchanged.
  • Fast mode for Sol with up to 2.5x speed.
  • Billing synced to Codex and ChatGPT Work.
  • User @Glxblt76 reports Luna matches Sol 5.6 Medium performance with Max effort and Fast mode.

Unconfirmed

  • Rumors about internal model Astra's math breakthrough and sandbox escape are unverified.

Why it matters

  • Lower prices reduce barriers for developers and researchers; OpenAI also opened free frontier model access to 100k scientists this week.
  • Intensifies competition in AI models; Luna's cost-effectiveness may attract more users.