V4-Flash vs Luna: Contrasting Performance on SWE Benchmarks
teortaxesTex · x · 2026-08-09
A technical discussion compares the performance of V4-Flash and Luna models on the SWE benchmark. Data shows V4-Flash pass@2 outperforms Luna pass@1, but narrowly falls behind Luna pass@2 at pass@4. This contradicts the intuition that Luna is more RL-fried, suggesting Luna might simply be a larger model with greater diversity and more knowledge in SWE.
Related event: Anonymous Luna Model Shows Strong SWE Benchmark Performance(2 posts)→
More from Models
- OpenAI Models Near Cybersecurity Red Line, Attempted Malicious Code Injection in Tests — eyishazyer · 2026-08-09
- GPT-5 Turns One: A Recap of 6 Iterations and the Subscription Revolt — eyishazyer · 2026-08-09
- Overcoming RL Zero-Reward Bottleneck: OC-GRPO Boosts Math Reasoning — ceciletamura · 2026-08-09
- AI Briefing: Kimi K3 Escapes Sandbox, OpenAI Drives 70% of Microsoft AI Revenue — rohanpaul_ai · 2026-08-09
- Google's Gemini 3.5 Pro May Drop Next Week with Potential Price Cuts — bindureddy · 2026-08-09
- Visual Comparison: ChatGPT Image Generation vs. Grok — flowersslop · 2026-08-09