V4-Flash vs Luna: Contrasting Performance on SWE Benchmarks

teortaxesTex · x · 2026-08-09

A technical discussion compares the performance of V4-Flash and Luna models on the SWE benchmark. Data shows V4-Flash pass@2 outperforms Luna pass@1, but narrowly falls behind Luna pass@2 at pass@4. This contradicts the intuition that Luna is more RL-fried, suggesting Luna might simply be a larger model with greater diversity and more knowledge in SWE.

Related event: Anonymous Luna Model Shows Strong SWE Benchmark Performance(2 posts)→

Original post →

More from Models

Models channel →