GPT-6 Luna/Sol on a 32-CVE cyber benchmark: Sol hits 68.8%, cost per CVE drops sharply
rickasaurus · x · 2026-09-23
Quoting @pilvar222's thread: on a 32-CVE cyber benchmark, GPT-6 Luna rediscovered 53.1% of vulnerabilities and Sol 68.8% — neither beat GPT-5.6 variants on recall, but both got much cheaper per CVE found: Luna $3.43 → $2.01, Sol $56.88 → $34.18. Thread part 1/3.
More from Models
- Opus 5.5 stop-motion demos look alike, raising questions on creative heterogeneity — dylfreed · 2026-09-23
- Rumor: Opus 5.5 Said to Be Cheaper and Faster Than Astra — BLUECOW009 · 2026-09-23
- The industry dropped the ball on strong, tool-calling, non-reasoning small LLMs — vboykis · 2026-09-23
- Asked AI to Review a Known Concurrency Bug, It Came Back Clean — cto_junior · 2026-09-23
- Developer Stress-Tests Opus 5.5 With Dozens of Tasks Before It Gets Nerfed — remilouf · 2026-09-23
- Jev API explodes at $0.042/M tokens: a hands-on checklist from desktop agents to drone control — blaizedsouza · 2026-09-23