Meta's Muse Spark 1.3 underwhelms on cybersecurity benchmark: 19/32 CVEs vs Grok 4.6's 23.3
teortaxesTex · x · 2026-09-03
Third-party testers ran Meta's newly launched Muse Spark 1.3 through their cybersecurity benchmark:
- At pass@1, the model rediscovered an average of 19/32 CVEs, trailing Grok 4.6's 23.3/32
- Pooling three runs (pass@3) yielded 24/32, still behind DeepSeek V4 Pro at 28/32
- Pricing is competitive with frontier models but hard to justify versus some Chinese models
- Cached input token rate was just 45%, lower than usual; results were also normalized assuming 90%
Verdict: Meta is slowly catching up — decent at coding, not yet there for security. Zuckerberg framed the launch as Meta's biggest coding/agentic jump, with open weights and a mysterious 🍉 release teased.
More from Models
- China's token plans priced far above overseas LLM subscriptions, per finance mag piece — yihui_indie · 2026-09-03
- Hands-On Ranking: Fable 5.1 Beats Claude Opus on RL Environment Creation — HarveenChadha · 2026-09-03
- Philosopher says Fable 5.1 found Arrhenius's favored population ethics impossibility theorem is false — willmacaskill · 2026-09-03
- Reddit users report mass OpenAI account bans hitting paying customers — Existing-Slide7395 · 2026-09-03
- Gemini Speech-to-Text Shines in Real-World Test—Dev Open-Sources Windows Tool AivoRelay — lvvy · 2026-09-03
- Meta's new Muse model suddenly outputs Chinese characters, sparking distillation rumors — zsakib_ · 2026-09-03