105 real bugs tested: Muse Spark 1.3 ties Fable 5.1 at 33 as Meta joins the frontier
PawelHuryn · x · 2026-09-13
- PawelHuryn ran a private benchmark: 105 planted bugs across two real repos, original harness and API. Scores: Muse Spark 1.3 (max) 33, Fable 5.1 (high) 33, Grok 4.6 (xhigh) 27, Opus 5 (max) 27, Muse Spark 1.3 (high) 19. Verdict: "Meta joined the frontier."
- Methodology notes: bugs are real issues frontier models genuinely struggled with in early 2026; blind cross-family model judges, calibration confirmed; unplanted bugs don't count; answer key stays private, only anonymized logs published.
- Quirk: OpenAI models "bugmaxx"—reporting many real but irrelevant issues, and fixing them often complicates the solution.
- More effort levels dropping in the thread; Muse Spark 1.3 was the most requested model from the previous post.
Related event: Muse Spark 1.3 Ties for First in 105-Real-Bug Benchmark(3 posts)→
More from coding & agent
- Building a hallucination detector taught us "hallucination" isn't one category: partition, don't boolean — Top-Shopping539 · 2026-09-13
- Fulmar open-sources a native macOS wrapper for the DeepSeek agent harness — absolutefunnyguy · 2026-09-13
- AFK Pilot Lets You Steer Grok, Codex and Claude Code from Any Browser — PawelHuryn · 2026-09-13
- AI Engineer Learning Path: Build First, Then Go Deep Where You Get Stuck — ashishllm · 2026-09-13
- Forma: Open-Source AI Tool Generates DevExpress Report Layouts from Images and PDFs — waqarsyd · 2026-09-13
- 15 days, 50 capabilities: dev turns Hermes Agent into a personal operating layer — Teknium · 2026-09-13