Muse Spark 1.3 ties Fable 5.1 in 105-bug repo test, Meta joins frontier
PawelHuryn · x · 2026-09-13
The author ran an original harness with raw API on 2 real repos containing 105 planted bugs, asking models to find and fix them. Results: Muse Spark 1.3 (max): 33, Fable 5.1 (high): 33, Grok 4.6 (xhigh): 27, Opus 5 (max): 27, Muse Spark 1.3 (high): 19. Takeaway: "Meta joined the frontier." More effort-level results dropping every 1.5h in the thread.
Related event: Muse Spark 1.3 Ties for First in Real-World Bug-Fixing Benchmark(2 posts)→
More from coding & agent
- Agent swarm leaderboard 'turkish-delight' dethroned; anyone can submit an agent in 2 minutes — mervenoyann · 2026-09-13
- Dev builds 'recursive self-improvement': scheduled automations with MCP access that fix codebases — DanielLockyer · 2026-09-13
- WTMemory: a 794KB Mac utility that hunts AI model weights and dev servers eating your RAM — saibharadwaj · 2026-09-13
- "AI won't kill me—but I'll be buried under an ever-growing review pile" — tokenbender · 2026-09-13
- ProTip: lock down code sections in agents.md to stop agents from breaking them — cyrus_zei · 2026-09-13
- Open-source thesys-core highlights exact paragraphs behind AI answers in 100+ page PDFs — Flat-Phone-1596 · 2026-09-13