105 planted bugs in real repos: Meta's Muse Spark 1.3 ties Fable 5.1, joins the frontier
PawelHuryn · x · 2026-09-14
Pawel Huryn ran an original harness over two real code repos with 105 planted bugs, asking frontier models to find and fix what they could.
Results:
- Muse Spark 1.3 (max): 33
- Fable 5.1 (high): 33
- Grok 4.6 (xhigh): 27
- Opus 5 (max): 27
- Muse Spark 1.3 (high): 19
The takeaway: Meta has joined the frontier. Scale AI CEO Alexandr Wang amplified the result, saying muse code + muse spark 1.3 max is "really good at real-world software engineering." More effort levels are being dropped in the thread every 1.5 hours.
Related event: Bug Hunt Bench: Muse Spark 1.3 Ties Fable 5.1 for the Top Spot(6 posts)→
More from coding & agent
- Open-source agent harness runs 63-hour autonomous Riemann hypothesis attempt on a single RTX 3090 — GuiltyBookkeeper4849 · 2026-09-14
- Agentic coding brings back human connection at this game studio, founder says — TAbrodi · 2026-09-14
- Clarifying the AI agent messaging drama: it just couldn't access a link — Kyrannio · 2026-09-14
- Cloud queue + local agent hits a wall: laptop off, all scheduled jobs die — PriorElephant9 · 2026-09-14
- Freelancer's Gmail task-extraction script worked — until the manual 'last mile' killed it — Thefounderman1 · 2026-09-14
- smolvm v1.16 adds incremental checkpoints: git-like time travel for agents, 10x faster branching — LoganGrasby · 2026-09-14