Placebo-controlled test: Claude Code ignores installed debugging skill, and forcing it doesn't help
Sufficient-Storage87 · reddit · 2026-10-09
- Does Claude use skills on its own? The author installed diagnosing-bugs from mattpocock/skills and ran 16 headless sessions (10 Sonnet 5.5, 6 Haiku 5.5) on a failing test. Claude never invoked the skill, even when prompts included its own trigger word "diagnose".
- Forced comparison: on a real python-humanize bug (precisedelta rounding), stripped git history, no web tools, 66 hidden tests, Haiku 5.5 in three groups of 5 runs scored: no skill 61.8, a same-length placebo skill 61.8, the real skill 62.6 — essentially identical.
- Claude followed the skill's method (ranked hypotheses, feedback loops) but missed the same 3 edge cases every run.
- Caveats: one bug, one model, headless only; an independent Claude session audited the raw transcripts and confirmed scores recompute exactly.
More from coding & agent
- YC partner shares multi-agent workflow: agents write briefs for each other like middle managers — ycombinator · 2026-10-10
- AI writing skill boosts output to 4 posts a week, human editors become the bottleneck — danshipper · 2026-10-10
- AI agent digs through Azure billing to recover nearly $2,000 in lost credits — pswider · 2026-10-10
- HQ Agent Tool Ships Major Update: Per-App Postgres, Full-Stack Next.js, Interactive Pages — jacob_posel · 2026-10-10
- Agent swarm converts GTA footage to 3D scenes in just a few hours — Daniel_Farinax · 2026-10-10
- OpenAI's Huet notes DevDay Chromatic handheld has Wi-Fi, enabling Codex-powered hacks — romainhuet · 2026-10-10