Claude Code Ships 'claude plugin eval' as Hamel Husain Details Four Key Shortcomings
sh_reya · x · 2026-09-12
ClaudeDevs shipped a new Claude Code feature, claude plugin eval: create test cases, run your plugin or skill against them, score the runs, then re-run each case without the plugin to measure its actual value.
Hamel Husain tried it and outlined improvements:
- In-situ feedback: the workflow quizzes you up front about your recollection of plugin use; better to collect feedback while you're actually using it
- Error discovery first: it auto-builds datasets and judges from upfront answers and docs; it should start by annotating real traces/session history with you to find what's broken
- Clunky setup: configuring it for a skill that wasn't a plugin took 10 minutes of asking Claude
- Report readability: the generated HTML takes a long stare to decode
Related event: Claude Code launches plugin eval to quantify plugin and skill value(10 posts)→
More from coding & agent
- Zed CEO on Meta's coding agent: fast and pleasant, but needs a lot of hand-holding — zeeg · 2026-09-12
- Token anxiety with Fable and Astra: dev neurotic-prompts and watches sessions to stop runaway spend — johnlindquist · 2026-09-12
- DSPy 3.4 RC drops litellm: faster imports and a much lighter dependency tree — lateinteraction · 2026-09-12
- GPT-6 Astra review: stunning at 3D games and computer use, still not a daily driver — petergyang · 2026-09-12
- zeeg: UIs are dead, use traces — recommends vitest-evals over UI-based evals — zeeg · 2026-09-12
- AI Is Reversing Developer Sentiment Toward the TypeScript Effect Library — ethanniser · 2026-09-12