Claude Code can now eval whether your skills and plugins actually help
JeremyNguyenPhD · x · 2026-09-22
Addy Osmani reminds devs that Claude Code can test whether a skill or plugin actually improves its answers. Run claude plugin eval init, describe what a good result looks like, and Claude writes the test cases into your plugin's folder; then run claude plugin eval . to execute them. A practical workflow for systematically validating skills/plugins instead of trusting vibes.
Related event: Claude Code Adds plugin eval to Test Plugin Effectiveness(2 posts)→
More from coding & agent
- Speaking UI layouts into existence with Jev, shadcn and a local transcription model — _AustinCalvert_ · 2026-09-22
- ai-copywriter hits 1k stars: an agent skill that writes converting copy with zero AI tells — tom_doerr · 2026-09-22
- ContextBridge: an open-source pool that routes AI tasks across local machines, APIs and shared hosting — IamAngusU · 2026-09-22
- Microsoft's 44-page playbook: 100+ internal AI projects show licenses alone fail — alex_verem · 2026-09-22
- signal-mcp: open-source local-first MCP server exposing Signal to AI clients — googlarz · 2026-09-22
- mcp-server-npm: an MCP server to search, compare and inspect npm packages — modelcontextprotocol · 2026-09-22