72-trial ablation: do agent skills actually help or just burn tokens?

rseroter · x · 2026-09-16

@kweinmeister ran an ablation study: 12 Cloud use cases across 72 trials to test whether loading skills into an AI coding agent actually helps, evaluating the Google Cloud Developer Plugin with NVIDIA's SkillEvaluator and Harbor.

The motivation: every time an agent guesses the wrong CLI command instead of checking syntax, you burn tokens. Do essential skills keep the agent on track? Worth reading for anyone building or using agent skills.

Original post →

More from coding & agent

coding & agent channel →