Shortening agent prompts broke semantic testing: keyword checks missed a P0-to-P1 rule change
Dinu_Dev · reddit · 2026-10-08
A developer shortened their coding-agent instructions and found their keyword-based coverage test passed four deliberately altered versions — one changed "stop only for P0" to "stop only for P1" — showing keyword checks can't catch meaning changes.
Their current fix: verify 15 required sentences verbatim and test intentionally corrupted copies, which caught those cases. But exact-text checks can't tell whether a shorter paraphrase preserves meaning.
They're asking the community for small behavior-based tests that catch a changed rule without rejecting every harmless rewording.
More from coding & agent
- 10 open-source GitHub repos to cut your AI agent's token bill — victor_explore · 2026-10-08
- Dev Argues Coding Is Solved by AI, the Hard Part Is Coming Up with Novel Ideas — Suspicious_Fun_6338 · 2026-10-08
- Paper models LLMs as oracles in a pushdown automaton, unifying Agents and Workflows — China666 · 2026-10-08
- Naveen Rao: Let Agents Probe Every PowerPoint Feature, Then Rewrite the App — NaveenGRao · 2026-10-08
- Coding agents can build funnels but can't grasp why humans buy, marketer argues — boringmarketer · 2026-10-08
- AWS open-sources Strands Box, an OS-level sandbox for AI agents — TheNickWalsh · 2026-10-08