CAVEAT testbed exposes how merchants can steer your shopping AI agent
ZacharyHuang12 · x · 2026-10-07
AI agents are about to spend your money, and merchants have every incentive to steer them away from your interests. Researcher YuxuanL introduced CAVEAT, a testbed for studying agent behavior when store-side incentives conflict with user intent.
Key releases:
- CAVEAT-Harness: a standardized harness for testing how agents respond to merchant-side manipulation
- CAVEAT-27B: a post-trained 27B model designed to keep shopping agents loyal to the user
The author's core claim: right now "the store usually wins" — merchants can influence agent decisions via page design and persuasive offers, and it takes explicit alignment work and benchmarks like CAVEAT to fix this. It's essentially user-side alignment research for agentic commerce, providing infrastructure for future evals and defenses.
More from coding & agent
- Claude Opus rebuilt Terraria in 3D: 346K lines of C# in 8 days, zero hand-written code — Promptmethus · 2026-10-07
- Dev puts the entire game of Minecraft inside Twitter with working multiplayer — jaivinwylde · 2026-10-07
- Scikit-learn Team Launches Skore, an Enterprise Platform for Tabular AI Built With AI Agents — GaelVaroquaux · 2026-10-07
- Amateur Team Discloses Identity-Constitutive Memory Protocol for AI Agents, Keeps Core Docs Secret — AdeptBathroom3318 · 2026-10-07
- Prompting agents to mimic Modal's GPU glossary styling yields polished HTML artifacts — sumitdotml · 2026-10-07
- Jordan and coauthors tackle when to stop generator-verifier loops while controlling false discovery — _onionesque · 2026-10-07