Agents lack a benchmark for discretion: all evals score completion, none score leakage
victor_explore · x · 2026-09-24
Citing Zuckerberg's point that personal AI agents need a skill coding agents never did — discretion (e.g., booking a table while working around a dietary restriction or pregnancy and telling the restaurant none of it) — developer @victorexplore flags a blind spot in agent evaluation:
- Every existing agent eval scores whether the task was finished; none score how much information the agent leaked along the way
- His practical prescription: start logging what your agent hands to third parties
- That leakage number, he argues, is the first thing users will ask about
A concrete call to treat privacy discretion as a measurable agent skill, not just a product talking point.
More from coding & agent
- The whole voice agent demo cost ~$0.02: KugelAudio at $0.035/min, Gladia $0.75/hr — tobowers · 2026-09-24
- A real-time voice agent with zero US servers: Gladia STT, Gemma 4 on Scaleway, KugelAudio TTS — tobowers · 2026-09-24
- Devs add 10x more test harnesses just to slow down AI coding agents — cjimti · 2026-09-24
- Screenshot-annotated web edits tested with Ling-3.0: only 2 of 3 runs passed — alifcoder · 2026-09-24
- Claude directory approved an MCP submission in just 5 minutes — devenbhooshan · 2026-09-24
- Usage resets land before rest days, forcing devs into all-night AI coding sessions — cjimti · 2026-09-24