Agents lack a benchmark for discretion: all evals score completion, none score leakage

victor_explore · x · 2026-09-24

Citing Zuckerberg's point that personal AI agents need a skill coding agents never did — discretion (e.g., booking a table while working around a dietary restriction or pregnancy and telling the restaurant none of it) — developer @victorexplore flags a blind spot in agent evaluation:

A concrete call to treat privacy discretion as a measurable agent skill, not just a product talking point.

Original post →

More from coding & agent

coding & agent channel →