Desktop agents: what evidence proves a task is actually done, from saved drafts to uploads?

Special-Ad8671 · reddit · 2026-10-06

A developer building a voice-first Windows assistant argues desktop agent testing can't stop at 'the input was sent' and asks what evidence suffices to call a task done.

Concrete ambiguous cases cited:

His proposed design: bind each action to a freshly observed window/control, inspect results before the next step, check explicit postconditions, and halt on ambiguous states rather than replaying input. The remaining gap is between local UI state and the app actually accepting a change.

He asks the community whether app-specific checks, accessibility-state checks, or a separate verifier would count as sufficient evidence for a saved email draft or completed upload — inviting concrete failure cases over benchmark scores.

Original post →

More from coding & agent

coding & agent channel →