Desktop agents: what evidence proves a task is actually done, from saved drafts to uploads?
Special-Ad8671 · reddit · 2026-10-06
A developer building a voice-first Windows assistant argues desktop agent testing can't stop at 'the input was sent' and asks what evidence suffices to call a task done.
Concrete ambiguous cases cited:
- Text field shows expected content, but autosave is unconfirmed
- File picker accepted a file, but upload completion is unknown
- A window changed after a click, so old control identity is unsafe to reuse
His proposed design: bind each action to a freshly observed window/control, inspect results before the next step, check explicit postconditions, and halt on ambiguous states rather than replaying input. The remaining gap is between local UI state and the app actually accepting a change.
He asks the community whether app-specific checks, accessibility-state checks, or a separate verifier would count as sufficient evidence for a saved email draft or completed upload — inviting concrete failure cases over benchmark scores.
More from coding & agent
- Zeroization can make things worse: how wiping secrets creates more copies — jedisct1 · 2026-10-06
- Indie dev launches agent-first creator marketing platform Clipatra, pays out $3,000+ — tibo_maker · 2026-10-06
- TensorFold bonds dual Thunderbolt 5 links for 83% throughput boost on Apple Silicon — AIFlow_ML · 2026-10-06
- A 45-minute visual tour of graph theory, taught through the author's hometown — TivadarDanka · 2026-10-06
- Open-weights Kolibri-1 plays Breakout with no fine-tuning at ~25ms per move — Nils_Reimers · 2026-10-06
- Gatana adds centralized skills sync to MCP Gateway across all agents — Gatana_Official · 2026-10-06