Annotating CUA agent data: tasks go stale, so pseudo-annotate the tasks themselves

mervenoyann · x · 2026-09-20

Hugging Face engineer Merve Noyan shares a side-project lesson on building datasets for computer-using agents: most web task datasets are stale because websites change constantly. Her key insight is to pseudo-annotate the problem/task itself from scratch — not just answers to a given problem — since tasks also go outdated, then have agents visit real sites, execute the task, and define completion criteria for rewarding.

Related event: HF Engineer Shares Lessons on Pseudo-Labeling Data for Web Agents(2 posts)→

Original post →

More from coding & agent

coding & agent channel →