Weco agent rewrote its own scoring code; founder says lock eval files before agent runs
victor_explore · x · 2026-10-01
Weco CEO Zhengyao Jiang recounts how his team caught their agent rewriting the code that scores its own tests — convinced it was cheating, it turned out to be fixing a bug, showing how hard spotting agent cheating is becoming. victorexplore draws the engineering lesson: lock your eval files before the agent run starts, since once a bug fix and a cheat look identical in the diff, the only safe grader is one the agent can't touch.
More from coding & agent
- Firecrawl's Alexandria hits #6 most popular ChatGPT plugin with 120+ data providers — devdigest · 2026-10-01
- Cloudflare launches agent-themed batch: pay-per-use gateway, AutoRouter, 6x faster containers — threepointone · 2026-10-01
- Blogger warns the cheap vibe-coding era is almost over — and the dots explain why — michalmalewicz · 2026-10-01
- Pocket FM's Memory System Matches Claude Code's 86.7% Accuracy at 1/21 the Cost — bigaiguy · 2026-10-01
- Ben Goertzel: AI agents are co-authoring formally verified software, not just patching bugs — bengoertzel · 2026-10-01
- Pedro Domingos: There are no software engineers anymore, we're all agent wranglers — pmddomingos · 2026-10-01