Introducing GDPevo: A Benchmark for Evaluating Agent Self-Evolution in Business Workflows

PrismShadow · hf · 2026-08-06

GDPevo is a novel benchmark designed to evaluate agent self-evolution capabilities, grounded in real-world enterprise workflows. Existing benchmarks often struggle to attribute test-time performance gains to prior training experiences and remain vulnerable to data contamination.

The benchmark's core mechanism, rule hybridization, decomposes enterprise workflows into atomic business rules, distributes them across training tasks, and recombines them in held-out test tasks. This ensures that test performance is genuinely attributable to the agent's accumulated experience. GDPevo spans multiple domains including CRM, ERP, finance, healthcare, and legal. Evaluations show that self-evolution consistently improves held-out accuracy by up to 16.44 percentage points, though the best-evolved agents remain far below the fully informed oracle ceiling of 91.6%.

Original post →

More from coding & agent

coding & agent channel →