ByteDance's TraceDance auto-builds agent behavior benchmarks from real deployment traces

ByteDance · hf · 2026-09-29

TraceDance is an automated system that turns real agent deployment traces into targeted behavior benchmarks for developer-specified undesirable behaviors, going beyond fixed benchmark suites.

Key techniques: Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash LLM, while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. Benchmarks use decision-point continuation to grade an LLM's next turn at a recorded decision point with a behavior-specific rubric—no reference answer or environment replay needed.

Experiments on coding and general tool use drew on 252,557 sessions, producing 107 benchmarks with 4,125 instances and fulfilling 95.3% of build requests. Human annotators confirmed 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments matched inter-annotator agreement. Nine frontier LLMs averaged only a 26.7% pass rate, revealing persistent behavior weaknesses at decision points. The authors position TraceDance as a key component of a recursive self-improvement (RSI) loop.

Original post →

More from coding & agent

coding & agent channel →