ByteDance's TraceDance auto-builds agent behavior benchmarks from real deployment traces
ByteDance · hf · 2026-09-29
TraceDance is an automated system that turns real agent deployment traces into targeted behavior benchmarks for developer-specified undesirable behaviors, going beyond fixed benchmark suites.
Key techniques: Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash LLM, while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. Benchmarks use decision-point continuation to grade an LLM's next turn at a recorded decision point with a behavior-specific rubric—no reference answer or environment replay needed.
Experiments on coding and general tool use drew on 252,557 sessions, producing 107 benchmarks with 4,125 instances and fulfilling 95.3% of build requests. Human annotators confirmed 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments matched inter-annotator agreement. Nine frontier LLMs averaged only a 26.7% pass rate, revealing persistent behavior weaknesses at decision points. The authors position TraceDance as a key component of a recursive self-improvement (RSI) loop.
More from coding & agent
- Community-maintained list tracks 51 active AI angel investors with sources and verification dates — TheMoonMidas · 2026-09-29
- 400+ LLM Agents Living in a 2004-era MMO Server, All Local on Qwen 4B — kristiantalley679 · 2026-09-29
- px0 release fixes LSP memory leaks, ReDoS vulnerability, adds CSV table rendering — arpit_bhayani · 2026-09-29
- Full Prompt Shared: Make Your Agent Build an Offline Progress Dashboard Before Long Tasks — myLifeintheStack · 2026-09-29
- Local gamedev AI stack: Meta's Muse Glimmer nails tool use in one try where Gemma 4 12b fails — draginol · 2026-09-29
- Founder: AI Can Now Do the Work of 10 Employees Before You Hire — FinanceYF5 · 2026-09-29