TimeEvo: failure-driven tool synthesis lifts time series agent accuracy on every task and backbone
Jie Yang · hf · 2026-09-29
Time series agents answer analytical questions by calling tools, but tool libraries are usually hand-picked in advance. The authors identify two failure modes:
- Human-agent tool misalignment: a library of 21 expert-curated tools helps some tasks and hurts others, dropping anomaly accuracy under every backbone tested;
- Silent harm: one round of generic self-revision changed 147 answers and broke 56, while the final score moved by less than a point.
The root cause: whether a tool helps is decided question-by-question at runtime, yet tools are supplied in advance and judged by a single average.
TimeEvo clusters diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools to fill them, and admits candidates through a paired admission gate.
Experiments on ten time series QA tasks across three backbones show improvements on every task and backbone starting from an empty library, and a library grown on a cheap model still transfers to stronger ones. Code is open-sourced.
More from coding & agent
- Solo dev wins YC hackathon with QM Motion, giving coding agents a sense of time — garrytan · 2026-09-29
- Million-MAU AI news aggregator AIHOT goes open source, dev shares full AI rewrite workflow — dotey · 2026-09-29
- Engineer's playbook: draft with Sonnet 5.5, build with Opus 5.5 to cut costs — every · 2026-09-29
- Recreating Famous UFO Encounters With Claude + Unreal Engine in Virtual Production — playertariat · 2026-09-29
- Claude refuses to speed up tests, says 'Running tests will take 35 hours' to respect rate-limit spec — julianharris · 2026-09-29
- Claude Opus 5.5 runs an AI game studio: builders, critics, revisers, verifiers — chrisfirst · 2026-09-29