Arize AI shows how agents can self-learn from production traces to fix their own bugs
AI Engineer · youtube · 2026-10-11
AI Engineer published a talk by Arize AI's Fuad Ali on building self-learning loops for agents:
- The demo uses a shopping assistant (Wonder Toys) built on the OpenAI Agents SDK that confidently returns products outside a user's budget due to a broken price filter.
- Instead of treating telemetry as a dashboard for humans, a coding agent consumes traces and the app repo to explain failures: Arize Signal clusters recurring failures, gathers evidence, opens investigations or PRs; Arize skills give coding agents direct access to traces, datasets, and evaluators.
- Evaluation design distinguishes a fixed LLM judge rubric from a harness that can inspect tool calls, code, and external context — the budget bug becomes an eval checking whether returned prices satisfy the requested range.
- Agent experiments replay failure datasets against a dev endpoint, comparing quality, latency, and cost before shipping, closing the loop of auto-detect → fix → validate on real failures → human review.
Resources: arize-skills and project-rosetta-stone are open source on GitHub.
More from coding & agent
- Developer turns X bookmarks into a personal library, launching free and open source in 48 hours — nutlope · 2026-10-12
- One question to your coding agent: 6x speedup and 8x cost cut — venturetwins · 2026-10-12
- LangChain launches Managed Deep Agents: agents as directories, fully hosted runtime — EdenEmarco177 · 2026-10-12
- Web Browsing Agent Runs Locally on a 9B Model, Tasks Done in ~11-17s — BangsFactory · 2026-10-12
- 'Vibe coding' should be retired — AI now writes more coherent code than most engineers — signulll · 2026-10-12
- AI scans satellite photos to find driveways and auto-sends $90 powerwash offers — nikitabier · 2026-10-12