Sharon Li's Simons talk: turn-by-turn metrics to trace and diagnose agent failures
SharonYixuanLi · x · 2026-10-10
Sharon Yixuan Li (UW-Madison) spoke at the Simons Institute on a growing gap: AI agents are becoming autonomous faster than our ability to monitor them. She presented three NeurIPS'26 papers quantifying agent progress and diagnosing failures at turn-by-turn granularity:
- Tracing Agentic Failure from the Flow of Success: tracing failure sources from success trajectories.
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents: a 'progress advantage' signal for measuring agent progress.
- Hide-and-Seek in Trajectories: discovering failure signals for VLA runtime monitoring.
Talk video is available.
More from coding & agent
- Rogue agentic deployments should be tracked as APTs, researcher argues — nitarshan · 2026-10-11
- Open Dots: MIT-licensed self-hosted agent workspace with 5.6k GitHub stars — matchaman11 · 2026-10-11
- From 'Vibe Coding Is Evil' to 'My New Game Is All Vibe-Coded' in 10 Months — DeryaTR_ · 2026-10-11
- NYC AI dinner leak claims all frontier labs run "RSI loops" (unverified) — Hesamation · 2026-10-11
- Garry Tan's multi-agent workflow: one thread runs the merge PR queue, only green CI hits master — garrytan · 2026-10-11
- Plane Agents Burn Massive Tokens Just Two Weeks After Launch — JosephJacks_ · 2026-10-11