TraceML compares 4,465 human vs 207 agent Kaggle trajectories: agents use narrower research skills
burny_tech · x · 2026-10-01
TraceML is a trajectory-level analysis tool for systematically comparing AI agents with human experts, accepted at NeurIPS 2026 E&D Track and the COLM Workshop.
- It annotates 4,465 human and 207 agent Kaggle competition trajectories.
- Comparing Codex agents with Kaggle Grandmasters reveals humans draw on a broader set of research skills, while agents concentrate on a narrower subset.
- Distilling the human-research skills uncovered by TraceML back into Codex agents substantially improves final outcomes.
- The work offers a quantifiable way to understand and improve auto-research agents.
More from coding & agent
- Posting feature maps to X and developing straight from the timeline — Baconbrix · 2026-10-01
- Linewise's video agent hits 31x GPU throughput at 1/15 cost on Inco inference infra — songhan_mit · 2026-10-01
- Enterprise AI startups should pivot to being fulltime OpenClaw FDE shops, says founder — heyneighbor · 2026-10-01
- Founder Moves All Coding Agents Back to Fable 5.1, Citing Least Supervision Needed — bindureddy · 2026-10-01
- Hugging Face engineer ranks coding agents: ChatGPT Desktop first, Cognition last — NielsRogge · 2026-10-01
- Developer builds an agent skill that hunts for coupons before checkout — VelheticaTheStore · 2026-10-01