EgoTools benchmark: AI lags humans at understanding tool use in egocentric video
liuziwei7 · x · 2026-10-06
NTU S-Lab and Ropedia released EgoTools, a diagnostic benchmark testing whether AI can understand why people pick and use tools in first-person video, not just recognize them.
- Scale: 100 hours of egocentric video, 1,000 diagnostic questions, plus an 8B reference model.
- Fine-tuning the same backbone lifts scores from 50.0 to 60.9, beating all open-source models.
- Yet spatial reasoning degrades by 11 points, revealing a capability trade-off.
- Human experts still score 83.2, far above the best model.
More from Research
- Jon Barron walks through backpropagation by hand on a tiny two-layer network — techNmak · 2026-10-06
- New method dissects only task-relevant weights, making interpretability cheap enough for daily debugging — CatAstro_Piyush · 2026-10-06
- Debating computational irreducibility: if you've computed the Mandelbrot set, is the program just compression? — ctjlewis · 2026-10-06
- Social media use explains just 0.4% of teen well-being variation, researcher argues studies fail policy — asusarla · 2026-10-06
- Stanford Open-Sources DITTO-X: Force-Feedback Teleop With Reverse Human Intervention — CyberRobooo · 2026-10-06
- Cisco Benchmarks Decision Models: Jev Nears 31B LLM Judge on Zero-Shot Safety Classification — aminkarbasi · 2026-10-06