EASEL benchmark: multimodal agents fail at dexterous, closed-loop visual tool use
EASEL-Bench · hf · 2026-08-31
EASEL benchmarks fine-grained visual tool use in multimodal agents via reference-guided reconstruction and semantic tasks. It reveals current agents struggle with closed-loop precision and trajectory stability — seeing and precisely acting remains a weak point.
More from Research
- Medusa from training to inference: a two-part guide to multi-token prediction acceleration — No_Progress_5399 · 2026-08-31
- LayerRecall: layer-wise memory routing for long-horizon consistency in video generation — zju · 2026-08-31
- StarHarness: evolving fixed-weight agent harnesses for enterprise tool use — ServiceNow-AI · 2026-08-31
- ECCV 2026 Paper: Fourier Self-Supervision for Category Discovery — y_m_asano · 2026-08-31
- Paper: Deterministic Horizon Limits Pure Neural Reasoning — aronchick · 2026-08-31
- RAG poisoning causes 'attention collapse', fooling confidence detectors — rohanpaul_ai · 2026-08-31