StreamGaze, first benchmark for gaze-guided temporal reasoning in streaming video, accepted at NeurIPS
mohitban47 · x · 2026-10-03
- StreamGaze, developed with Adobe Research, has been accepted to the NeurIPS 2026 E&D track (scores 5/5/4). It is the first benchmark evaluating whether MLLMs can use human gaze signals for temporal reasoning over streaming egocentric videos.
- The benchmark spans past, present, and proactive reasoning across 10 tasks, from gaze-sequence matching to alerting when objects enter the field of view.
- A data construction pipeline aligns egocentric videos with raw gaze trajectories via fixation extraction, region-specific visual prompting, and scanpath construction to produce spatio-temporally grounded QA pairs.
More from Multimodal
- Codex rebuilds an After Effects video from scratch in ~1 hour via Higgsfield MCP — CSProfKGD · 2026-10-03
- invideo launches agentic video editor with multitrack timeline and custom AI agents — azed_ai · 2026-10-03
- Krea2 outputs coming out too soft: users struggle to fix despite parameter tweaks — maxiedaniels · 2026-10-03
- Reviewing AI video with separate checks for motion, subject, and background via VL models — professr_dumbledore · 2026-10-03
- ByteDance ships 4-step DMAD LoRAs for H3, cutting generation steps further — linoy_tsaban · 2026-10-03
- Sphere Encoder 2: Turning an Autoencoder into a 1-4 Step Image Generator — kastnerkyle · 2026-10-03