TennisVAR grounds tennis tactical reasoning in stroke evidence, crushing GPT-5.5 on localization
量子位 · wechat · 2026-08-20
A team from Xiamen University, Shanghai Innovation Institute and SJTU propose TennisVAR, tackling the gap where sports multimodal models "sound right" without truly understanding video. It reconstructs tactical logic along stroke chains and cites the exact strokes supporting each conclusion.
- New task & dataset: stroke-evidence-grounded tactical reasoning with benchmark TRACE — 11,189 rally videos, 41,485 strokes, 25,429 tactical units from 109 pro matches; models must output answers, hierarchical tactical labels, stroke sequences, key strokes and reasoning, with evidence grounded to racquet-contact frames.
- Method: Event→Relation→Evidence→Tactic pipeline; an event parsing module fuses DINOv3 features, short-term motion cues and TrackNet ball tracks, while a tactical-graph temporal reasoner links strokes by time and per-player decision continuity.
- Results: GPT-5.5 zero-shot scores T-F1@8 of 37.03 vs 73.04 for TennisVAR (T-IoU@4: 56.19 vs 24.54). Ablations show removing graph reasoning drops T-F1@8 by 17.14 points, confirming structured modeling matters.
Paper: arxiv.org/abs/2608.12920; code open-sourced.
More from Multimodal
- GTA: San Andreas Reimagined in Claymation Style — ShaKodemon · 2026-08-20
- Runway announces 15 winners for 'Another Big Ad Contest for Products That Don't Exist' — Kyrannio · 2026-08-20
- DayZ Reimagined as a Movie with AI — blufattanza89 · 2026-08-20
- AI-Generated Video: Cooking Japanese Tamagoyaki — Independent-Arm-7397 · 2026-08-20
- AI Recognizes Song in Just 1 Second — No_Tower_ai · 2026-08-20
- Krea 2 excels at generating stunning pixel art video game levels — Neggy5 · 2026-08-20