VepAgent Uses Tool-Augmented RL to Fix Video Event Prediction's Causal Gap
UCSC-VLAA · hf · 2026-10-08
UCSC-VLAA proposes VepAgent, an agentic framework combining causal-transition reasoning with tool-augmented reinforcement learning for Video Event Prediction (VEP).
- Problem: MLLMs rely on retrospective summarization and text-centric priors, limiting their ability to bridge unobserved causal transitions when predicting future video events.
- Method: builds futurebench-4K, a chain-of-thought SFT dataset structuring deduction of unobserved intermediate states; adds a diagnostic tool library (state tracking, frame retrieval, region magnification) for agents to recover missing spatio-temporal evidence; and a composite reward jointly optimizing prediction accuracy, causal coherence, and reliable priors to force genuine visual grounding.
- Results: state-of-the-art on FutureBench and NEPBench, significantly outperforming larger MLLMs.
More from Research
- FractAL introduces soft acquisition-strategy selection for batch-mode active learning — anshulkundaje · 2026-10-08
- DeLM's decentralized multi-agent system runs 2.49x faster, but MAS evals pick wildly different metrics — jyangballin · 2026-10-08
- Study: RAG retrieval diversification helps only on redundant multi-evidence pools, paper proposes per-query rule — _reachsumit · 2026-10-08
- Quantize by Drift: label-free mixed-precision quantization for text embedders hits 0.911 Spearman — _reachsumit · 2026-10-08
- RunningTab: environment-side task ledger consistently beats in-model tracking across 3 benchmarks and 3 LLMs — RexDouglass · 2026-10-08
- ColPali-style visual document indices can be inverted: 47% of words recovered, source page ranked first 98.4% — _reachsumit · 2026-10-08