AVIS paper accepted at NeurIPS 2026: jointly scaling visual context and reasoning per query

CSProfKGD · x · 2026-09-25

AVIS: Adaptive Visual Inference Scaling for Vision-Language Models has been accepted at NeurIPS 2026, by a team of 11 authors including Ahmadreza Jeddi.

Key idea: existing test-time scaling for VLMs optimizes a single axis of compute — either scaling visual context or doing more reasoning rollouts. AVIS adaptively allocates compute across both coupled axes on a per-query basis:

It's deployment-friendly: all rollouts share a single prefill pass and reuse the KV cache. Across diverse image and video reasoning benchmarks, AVIS achieves a better accuracy-compute trade-off than single-axis scaling methods. Paper, code and website are public.

Original post →

More from Models

Models channel →