Kaiming He team's VISTA paper hits arXiv: Claude aces all 25 ARC-AGI-3 games with 57.4% fewer actions
GregKamradt · x · 2026-10-04
The VISTA paper is now on arXiv from Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He, expanding on the team's August blog post with new ablations and evaluations beyond ARC-AGI-3 on visual games and reasoning tasks.
Key points:
- VISTA is a "visual harness" that gives a general multimodal model long-horizon vision: it perceives environments through raw visual observations and keeps a lossless visual memory the model can actively retrieve from while reasoning.
- On ARC-AGI-3, VISTA lifts Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, clearing all 25 public games with 57.4% fewer actions than first-time human players.
- Across three additional visual game and puzzle benchmarks, VISTA substantially outperforms baselines using the same underlying model with minimal harnesses.
The authors argue multimodal models already possess strong reasoning ability, and that the right harness is what unlocks it. Paper and code are both public.
Related event: MIT's VISTA Helps Claude Clear All ARC-AGI-3 Levels Without Training(2 posts)→
More from Research
- BF16 rounding breaks a conservation law, blowing up FlashAttention gradients late in training — HongyiWang10 · 2026-10-04
- Microsoft's ActiveSaddler adapts agent harness training scenarios, boosting Pass@1 by up to 7.5 points — dair_ai · 2026-10-04
- Anthropic's circuit tracing paper reverse-engineers how Claude 3.5 Haiku reasons internally — austinc3301 · 2026-10-04
- V-Rubrics: 50k visual samples split into 353k checkable criteria fix multimodal RL credit assignment — jiqizhixin · 2026-10-04
- Looped-DiT: 260M looped model beats 6.5x larger text-to-image rival with 4.9x less compute — Apprehensive_Sky892 · 2026-10-04
- Schmidhuber says he published the first concrete RSI algorithm back in 1987 — SchmidhuberAI · 2026-10-04