MIT open-sources VISTA visual harness, enabling Claude Opus 5.0 to solve all 25 ARC-AGI-3 games
JFPuget · x · 2026-09-06
MIT researchers (with Kaiming He as co-author) have open-sourced VISTA, a visual harness that gives general-purpose multimodal models a continuous visual interface for long-horizon reasoning in interactive environments.
- How it works: every environment frame is archived as visual memory, letting the agent revisit original evidence while reasoning and acting in an observe → reason → act → observe loop.
- Results: paired with Claude Opus 5.0 (via Claude Code at xhigh effort), VISTA completes all 25 public ARC-AGI-3 games with a 100% win rate and a perfect Relative Human Action Efficiency (RHAE) score of 100. Codex CLI with GPT-5.6 Sol reaches 99.
- Code is available on GitHub, with an accompanying blog post.
More from coding & agent
- Rork claims App Store publishing is now 2.5x to 4x faster — rudrank · 2026-09-07
- Hands-on test: AI ops agent Ghost diagnoses and fixes three simultaneous service failures — BLUECOW009 · 2026-09-07
- Dev runs a live show produced entirely by an agentic AI research team built on a wiki — SurvivedDravoswater · 2026-09-07
- Open-source Codex Skill turns one sentence into a holographic 3D card with Blender and Three.js — Promptmethus · 2026-09-07
- Jim Koppel: the vibe coder's way to understand code is running everything; the agentic engineer reads it — jimmykoppel · 2026-09-07
- OpenAI reportedly testing ChatGPT feature that mimics your writing style from emails and docs — VraserX · 2026-09-07