The VISTA visual harness allows multimodal models to achieve a perfect score on all 25 public games within the ARC-AGI-3 benchmark, completing tasks with 57.4% fewer actions than human participants. This development, detailed in a paper from MIT researchers Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, and Kaiming He, demonstrates that appropriate harnessing can unlock advanced reasoning in diverse interactive environments.

VISTA provides a general-purpose multimodal model with "long-horizon vision," allowing it to directly observe environments through high-dimensional visual input, such as raw PNG images. A key component of VISTA is its lossless visual memory, which stores every frame from the environment in its original form, complete with turn and frame indices. This contrasts with standard Vision-Language Models (VLMs) and Large Language Models (LLMs), where memory mechanisms like the KV cache are often compressed, lossy, and have limited horizons, potentially losing past visual states. VISTA's explicit visual memory ensures that the complete pixel history remains accessible, allowing the model to actively retrieve and reorganize past observations as it reasons. The model can request specific views for comparison or use `read_pixels` to obtain exact color samples from selected regions.

The researchers used Claude Opus 5.0 as the base model for their experiments. Prior to VISTA, Claude Opus 5.0 achieved a Relative Human Action Efficiency (RHAE) score of 40.68 on ARC-AGI-3. With VISTA, this score improved to a perfect 100.00. This efficiency gain is also reflected in action count, with the model using 57.4% fewer actions than first-time human players. The ARC-AGI-3 benchmark, developed by AI researcher Francois Chollet, assesses a system's ability to handle unfamiliar situations by requiring it to infer rules and goals in game-like environments without explicit instructions.

The VISTA framework emphasizes a minimalist design, avoiding complex systems or task-specific engineering. Its simple architecture allows for natural extension to various visual environments with minimal adaptation, outperforming baselines that use the same underlying models with less sophisticated harnesses across other visual games and puzzles. For instance, on GPT-5.6 Sol, replacing text grids with images and incorporating VISTA's memory and inspection tools improved RHAE from 13.33 to 99.00. This approach also significantly reduces token consumption; a 64x64 numerical grid consumes approximately 4,000 text tokens, whereas a rendered 512x512 image uses around 308 visual tokens, cutting token usage per game from 71.9 million to 30.7 million.

The work suggests that direct visual reasoning combined with active inspection and lossless visual memory can be more effective than symbolic or code-based world models for complex interactive visual domains like ARC-AGI-3. The researchers indicate that VISTA's effectiveness across diverse visual environments highlights the importance of visual harness design in advancing multimodal agents.