VisionHOPE: First Visual Backbone Formulated as a Self-Modifying Learning System

CASIA · hf · 2026-09-30

Researchers from CASIA introduce VisionHOPE, the first generic visual backbone formulated as a self-modifying learning system. Building on Nested Learning's self-referential construction, what the model remembers and how it learns co-evolve within an image via five coupled memories governing content, K/V representations, learning rate, and retention. To stabilize the unconstrained self-referential updates, the authors derive a stability-matched step-size control scheme (soft cap on injection plus spectral clamp on memory transitions) and prove non-expansive memory dynamics per scan. Chunks are aligned to image rows and columns across four directional scans for 2D feature maps. VisionHOPE achieves competitive results on ImageNet-1K, COCO, and ADE20K, with code released on GitHub.

Original post →

More from Research

Research channel →