VisionHOPE: First Visual Backbone Formulated as a Self-Modifying Learning System
CASIA · hf · 2026-09-30
Researchers from CASIA introduce VisionHOPE, the first generic visual backbone formulated as a self-modifying learning system. Building on Nested Learning's self-referential construction, what the model remembers and how it learns co-evolve within an image via five coupled memories governing content, K/V representations, learning rate, and retention. To stabilize the unconstrained self-referential updates, the authors derive a stability-matched step-size control scheme (soft cap on injection plus spectral clamp on memory transitions) and prove non-expansive memory dynamics per scan. Chunks are aligned to image rows and columns across four directional scans for 2D feature maps. VisionHOPE achieves competitive results on ImageNet-1K, COCO, and ADE20K, with code released on GitHub.
More from Research
- Tempo (UIST 2026): A computer-use agent that reasons about your long-term goals — kenziyuliu · 2026-09-30
- UMAP Author Teases New Release Built for Much Larger Datasets — leland_mcinnes · 2026-09-30
- Chollet: The Test for Human-Level AGI Is Passing ARC-AGI-(n+1) on Release — fchollet · 2026-09-30
- Jev as a Reranker: a Solid Cost/Speed Default, but Not a Drop-In Silver Bullet — amaarora · 2026-09-30
- Reality Check: 4 Open-Source Models, 14,400 Real-World Rollouts to Benchmark Robot Policies — YuXiang_IRVL · 2026-09-30
- Agent kept misdiagnosing complaints — team shares 4 memory design decisions that fixed it — Financial-Shirt2304 · 2026-09-30