VisionHOPE: Visual Backbones as Self-Modifying Learning Systems
The authors introduce VisionHOPE, the first generic visual backbone formulated as a self-modifying learning system, enabling memory content and learning rules to co-evolve within an image. Building on Nested Learning (NL), it integrates five coupled memories storing content, generating keys and values, and governing learning rate and retention. To avoid instability, the authors derive a stability-matched step-size control scheme with a spectral clamp ensuring non-expansive memory dynamics. VisionHOPE achieves competitive results on ImageNet-1K, COCO, and ADE20K.
VisionHOPE is the first generic visual backbone formulated as a self-modifying learning system.
The system integrates five coupled memories that co-evolve along each scan of an image.
VisionHOPE achieves competitive results on ImageNet-1K, COCO, and ADE20K.