FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
QUESTION — How can the reconstruction-generation gap in representation autoencoders be mitigated without modifying the pretrained visual encoder?
The authors introduce FuseReg, a regularization method that replaces heuristic layer selection with training over random subsets of encoder layers. By penalizing cross-layer disagreement sensitivity, this technique yields a single decoder capable of reconstructing from various layer fusions without retraining. Experiments on ImageNet-256 using DINOv3-L show that decoder replacement alone reduces unguided gFID by 27%, while joint regularization during diffusion training reduces unguided gFID by 29% on DiT-Base.
Decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator.
Joint regularization of both stages reduces unguided gFID by 29% on DiT-Base.
Hongyang-Du · 25 Sept 2026
read the original ↗