CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 7 upvotes

On the Diffusibility of High-Dimensional Latents

QUESTION — Why does finetuning visual encoders for image reconstruction reduce effective dimensionality and impact generation?

The paper investigates Representation Autoencoders (RAEs) for diffusion models operating in visual encoder feature spaces. The authors find that finetuning these encoders for image reconstruction reduces effective dimensionality, forcing models to fit orthogonal noise directions outside the signal manifold. To fix this optimization inefficiency, they propose using clean data parameterization (x_{0}-prediction) instead of standard velocity prediction in flow matching, which consistently improves text-to-image generation performance.

Using clean data parameterization (x_{0}-prediction) consistently improves text-to-image generation performance across strong-reconstruction encoders.

chfeng · 23 Sept 2026 read the original ↗
↑