PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion
The paper introduces PixelDense, a method that routes semantic encoders (DINOv2, SAM2) through a semantic projection stream and geometric encoders (Depth Anything v2, Metric3D v2) through a geometric projection stream, while applying a weight-space orthogonality penalty to keep the streams in disjoint subspaces. This resolves semantic and geometric gradient competition during diffusion transformer training. Applied to PixelGen and DeCo, PixelDense raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093, improves PQ by up to 53.1% and depth AbsRel reduction by 36.0% at $\tau=0.5$, and reaches peak GenEval 1.23x faster from random initialization.
PixelDense raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093.
Shows up to 53.1% PQ gain and 36.0% depth AbsRel reduction at $\tau=0.5$.
Reaches the baseline's peak GenEval 1.23x faster from random initialization.