CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 19 upvotes

PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion

QUESTION — How can dense prediction foundation models be simultaneously leveraged as alignment targets for pixel diffusion without causing gradient conflicts?

The paper introduces PixelDense, a method that routes semantic encoders (DINOv2, SAM2) through a semantic projection stream and geometric encoders (Depth Anything v2, Metric3D v2) through a geometric projection stream, while applying a weight-space orthogonality penalty to keep the streams in disjoint subspaces. This resolves semantic and geometric gradient competition during diffusion transformer training. Applied to PixelGen and DeCo, PixelDense raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093, improves PQ by up to 53.1% and depth AbsRel reduction by 36.0% at $\tau=0.5$, and reaches peak GenEval 1.23x faster from random initialization.

PixelDense raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093.

Shows up to 53.1% PQ gain and 36.0% depth AbsRel reduction at $\tau=0.5$.

Reaches the baseline's peak GenEval 1.23x faster from random initialization.

xsourse · 30 Sept 2026 read the original ↗
↑