GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
The paper proposes the geometry-native autoencoder (GAE) to address the limitation of visual generators that focus on appearance and lack 3D consistency. Rather than treating geometry as an additional output, GAE reparameterizes features from a geometry foundation model into a compact latent space that decodes jointly to appearance, depth, cameras, and point maps. Controlled experimental comparisons demonstrate that replacing standard latents with GAE improves visual quality and independently measured 3D coherence, reducing FVD on RealEstate10K and DL3DV while halving camera-trajectory error on RealEstate10K.
Replacing the latent with GAE improves visual quality and independently measured 3D coherence: FVD falls by 12.7% and 23.1% on RealEstate10K and DL3DV.
Camera-trajectory error is halved on RealEstate10K when using GAE.
The GAE latent is jointly decodable to appearance, depth, cameras, and point maps.