DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence
The paper introduces DC-SAE (Decoupled Compact Semantic Autoencoder), a decoupled tokenizer combining macro-level semantic encoders for high compression ratios and pixel-level encoders for fine details, addressing the trade-off between compression fidelity and diffusion convergence speed. Evaluated on the ImageNet dataset at 512x512 resolution, DC-SAE achieves 32times spatial compression with 29.79 PSNR and 3.37 gFID, outperforming previous baseline DC-AE by 13.5% on PSNR and 54.9% on gFID. Additionally, a 1.6B-parameter DiT using DC-SAE achieves 0.84 on GenEval and 86.007 on DPG-Bench at 1024x1024 resolution.
On the ImageNet dataset with 512 times 512 resolution, DC-SAE achieves 32times spatial compression, with 29.79 PSNR and 3.37 gFID.
DC-SAE substantially outperforms previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5% and 54.9% on PSNR and gFID, respectively.
A 1.6B-parameter DiT using DC-SAE achieves 0.84 on GenEval and 86.007 on DPG-Bench for text-to-image generation at 1024 times 1024 resolution.