StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training
The paper points out that training instability in vector-quantized (VQ) discrete visual tokenizers stems from the entanglement of Encoder-Decoder and Codebook training, where modules fail to fulfill their roles in isolation. The authors propose StableVQ, comprising three components: Dynamic STE to correct Encoder learning instability, Region VQ Loss to reconceive the Codebook's learning objective for tracking the encoder output distribution independently, and Decoupled Schedule to assign independent learning rate schedules to each module. StableVQ is lightweight and introduces no learnable parameters.
Built on top of shared-projection codebooks, StableVQ is lightweight and introduces no learnable parameters.
Experiments on ImageNet demonstrate consistent improvements in training stability, codebook utilization, and reconstruction quality across diverse codebook sizes and initialization settings.