SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
QUESTION — How can scene tokenizers be designed to achieve predictable semantics and high video fidelity in autoregressive video generation?
This work addresses limitations in flexible tokenizers that only apply representation-alignment loss on early decoder hidden states. The authors introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads to reconstruct them from each retained token prefix alone. Experimental results demonstrate that a 201M SemanTok AR model matches or outperforms a VideoFlexTok AR model 3.4times its size, while maintaining high semantic alignment and generation fidelity across out-of-distribution classes and noise levels.
A 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4times its size.
mboss · 30 Sept 2026
read the original ↗