CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — architecture 11 upvotes

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

QUESTION — How can scene tokenizers be designed to achieve predictable semantics and high video fidelity in autoregressive video generation?

This work addresses limitations in flexible tokenizers that only apply representation-alignment loss on early decoder hidden states. The authors introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads to reconstruct them from each retained token prefix alone. Experimental results demonstrate that a 201M SemanTok AR model matches or outperforms a VideoFlexTok AR model 3.4times its size, while maintaining high semantic alignment and generation fidelity across out-of-distribution classes and noise levels.

A 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4times its size.

mboss · 30 Sept 2026 read the original ↗
↑