FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
The paper introduces FLAT (Flexible-Length Aligned Transmodal representations), a pre-training framework that jointly optimizes a shared multimodal encoder with text-to-image (T2I) and image-to-text (I2T) decoders. By combining contrastive alignment with bidirectional cross-modal generative objectives, FLAT maps inputs into a unified continuous 1D sequence space, applying nested dropout to support dynamic output lengths. FLAT achieves a T2I GenEval score of 71.1 pre-trained, reaching 83.1 GenEval with task-specific fine-tuning, alongside strong retrieval results on MS-COCO and Flickr30K.
FLAT achieves a T2I GenEval score of 71.1.
Task-specific fine-tuning achieves 83.1 GenEval on T2I generation.
It achieves 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO image captioning.