CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — architecture 16 upvotes

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

QUESTION — How can joint multimodal representation learning and generation be achieved without bottlenecking performance behind frozen embeddings?

The paper introduces FLAT (Flexible-Length Aligned Transmodal representations), a pre-training framework that jointly optimizes a shared multimodal encoder with text-to-image (T2I) and image-to-text (I2T) decoders. By combining contrastive alignment with bidirectional cross-modal generative objectives, FLAT maps inputs into a unified continuous 1D sequence space, applying nested dropout to support dynamic output lengths. FLAT achieves a T2I GenEval score of 71.1 pre-trained, reaching 83.1 GenEval with task-specific fine-tuning, alongside strong retrieval results on MS-COCO and Flickr30K.

FLAT achieves a T2I GenEval score of 71.1.

Task-specific fine-tuning achieves 83.1 GenEval on T2I generation.

It achieves 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO image captioning.

imguangyu · 15 Sept 2026 read the original ↗
↑