Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
The paper introduces Imagine3D-LLM, an MLLM that learns to assemble a compact 3D representation from multi-view images before formulating text answers. By appending learnable summary tokens and decoding them into a 3D Gaussian Splatting representation supervised by photometric reconstruction loss, the model induces cross-frame correspondence across underlying image features. This method consistently outperforms prior approaches on spatial reasoning and 3D understanding benchmarks.
Imagine3D-LLM learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation.
Reconstruction supervision induces stronger cross-frame correspondence within the LLM's underlying image features.
Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks.