Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
The paper proposes the Mind2Dialogue framework to bridge the supervision gap regarding unspoken user beliefs and goals in LLM assistant training. The system employs a psychology-guided simulator to update evolving mental states through interaction, generating coherent conversations and informed responses from an Oracle assistant. Privileged distillation then trains models on these well-informed responses. Evaluations demonstrate that training on the full Mind2Dialogue corpus improves every reported personalization metric over the corresponding instruction-tuned baseline models like Qwen, Llama, and OLMo.
Mind2Dialogue improves every reported personalization metric over the corresponding Qwen, Llama, and OLMo instruction-tuned baselines.
Gains include improvements of 26.6 to 40.9 percentage points in preference-following generation.