When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents
This work studies intent drift, a multi-turn failure mode where superseded parts of user intent continue to influence final answers or tool actions in LLM agents. The authors introduce IntentFlux, an executable benchmark converting verifiable tasks into dialogues with controlled intent changes, demonstrating that mean task scores drop as dialogues accumulate withdrawn information. They further introduce StateForge, a mechanism that explicitly maintains active requirements before generation, improving mean task scores on General-Test from 0.367 to 0.467 and establishing explicit state maintenance as a partial mitigation.
In a 627-case calibration, mean task score falls from 0.476 to 0.384 as dialogues contain more superseded and withdrawn information.
On General-Test, StateForge improves mean task score from 0.367 to 0.467.