Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
This work demonstrates that Large Language Models exhibit fundamental linearity: when distinct text stream inputs are linearly combined, the model outputs a superposition of individual next-token distributions. This property tends to diminish during pretraining but can be substantially restored via lightweight fine-tuning. Additionally, the authors introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.
LLMs exhibit superposition linearity when input streams are linearly combined.
Linearity can be substantially restored through lightweight fine-tuning.
A guided decoding procedure enables simultaneous generation of two coherent continuations from a single forward pass.