Disentangling Representation Evolution in Transformers through Directional Decomposition
This study investigates the functional geometry of Transformer representations by decomposing learned updates into parallel and perpendicular components relative to the current hidden state. Analysis across pretrained models reveals substantial parallel components beyond the residual identity path. Applying this decomposition to attention and MLP updates, the authors uncover a strong spatial asymmetry in editing robustness: exclude-self value-space parallel manipulation is markedly more robust. Furthermore, full-aggregate parallel suppression during from-scratch pretraining successfully lowers validation-loss trajectories and improves downstream performance.
Exclude-self value-space parallel manipulation is markedly more robust than residual-space and perpendicular counterparts.
Full-aggregate parallel suppression during from-scratch pretraining lowers validation-loss trajectories and improves downstream averages.