CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — architecture 9 upvotes

Disentangling Representation Evolution in Transformers through Directional Decomposition

QUESTION — How can representation evolution in Transformers be decomposed into parallel and perpendicular components to improve editing and compression?

This study investigates the functional geometry of Transformer representations by decomposing learned updates into parallel and perpendicular components relative to the current hidden state. Analysis across pretrained models reveals substantial parallel components beyond the residual identity path. Applying this decomposition to attention and MLP updates, the authors uncover a strong spatial asymmetry in editing robustness: exclude-self value-space parallel manipulation is markedly more robust. Furthermore, full-aggregate parallel suppression during from-scratch pretraining successfully lowers validation-loss trajectories and improves downstream performance.

Exclude-self value-space parallel manipulation is markedly more robust than residual-space and perpendicular counterparts.

Full-aggregate parallel suppression during from-scratch pretraining lowers validation-loss trajectories and improves downstream averages.

shwai-he · 14 Sept 2026 read the original ↗
↑