Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
QUESTION — How can we overcome the limitation of transformers stopping contextual computation too early without retraining the entire model?
The study demonstrates that pretrained transformers utilize little of their depth for in-context referencing, with 13 base models reliably following only 1.4-3.6 lines. To address this, the authors introduce a task-trained rank-8 LoRA at an early layer while keeping all model weights frozen. This mechanism triggers a relay effect across intermediate layers, substantially extending context processing depth to handle chains ranging from 50 lines up to over 160 lines depending on the configuration.
Thirteen base models reliably follow only 1.4-3.6 lines.
Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains.
Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight.
Lunamos · 29 Sept 2026
read the original ↗