CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — architecture 4 upvotes

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

QUESTION — How can we prune the encoder of large ASR models like Whisper without requiring custom inference code?

Pruning large ASR transformer models like Whisper has focused on decoders to speed up transcription, while encoder pruning remains rare due to custom inference code requirements. The authors present an approach ranking encoder layers by leave-one-layer-out Word Error Rate change. The six layers causing the least change, representing 18.5% of the encoder stack, are removed. Unlabeled monolingual speech data is used for distillation to recover performance degradation from zero-shot pruning, bringing Mean WER to 20.1% compared to 21.9% zero-shot.

The six layers that cause the least change are removed, corresponding to 18.5% of the encoder stack.

Mean WER across four languages increases to 20.1% after distillation, compared to 21.9% zero-shot, going from a baseline of 18.2%.

rasgaard · 23 Sept 2026 read the original ↗
↑