Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery
Pruning large ASR transformer models like Whisper has focused on decoders to speed up transcription, while encoder pruning remains rare due to custom inference code requirements. The authors present an approach ranking encoder layers by leave-one-layer-out Word Error Rate change. The six layers causing the least change, representing 18.5% of the encoder stack, are removed. Unlabeled monolingual speech data is used for distillation to recover performance degradation from zero-shot pruning, bringing Mean WER to 20.1% compared to 21.9% zero-shot.
The six layers that cause the least change are removed, corresponding to 18.5% of the encoder stack.
Mean WER across four languages increases to 20.1% after distillation, compared to 21.9% zero-shot, going from a baseline of 18.2%.