The Router Within: Eliciting Native Skill Routing from a Frozen LLM
This paper introduces Gavel, an architecture that extracts skill-routing signals directly from a frozen LLM's mid-layer states without requiring skill text in the context. Gavel operates in two steps: a glance phase that scores the full skill library against compact per-skill banks, and a verdict phase that confirms the selection using the model's own likelihood. Experiments show that this lightweight approach outperforms heavy progressive disclosure and retrieve-and-rerank pipelines across multiple public benchmarks and agent trajectories, while scaling effectively with backbone improvements.
Gavel outperforms progressive disclosure and retrieve-and-rerank pipelines that add 1.2B to 16B external parameters.
It achieves improvements of up to 13.4 points on written tasks.
It achieves improvements of up to 21.9 points when the need for a skill arises mid-rollout.