Persona Dosing: Calibrated Activation Steering for Graded Trait Control
This work introduces PersonaDose, a method for controlling language models via a trait description and a requested mean intensity. By specializing a shared FLAS controller on persona responses and calibrating its flow time against measured trait expression without paired training targets, the approach raises core-trait expression over contrastive activation addition. Evaluated across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, calibration-selected settings maintain an expression advantage on held-out questions and achieve small mean targeting errors across calibration-reachable targets per model.
PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition.
Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model.