Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models
This work investigates how large language models handle clashing moral values using a 12,000-instance dataset covering pairwise value conflicts across multiple languages. To decouple dataset correlation learning from abstract values, the authors propose a task vector transfer experiment. By orthogonalizing the computed task vector with respect to the general instruction-following vector, the method successfully isolates the direction of specific value preferences, which can then be used to conduct task arithmetic and produce a model with the opposite stance.
GPT-5-mini consistently favors Honesty over Autonomy across all five languages when no policy is given.
Both plain fine-tuning and Direct Preference Optimization fine-tuning effectively remove first-option bias, increasing accuracy to greater than 98% on Llama-3.2-1/3B models.
Task vector transfer based on orthogonalization effectively isolates the direction of specific value preference for task arithmetic.