From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
Researchers present 'weights to words', a method that automatically discovers domain-relevant preference dimensions described in natural language from choice data. This approach aims to make preference models more interpretable and editable by allowing users to inspect and modify model inferences in real time. Two pre-registered human-subjects experiments demonstrate that regularizing a preference model toward the learned basis improves prediction accuracy, and that incorporating user edits further enhances performance.
Why it matters: This work advances the interpretability and human-editability of preference models, potentially increasing trust and accuracy in AI systems that learn from human choices.
Full story at: arXiv Machine Learning ↗