Rushes: Dataset Reveals Limits of LLMs in Adapting to Individual User Preferences
A new arXiv preprint introduces Rushes, a dataset of over 44,000 decision events from interactive narrative games, capturing how thousands of users make sequential choices. The study finds that leading large language models, including GPT-5, do not outperform simple baselines in predicting individual user choices, highlighting a persistent 'Engagement Gap.' This suggests that current alignment methods, which optimize for population-level preferences, may be inadequate for capturing diverse, context-dependent user behaviors.
Why it matters: The findings raise questions about the ability of current AI alignment techniques to personalize responses and adapt to individual users, a key challenge for future AI systems.
Full story at: arXiv Computation and Language ↗