← Back to brief
ResearchOfficialPreprintarXiv Statistical ML

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

Researchers propose SARA, a contrastive framework for preference-based reinforcement learning (RL) that is designed to be robust to noisy labels and adaptable to various feedback formats. SARA learns a latent representation of preferred samples and computes rewards based on similarity to this latent, demonstrating competitive and more stable performance on continuous control offline RL benchmarks. The method shows statistically significant improvements over baselines and maintains higher correlation with environment rewards across different noise rates.

Why it matters: This work advances the robustness and versatility of preference-based RL, addressing the challenge of aligning AI with human intent when feedback is noisy or inconsistent.

Full story at: arXiv Statistical ML