← Back to brief
ResearchOfficialPreprintarXiv AI/ML

Masked Diffusion Language Models Enable Strong and Steerable Text-Based World Models for Agentic RL

A new study introduces masked diffusion language models (MDLMs) as world models for reinforcement learning (RL), addressing limitations of autoregressive models such as left-to-right bias. The researchers curated over 239,000 trajectories across nine environments and demonstrated that MDLMs provide greater coherence, groundedness, and rollout diversity than much larger autoregressive language models. Their plug-and-play GRPO training framework led to up to 47% absolute improvement in zero-shot transfer to out-of-distribution environments, without environment-specific fine-tuning.

Why it matters: This work advances scalable and steerable world modeling for RL agents, potentially reducing reliance on hand-crafted training environments and improving generalization.

Full story at: arXiv AI/ML