RoboTTT: Scaling Robot Policy Context to 8K Timesteps
Researchers present RoboTTT, a robot policy model that scales visuomotor context to 8,000 timesteps—three orders of magnitude beyond prior approaches—without increasing inference latency. RoboTTT enables new capabilities such as one-shot in-context imitation from human video, robust long-horizon task completion, and on-the-fly policy improvement. The model achieves an 87% performance improvement over single-step baselines and fully completes complex, multi-stage real-robot manipulation tasks.
Why it matters: This work establishes context length as a powerful new scaling axis for robot foundation models, unlocking significant advances in long-horizon and in-context robotic capabilities.
Full story at: arXiv Robotics ↗