RLPD: offline-to-online RL on locomotion and Humanoid
A three-person PyTorch reproduction and critical evaluation of RLPD (Ball et al., ICML 2023). Across the complete locomotion matrix, RLPD finishes at 88–90 normalized on all three tasks. On Humanoid-v5, only 6.6% of online states are covered by the offline dataset—and online-only beats the 50/50 mix by +21.9 points at the matched 500k horizon.
Online states covered by offline data
Humanoid is the outlier: sparse coverage, 6.64× normalized distance.






