Equal-source replay sampling
Each update samples 128 online and 128 offline transitions. Offline data therefore enter through replay composition rather than a separate RLPD pretraining phase.
ratio = 0.5 · batch = 256We implemented RLPD in PyTorch, evaluated it against IQL and SACfD on three locomotion tasks, and extended the comparison to Humanoid-v5. Matched replay ablations test whether a fixed offline-data mixture helps online adaptation in that setting.
The locomotion reproduction uses Minari-v5 medium datasets. RLPD combines each fixed dataset with online transitions; the critic ensemble, LayerNorm, and high update-to-data ratio follow the original method. The Humanoid-v5 extension tests the same design on a larger state space.
Each update samples 128 online and 128 offline transitions. Offline data therefore enter through replay composition rather than a separate RLPD pretraining phase.
ratio = 0.5 · batch = 256Layer normalization was included to mitigate critic-scale instability under distribution shift. In the Humanoid ablation without it, mean Q reached 8.9×10¹⁰ by 15k steps.
layernorm = trueThe implementation uses ten critics and twenty gradient updates per environment step, increasing optimization effort per newly collected transition.
ensemble = 10 · utd = 20Normalized returns use this project's measured random-policy and Minari-v5 expert anchors (0 and 100), not the original paper's D4RL scale. Each evaluation averages ten deterministic episodes; locomotion results use the final 245k evaluation of a 250k-step budget. IQL receives one million offline updates before online step 0, so equal interaction counts do not mean equal optimization compute.
Implementation: training loop · replay sampling · state-coverage analysis.
At 245k online steps, mean normalized return was 88.0 on Hopper, 89.6 on Walker2d, and 88.6 on HalfCheetah. Each value is the mean across three seeds; RLPD exceeded the SACfD mean on all three tasks.
At 245k steps, the RLPD mean exceeds the IQL mean by 22.4 normalized-return points (n = 3 seeds per method).
The figure shows five-evaluation rolling means with seed-level uncertainty; the displayed endpoint scores come from unsmoothed evaluation CSVs. Recorded rollouts are qualitative examples, not estimates of average performance.
At 245k steps, the recorded SACfD mean Q was approximately 85,300 versus 545 for RLPD. This 156× endpoint gap indicates value-scale instability in the tested SACfD path; it does not by itself explain the return difference.
State-based Humanoid-v5 has 348 observation dimensions and 17 actions. At the 995k evaluation, IQL had the highest three-seed mean return. RLPD remained numerically stable but achieved a lower return; two of three SACfD seeds diverged before the full horizon.
IQL received one million offline gradient updates before online step 0; the methods share an online interaction budget, not a compute budget. The rollout is a selected high-performing seed, not an estimate of expected performance. SACfD diverged in two seeds at 430k and 550k, so its reported endpoint summary combines unequal horizons.
At the 495k evaluation of a 500k-step Humanoid-v5 budget, we changed replay composition while keeping the critic architecture, ensemble, LayerNorm, update ratio, and batch size fixed. Each primary condition has three seeds.
128 offline + 128 online per update
256 online samples per update
Seed 0 produced non-finite values by about 15k steps.
Seed-0 last-five return at the 495k evaluation.
Seed-0 last-five return versus 7.24 for the 50/50 reference.
Seed-0 last-five return versus 7.24 for the 50/50 reference.
Online-only seed means were 23.11, 44.97, and 15.81; the matched 50/50 seed means were 7.24, 3.79, and 6.74. The resulting +22.0-point mean difference is descriptive given n = 3 and substantial variance. The replay-ratio figure is exploratory: most ratios have one seed, and its solid 90%-online point uses seed 1 while other solid points use seed 0. It is not a seed-uniform dose-response curve.
We compared states collected in online replay buffers with each task's fixed offline dataset. PCA provides a two-dimensional view; the reported coverage metric uses all standardized state dimensions and a nearest-neighbor threshold derived from offline states.
For each task, PCA is fitted only on standardized offline states. Offline and online states are then transformed through that same basis, so their relative position is directly comparable within a task.
PCA is fitted on 20,000 standardized offline states for each environment; online states do not influence the axes.
Four thousand offline and four thousand online states are displayed for the seed-0 RLPD run in each task.
Locomotion clouds retain substantial overlap, whereas Humanoid online replay states are displaced from the main offline cloud.
The projection can hide variance in omitted components. Three-seed nearest-neighbor metrics in the full state space provide the primary evidence.
Across three seeds per task, the full-space metric counts online replay states inside the offline 95th-percentile nearest-neighbor radius. It samples states accumulated during training, not final-policy rollouts. The offline/offline distance control is 1.04–1.16.
The results depend on the stated datasets, environment versions, compute budgets, and normalization anchors. Completed and interrupted runs are separated below; figure viewers retain the source plots at full resolution.
The research repository includes configuration, evaluation CSVs, code, and generated figures. It does not include trained checkpoints or raw replay buffers, so the full training and state-coverage results cannot be regenerated from that checkout alone. The project-specific Minari-v5 normalization also prevents direct numerical comparison with D4RL-normalized scores in the original paper.