RRLPD
PyTorch reproduction · MuJoCo / Minari-v5 · 2026

RLPD: offline-to-online reinforcement learning

We implemented RLPD in PyTorch, evaluated it against IQL and SACfD on three locomotion tasks, and extended the comparison to Humanoid-v5. Matched replay ablations test whether a fixed offline-data mixture helps online adaptation in that setting.

Research teamKaran AnchanPranav Prakash MenonKandi Sridhar
RLPD locomotion · n = 388.0–89.6normalized return at 245k steps
Matched Humanoid ablation · n = 3+22.0online-only minus 50/50 at 495k
Humanoid offline-state coverage6.6%post-hoc nearest-neighbor estimate
Core locomotion matrix27 runs3 methods × 3 tasks × 3 seeds
01Experimental setup

Data, replay, and optimization

The locomotion reproduction uses Minari-v5 medium datasets. RLPD combines each fixed dataset with online transitions; the critic ensemble, LayerNorm, and high update-to-data ratio follow the original method. The Humanoid-v5 extension tests the same design on a larger state space.

01Minari medium datafixed offline transitions
02Replay batch128 offline + 128 online
03Environment stepcollect one new transition
04Optimization20 critic updates + 1 actor update
01 / Replay mix50 / 50

Equal-source replay sampling

Each update samples 128 online and 128 offline transitions. Offline data therefore enter through replay composition rather than a separate RLPD pretraining phase.

ratio = 0.5 · batch = 256
Experimental roleFixes the per-update contribution of each replay source.
02 / Value controlLN

Layer-normalized critic

Layer normalization was included to mitigate critic-scale instability under distribution shift. In the Humanoid ablation without it, mean Q reached 8.9×10¹⁰ by 15k steps.

layernorm = true
Experimental roleProvides an explicit control on critic activation scale.
03 / Update pressure10 × 20

Critic ensemble and high UTD

The implementation uses ten critics and twenty gradient updates per environment step, increasing optimization effort per newly collected transition.

ensemble = 10 · utd = 20
Experimental roleDefines the optimization budget and ensemble estimator.
PyTorch 2.11Gymnasium MuJoCo v5Minari-v5 medium datasetsWeights & BiasesRTX 5070 · 12 GB

Normalized returns use this project's measured random-policy and Minari-v5 expert anchors (0 and 100), not the original paper's D4RL scale. Each evaluation averages ten deterministic episodes; locomotion results use the final 245k evaluation of a 250k-step budget. IQL receives one million offline updates before online step 0, so equal interaction counts do not mean equal optimization compute.

Implementation: training loop · replay sampling · state-coverage analysis.

02Locomotion results

RLPD on three MuJoCo tasks

At 245k online steps, mean normalized return was 88.0 on Hopper, 89.6 on Walker2d, and 88.6 on HalfCheetah. Each value is the mean across three seeds; RLPD exceeded the SACfD mean on all three tasks.

Hopper-v5normalized return · mean ± seed std
050100 · expert
RLPD
88.0 ± 6.8
IQL
65.6 ± 29.2
SACfD
41.9 ± 11.3

At 245k steps, the RLPD mean exceeds the IQL mean by 22.4 normalized-return points (n = 3 seeds per method).

Recorded RLPD rollout · hopper-v5
Qualitative example; the table reports three-seed aggregates.

The figure shows five-evaluation rolling means with seed-level uncertainty; the displayed endpoint scores come from unsmoothed evaluation CSVs. Recorded rollouts are qualitative examples, not estimates of average performance.

Critic-scale diagnostic

SACfD critic values grew sharply on Walker2d

At 245k steps, the recorded SACfD mean Q was approximately 85,300 versus 545 for RLPD. This 156× endpoint gap indicates value-scale instability in the tested SACfD path; it does not by itself explain the return difference.

SACfD endpoint
85,300
RLPD endpoint
545
Scale gap
156×
03Humanoid-v5 extension

Humanoid results

State-based Humanoid-v5 has 348 observation dimensions and 17 actions. At the 995k evaluation, IQL had the highest three-seed mean return. RLPD remained numerically stable but achieved a lower return; two of three SACfD seeds diverged before the full horizon.

Recorded rollout · selected IQL seed 2
87.8selected-seed last-five return; IQL three-seed mean: 70.1 ± 16.2
IQL · n = 3highest mean
70.1 ± 16.2
training context1M offline + 1M online
RLPD · n = 3finite estimates
13.0 ± 13.8
training contextno pretraining + 1M online
SACfD · n = 3divergent
2 / 3 NaN
training contextnon-finite runs retained

IQL received one million offline gradient updates before online step 0; the methods share an online interaction budget, not a compute budget. The rollout is a selected high-performing seed, not an estimate of expected performance. SACfD diverged in two seeds at 430k and 550k, so its reported endpoint summary combines unequal horizons.

04Matched replay ablation

Online-only versus 50/50 replay

At the 495k evaluation of a 500k-step Humanoid-v5 budget, we changed replay composition while keeping the critic architecture, ensemble, LayerNorm, update ratio, and batch size fixed. Each primary condition has three seeds.

RLPD · 50/505.9 ± 1.9

128 offline + 128 online per update

Online-only28.0 ± 15.2

256 online samples per update

No LayerNormdiverged

Seed 0 produced non-finite values by about 15k steps.

Offline-only replay−0.65

Seed-0 last-five return at the 495k evaluation.

UTD 1 instead of 203.26

Seed-0 last-five return versus 7.24 for the 50/50 reference.

2 critics instead of 104.61

Seed-0 last-five return versus 7.24 for the 50/50 reference.

Scope of inference

Online-only seed means were 23.11, 44.97, and 15.81; the matched 50/50 seed means were 7.24, 3.79, and 6.74. The resulting +22.0-point mean difference is descriptive given n = 3 and substantial variance. The replay-ratio figure is exploratory: most ratios have one seed, and its solid 90%-online point uses seed 1 while other solid points use seed 0. It is not a seed-uniform dose-response curve.

05State-distribution analysis

Offline-state coverage

We compared states collected in online replay buffers with each task's fixed offline dataset. PCA provides a two-dimensional view; the reported coverage metric uses all standardized state dimensions and a nearest-neighbor threshold derived from offline states.

Qualitative projection / seed 0

Humanoid online replay states separate in the offline-fitted PCA view

For each task, PCA is fitted only on standardized offline states. Offline and online states are then transformed through that same basis, so their relative position is directly comparable within a task.

Projection basisOffline fit

PCA is fitted on 20,000 standardized offline states for each environment; online states do not influence the axes.

Displayed sample4k + 4k

Four thousand offline and four thousand online states are displayed for the seed-0 RLPD run in each task.

Projected patternHumanoid separates

Locomotion clouds retain substantial overlap, whereas Humanoid online replay states are displaced from the main offline cloud.

Inference boundary2D is descriptive

The projection can hide variance in omitted components. Three-seed nearest-neighbor metrics in the full state space provide the primary evidence.

Full-space verification / three seeds

Across three seeds per task, the full-space metric counts online replay states inside the offline 95th-percentile nearest-neighbor radius. It samples states accumulated during training, not final-policy rollouts. The offline/offline distance control is 1.04–1.16.

06Reporting and reproducibility

Limitations

The results depend on the stated datasets, environment versions, compute budgets, and normalization anchors. Completed and interrupted runs are separated below; figure viewers retain the source plots at full resolution.

#Experiment groupCoverageReporting note
01Locomotion · medium27 completed runs3 algorithms × 3 tasks × 3 random seeds
02Locomotion · expertseed 0 completedseeds 1–2 terminate at 57.5k and are reported as incomplete
03Humanoid · IQL / RLPDn = 3 · 1M stepsIQL additionally receives one million offline updates before online training
04Humanoid · SACfD2 of 3 divergentnon-finite seeds stop at 430k and 550k; later curve values reflect the surviving seed
05Replay ablationn = 3 per conditiononline-only and 50/50 replay compared at the same 495k evaluation
06Expert Humanoid RLPDn = 1 · 1M stepssingle-seed result; excluded from aggregate conclusions
07State-distribution analysis4 tasks · n = 3PCA is descriptive; full-space nearest-neighbor metrics are primary

The research repository includes configuration, evaluation CSVs, code, and generated figures. It does not include trained checkpoints or raw replay buffers, so the full training and state-coverage results cannot be regenerated from that checkout alone. The project-specific Minari-v5 normalization also prevents direct numerical comparison with D4RL-normalized scores in the original paper.

Supplementary figures

Dataset quality and state coverage