Scaling distributed reinforcement learning past the sync bottleneck
Running more environments in parallel does not scale RL training, it just leaves the learner waiting. This post covers the actor-learner split, using Redis to move trajectories between them, and how V-trace corrects for stale off-policy data.