PyTorch distributed operations: where multi-GPU training stalls
Notes on PyTorch distributed operations: what NCCL is doing under the hood, why synchronous and asynchronous do not mean what you would guess, and how a missing wait call hangs a rank forever.