Why your GPU idles between inference batches
An L40S idling a full second between batches is a data transfer problem, not a compute problem. How multi-worker output processing, pre-allocated pinned buffers and dedicated CUDA streams took one PyTorch inference pipeline to about 4X throughput.