Distributing different layers of a model across different GPUs in sequence, with micro-batches flowing through the pipeline to keep all GPUs busy.
Distributing different model layers across GPUs in sequence with micro-batches flowing through.
Distributing different layers of a model across different GPUs in sequence, with micro-batches flowing through the pipeline to keep all GPUs busy.