NVLink provides a direct GPU-to-GPU communication channel with bandwidth far exceeding PCIe — fifth-generation NVLink on Blackwell delivers 1.8 TB/s bidirectional per GPU. This bandwidth is essential for tensor parallelism, where a single layer's computation is split across GPUs and intermediate activations must be exchanged every forward pass. Without NVLink, the PCIe bottleneck makes multi-GPU training and inference significantly slower.