AI clusters combine GPU nodes (each typically containing 4–8 GPUs connected via NVLink) with high-speed networking (InfiniBand or RoCE) and shared storage. Orchestration software handles job scheduling, fault tolerance, and checkpointing. Frontier training clusters like xAI's Colossus (100K H100s) and Meta's GPU farms represent the largest, while inference clusters optimize for throughput and latency rather than raw interconnect bandwidth.