Pioneered by Cerebras, wafer-scale chips keep data on a single piece of silicon with on-die interconnect rather than routing it across separate chips on a board or network, which cuts communication latency by orders of magnitude versus conventional GPU clusters. Cerebras's WSE-3 Turbo, used in three-wafer configurations in the CS-4 system announced August 19, 2026, is the current generation; Cerebras has reported inference speeds up to 30x faster than GPU-based systems on models exceeding 10 trillion parameters.