It represents inputs as vectors, mixes information across positions with attention, and transforms each position with feed-forward layers. Stacking these blocks lets the model learn contextual representations that can be used for generation, classification, and other sequence tasks.