The image is divided into patches, projected into embeddings, and combined with positional information. Self-attention then models relationships between patches, and the resulting representations support classification or feed downstream vision tasks.