Separate encoders may map each modality into compatible representations before fusion, or one shared model may tokenize several modalities directly. Cross-attention, projection layers, or shared latent spaces allow information from one modality to condition another.