Encoders or shared token representations map modalities into spaces where information can be combined. Training on aligned examples lets the model connect concepts across modalities and condition one form of output on another.
A multimodal model can process or generate more than one kind of data, such as text, images, audio, or video.
Encoders or shared token representations map modalities into spaces where information can be combined. Training on aligned examples lets the model connect concepts across modalities and condition one form of output on another.