The model predicts each target output under teacher forcing and receives a token-level or task-specific loss. Curated demonstrations can teach instruction following, formatting, domain behavior, or a particular task.
Supervised fine-tuning trains a pretrained model on labeled input-output examples that demonstrate desired responses.
The model predicts each target output under teacher forcing and receives a token-level or task-specific loss. Curated demonstrations can teach instruction following, formatting, domain behavior, or a particular task.