Examples establish the task and output format without updating model weights. Selection, order, and representativeness of demonstrations can materially change the score and should remain fixed for comparison.
Few-shot evaluation measures performance when the model is given a small number of demonstrations in the prompt.
Examples establish the task and output format without updating model weights. Selection, order, and representativeness of demonstrations can materially change the score and should remain fixed for comparison.