Choosing a metric is choosing what you will optimize. BLEU misses meaning; accuracy ignores calibration; Elo depends on the opponent pool. Good eval stacks several metrics plus human review instead of betting the product on one number.
A quantitative score for model outputs — accuracy, F1, BLEU, pass@k, Elo — each with known blind spots.
Choosing a metric is choosing what you will optimize. BLEU misses meaning; accuracy ignores calibration; Elo depends on the opponent pool. Good eval stacks several metrics plus human review instead of betting the product on one number.