A preference model or comparable feedback signal is learned from comparisons, then an optimization algorithm updates the policy against that signal. Constraints are often used to keep the policy from drifting too far from the supervised model.
Reinforcement learning from human feedback uses human preferences to train a model toward responses people judge more desirable.
A preference model or comparable feedback signal is learned from comparisons, then an optimization algorithm updates the policy against that signal. Constraints are often used to keep the policy from drifting too far from the supervised model.