Glossary · Math & training

RLHF (Reinforcement Learning from Human Feedback)

A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal. Implementations vary and need not all use the same reinforcement-learning algorithm.

Common confusion

RLHF optimizes a proxy learned from collected feedback. It does not guarantee broad alignment with every user or situation.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.