Glossary · Math & training
RLHF (Reinforcement Learning from Human Feedback)
A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal. Implementations vary and need not all use the same reinforcement-learning algorithm.
Common confusion
RLHF optimizes a proxy learned from collected feedback. It does not guarantee broad alignment with every user or situation.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.