>

Clicked Gallery

What is RLHF (Reinforcement Learning from Human Feedback)?

Highlighted from a real engineering doc. Explained by Clicked.

Used in a sentence

Engineering Notes · AI Systems

After pre-training, the model was aligned using RLHF, with human raters comparing candidate responses.

The reader highlighted one word in the docs. Clicked explained the technical term “RLHF” in simple terms:

Explained in three depths

Same facts, different vibe — Slang mode 😎

Formal definition — The same term, explained the usual way

Reinforcement learning from human feedback is a post-training alignment procedure in which human annotators express preferences between candidate model outputs, those preferences are used to fit a reward model approximating human judgment, and the base model is then optimized against that reward signal. It substantially improves instruction-following and response quality relative to pre-training alone, while introducing failure modes including sycophancy, over-refusal, and the encoding of annotator bias.

Want Clicked to explain terms like “RLHF” directly in your browser — including on PDFs?

Add to Chrome — Free

50 free Explanations · No credit card required