Skip to the document
Madhuopen lab
Outcome School · Training and alignment19 min read

by Amit Shekhar · 14 May 2026

Reinforcement Learning from Human Feedback (RLHF)

Reinforcement Learning from Human Feedback (RLHF), the training technique that turns a raw pre-trained LLM into a helpful, honest, and safe assistant by teaching it from human preferences.

3,780 words#llm#ai#machine-learning4 recall cards

Reinforcement Learning from Human Feedback (RLHF)
Read on Outcome School ↗then come back to lock it in
Before you read, guess

What three components make up RLHF?

Ten seconds, a guess, then read — a wrong guess still makes the answer stick.

What this article covers

  1. What is RLHF
  2. Why we need RLHF
  3. The Big Picture
  4. Stage 1: Supervised Fine-Tuning (SFT)
  5. Stage 2: Training the Reward Model
  6. Stage 3: RL Fine-Tuning with PPO
  7. The KL Penalty
  8. Putting It All Together
  9. Reward Hacking
  10. Common Mistakes
  11. Best Practices

The article lives on outcomeschool.com. Read it there, then come back: the tutor in the margin has read it and will answer questions, and the questions below check what stayed.

Before you go

In one sentence, what was this chapter about?

From memory, without scrolling up. Writing it is what makes it yours; the grade is only to show you what you had.

How sure?