Skip to the document
Madhuopen lab
Outcome School · Training and alignment13 min read

by Amit Shekhar · 17 May 2026

Direct Preference Optimization (DPO)

In this blog, we are going to learn about Direct Preference Optimization (DPO). We will also see how DPO works step-by-step and how it differs from RLHF (PPO).

2,478 words#llm#ai#machine-learning4 recall cards

Direct Preference Optimization (DPO)
Read on Outcome School ↗then come back to lock it in
Before you read, guess

What method is described as simple and powerful for training LLMs on human preferences?

Ten seconds, a guess, then read — a wrong guess still makes the answer stick.

What this article covers

  1. What is RLHF and Why Do We Need It?
  2. The Problem with RLHF
  3. What is Direct Preference Optimization (DPO)?
  4. What is Preference Data?
  5. The Key Idea Behind DPO
  6. The DPO Loss Function in Simple Words
  7. How DPO Works Step-by-Step
  8. DPO vs RLHF (PPO)
  9. Advantages of DPO
  10. Disadvantages of DPO

The article lives on outcomeschool.com. Read it there, then come back: the tutor in the margin has read it and will answer questions, and the questions below check what stayed.

Before you go

In one sentence, what was this chapter about?

From memory, without scrolling up. Writing it is what makes it yours; the grade is only to show you what you had.

How sure?