Skip to the document
Madhuopen lab
Outcome School · Transformers and architecture15 min read

by Amit Shekhar · 17 August 2026

Decoding Deep RL from Human Preferences

In this blog, we are going to learn about Deep Reinforcement Learning from Human Preferences, the 2017 paper that started it all. It taught machines what we want by simply asking us to pick which of two behaviors looks better. This is the origin of RLHF, the technique behind ChatGPT.

2,867 words#ai#machine-learning4 recall cards

Decoding Deep RL from Human Preferences
Read on Outcome School ↗then come back to lock it in
Before you read, guess

Why is defining objectives as formulas problematic for complex tasks?

Ten seconds, a guess, then read — a wrong guess still makes the answer stick.

Before you go

In one sentence, what was this chapter about?

From memory, without scrolling up. Writing it is what makes it yours; the grade is only to show you what you had.

How sure?