by Amit Shekhar · 13 August 2026
Decoding InstructGPT
In this blog, we are going to learn about InstructGPT, the model that taught GPT-3 to actually follow our instructions, and the work that led directly to ChatGPT.
Read on Outcome School ↗then come back to lock it in
Before you read, guessWhat is the core mismatch in InstructGPT's alignment problem?
Ten seconds, a guess, then read — a wrong guess still makes the answer stick.
What this article covers
- What is the InstructGPT paper?
- The building blocks we must know first
- The big picture: what InstructGPT does
- Why GPT-3 was not enough
- Helpful, Honest, and Harmless
- The three-step method
- Step 1: Supervised Fine-Tuning
- Step 2: The Reward Model
- Step 3: Reinforcement Learning with PPO
- The alignment tax
- The results
- What alignment looks like today
The article lives on outcomeschool.com. Read it there, then come back: the tutor in the margin has read it and will answer questions, and the questions below check what stayed.
