by Amit Shekhar · 5 April 2026
Math behind √dₖ Scaling Factor in Attention
Why we scale the dot product attention by √dₖ in the Transformer architecture with a step-by-step numeric example.
Read on Outcome School ↗then come back to lock it in
Before you read, guessWhat characteristic do the outputs of the softmax function exhibit?
Ten seconds, a guess, then read — a wrong guess still makes the answer stick.
What this article covers
- The Attention Formula (Quick Recap)
- What Happens Without Scaling?
- Why Do Dot Products Grow with dₖ?
- Understanding Variance of the Dot Product
- Proving It Step by Step: Variance of the Dot Product is dₖ
- What Large Dot Products Do to Softmax
- Why √dₖ is the Right Scaling Factor
- Seeing It with Real Numbers
- Putting It All Together
The article lives on outcomeschool.com. Read it there, then come back: the tutor in the margin has read it and will answer questions, and the questions below check what stayed.
