Skip to the document
Madhuopen lab
Outcome School · Transformers and architecture16 min read

by Amit Shekhar · 5 April 2026

Math behind √dₖ Scaling Factor in Attention

Why we scale the dot product attention by √dₖ in the Transformer architecture with a step-by-step numeric example.

3,059 words#math#llm#ai#machine-learning4 recall cards

Math behind √dₖ Scaling Factor in Attention
Read on Outcome School ↗then come back to lock it in
Before you read, guess

What characteristic do the outputs of the softmax function exhibit?

Ten seconds, a guess, then read — a wrong guess still makes the answer stick.

Before you go

In one sentence, what was this chapter about?

From memory, without scrolling up. Writing it is what makes it yours; the grade is only to show you what you had.

How sure?