Skip to the document
Madhuopen lab
Outcome School · Transformers and architecture17 min read

by Amit Shekhar · 22 April 2026

Grouped Query Attention

Grouped-Query Attention (GQA) and how it differs from Multi-Head Attention (MHA).

3,292 words#llm#ai#machine-learning4 recall cards

Grouped Query Attention
Read on Outcome School ↗then come back to lock it in
Before you read, guess

What memory issue arises from storing Key and Value vectors for every head during text generation?

Ten seconds, a guess, then read — a wrong guess still makes the answer stick.

Before you go

In one sentence, what was this chapter about?

From memory, without scrolling up. Writing it is what makes it yours; the grade is only to show you what you had.

How sure?