Skip to the document
Madhuopen lab
Outcome School · Transformers and architecture8 min read

by Amit Shekhar · 1 September 2026

Decoding EAGLE

EAGLE, a state-of-the-art way to speed up language model generation by drafting tokens at the feature level instead of the token level.

1,516 words#llm#ai#machine-learning4 recall cards

Decoding EAGLE
Read on Outcome School ↗then come back to lock it in
Before you read, guess

How does the accuracy of token predictions impact the performance gains in speculative decoding?

Ten seconds, a guess, then read — a wrong guess still makes the answer stick.

What this article covers

  1. What is EAGLE
  2. A quick recap of speculative decoding
  3. The problem with token-level drafting
  4. The big idea: draft at the feature level
  5. Resolving the uncertainty by feeding back the token
  6. The math behind the speedup with small numbers
  7. EAGLE-2 and dynamic draft trees
  8. How EAGLE lives on today

The article lives on outcomeschool.com. Read it there, then come back: the tutor in the margin has read it and will answer questions, and the questions below check what stayed.

Before you go

In one sentence, what was this chapter about?

From memory, without scrolling up. Writing it is what makes it yours; the grade is only to show you what you had.

How sure?