Skip to the document
Madhuopen lab
Outcome School · Transformers and architecture19 min read

by Amit Shekhar · 24 April 2026

Decoding DeepSeek-V4

DeepSeek-V4, the new family of open Mixture-of-Experts language models that natively supports a one-million-token context with dramatically lower inference cost.

3,785 words#llm#ai#machine-learning4 recall cards

Decoding DeepSeek-V4
Read on Outcome School ↗then come back to lock it in
Before you read, guess

What mechanism expands the residual stream and restricts the mixing matrix to ensure a spectral norm of 1 for stable deep training?

Ten seconds, a guess, then read — a wrong guess still makes the answer stick.

What this article covers

  1. The Big Picture
  2. Two Models: DeepSeek-V4-Pro and DeepSeek-V4-Flash
  3. Hybrid Attention with CSA and HCA
  4. Manifold-Constrained Hyper-Connections (mHC)
  5. Muon Optimizer
  6. FP4 Quantization-Aware Training
  7. Pre-Training
  8. Post-Training: Specialist Training and On-Policy Distillation
  9. Reasoning Modes
  10. Putting It All Together

The article lives on outcomeschool.com. Read it there, then come back: the tutor in the margin has read it and will answer questions, and the questions below check what stayed.

Before you go

In one sentence, what was this chapter about?

From memory, without scrolling up. Writing it is what makes it yours; the grade is only to show you what you had.

How sure?