Build the cache from one token upward: the two vectors that get stored, why a head is a view and head_dim is its width, what GQA and MLA trade away, and why a 128-token window is not cheap.
DeepSeek-V4.1 architecture. Chapter 1 ยท 24 min.
This page runs the interactive lesson in your browser: figures, checkpoints, and labs that compute their own numbers. Turn JavaScript on to use it.