official deepseek-ai/DeepSeek-V2 repository, paper and config
DEEPSEEK / MLA ORIGIN / 2024
DeepSeek-V2
TOTAL236Bparameters
ACTIVE21Bper token
DEPTH60repeated blocks
HIDDEN5,120main trunk
CONTEXT128Knative tokens
CHECKPOINT≈472GB48 shards · DeepSeek License
INPUT100.0KTEXT
TRANSFORMER STACK60 DECODER LAYERSMLA in every layer · dense first layer · MoE thereafter
RMSNorm Residual
++
OUTPUT1.0KTOKENS
wraps every repeated block
MULTI-HEAD LATENT ATTENTION · DECOUPLED ROPE MLA内容 latent 与位置分量分开,是后续 DeepSeek 系列注意力的技术起点。
DeepSeekMoE · 236B CAPACITY21B active · total capacity ≠ active compute
+
DEPTH60decoder
HIDDEN5,120main trunk
KV RANK512MLA
EXPERTS160top-6
SHARED2experts
CONTEXT128Ktokens