main
official config + model card · 163 FP8/BF16 shards
DEEPSEEK / SPARSE ATTENTION / 2025

DeepSeek-V3.2

TOTAL685Bparameters
ACTIVE37Bper token
DEPTH61+1repeated blocks
HIDDEN7,168main trunk
CONTEXT160Knative tokens
CHECKPOINT689.5GB163 shards
INTERACTIVE SYSTEM BLUEPRINT

MODEL MRI

DeepSeek-V3.2 · Sparse Attention Explorer

100K INPUT + 1K OUTPUT · 13.59P arithmetic FLOPs · click any module for evidence

VERIFIED ADAPTER
INPUT100.0KTOKENS
TRANSFORMER STACK61 MAIN LAYERS3 dense FFN · 58 sparse MoE · DSA throughout
pre-norm additive residual · no multi-stream expansion
++
OUTPUT1.0KTOKENS
wraps every repeated block
DEEPSEEK SPARSE ATTENTIONmain attention sparse; index scoring scans history
DEEPSEEK MoE · 685B CAPACITY37B active path · total capacity ≠ active compute
+
DEPTH61 + 1main + MTP
HIDDEN7,168main trunk
INDEX64 × 128top-2,048
EXPERTS256routed
ACTIVE8 + 1per token
PRECISIONFP8BF16 critical