official Mistral Mixtral model card, blog and config
MISTRAL / CLASSIC SMOE / 2023
Mixtral 8×7B
TOTAL46.7Bparameters
ACTIVE12.9Bper token
DEPTH32repeated blocks
HIDDEN4,096main trunk
CONTEXT32Knative tokens
CHECKPOINT≈93.4GB19 shards · Apache-2.0
INPUT31.8KTEXT
TRANSFORMER STACK32 SMOE BLOCKSGQA / sliding window + top-2 expert FFN
RMSNorm Residual
++
OUTPUT1.0KTOKENS
wraps every repeated block
SLIDING-WINDOW GQA · COARSE-GRAINED TOP-2每个专家都是完整大 FFN,结构直观但粒度比 DeepSeekMoE 粗。
Sparse MoE · 46.7B CAPACITY12.9B active · total capacity ≠ active compute
DEPTH32decoder
HIDDEN4,096main trunk
HEADS32 / 8Q / KV
EXPERTS8top-2
CONTEXT32Ktokens
ACTIVE12.9Bper token