official config + model card · 414 BF16 shards
MINIMAX / HYBRID ATTENTION / 2025
MiniMax M1
TOTAL456Bparameters
ACTIVE45.9Bper token
DEPTH80repeated blocks
HIDDEN6,144main trunk
CONTEXT4M+native tokens
CHECKPOINT912.2GB414 shards · Apache-2.0
MiniMax-01linear attentionMiniMax M1Lightning + RLM2compact active pathM3MSA + vision
INPUT100.0KTEXT
TRANSFORMER STACK80 HYBRID LAYERS70 Lightning Attention · 10 full Softmax · every 8th global
pre-norm additive residual · linear state plus periodic global refresh
++
OUTPUT1.0KTOKENS
wraps every repeated block
LIGHTNING ⇄ SOFTMAX · LIGHTNING ATTENTIONLightning Attention 避免为大多数层构造完整 n×n attention matrix,使长输入和长推理链的成本更平缓。
MiniMax MoE · 456B CAPACITY45.9B active path · total capacity ≠ active compute
DEPTH8070 + 10
HIDDEN6,14464 heads
ATTN7 : 1linear / full
EXPERTS32top-2
CONTEXT4M+native family
CHECKPOINT912GBBF16