official config + model card · 125 mixed FP8/BF16 shards
MINIMAX / AGENTIC MOE / 2026
MiniMax M2 Family
TOTAL229Bparameters
ACTIVE≈10Bper token
DEPTH62 + MTP3repeated blocks
HIDDEN3,072main trunk
CONTEXT196Knative tokens
CHECKPOINT230.1GB125 shards · MiniMax Community
M2family rootM2.1agent post-trainM2.5codingM2.7family representative
INPUT100.0KTEXT
TRANSFORMER STACK62 MAIN LAYERSdense GQA · 256 experts · top-8 · 3 MTP modules
pre-norm additive residual · narrower trunk than M1
++
OUTPUT1.0KTOKENS
wraps every repeated block
DENSE GQA · COMPACT ACTIVE PATHM2 用更窄主干和更细 expert bank,把 active path 从 M1 的 45.9B 降到约 10B。
MiniMax M2 MoE · 229B CAPACITY≈10B active path · total capacity ≠ active compute
DEPTH62main layers
HIDDEN3,07248 heads
GQA48 / 8Q / KV
EXPERTS256top-8
MTP3modules
CONTEXT196Ktokens