main
official config + technical report · 145 mixed-precision shards
XIAOMI MIMO / EFFICIENT AGENTIC / 2026

MiMo-V2-Flash

TOTAL309.8Bparameters
ACTIVE15Bper token
DEPTH48 + MTP3repeated blocks
HIDDEN4,096main trunk
CONTEXT256Knative tokens
CHECKPOINT313.1GB145 shards · MIT
INTERACTIVE SYSTEM BLUEPRINT

MODEL MRI

MiMo-V2-Flash · Hybrid Attention Explorer

100K INPUT + 1K OUTPUT · 4.93P arithmetic FLOPs · click any module for evidence

VERIFIED ADAPTER
INPUT100.0KTEXT
TRANSFORMER STACK48 MAIN LAYERS39 SWA · 9 full · 1 dense FFN + 47 MoE
pre-norm additive residual · attention sink on SWA layers
++
OUTPUT1.0KTOKENS
wraps every repeated block
SWA ⇄ FULL ATTENTION · ATTENTION SINK滑窗层保留 sink bias,使局部窗口仍能聚合到稳定的全局锚点,缓解纯窗口切断远距离信息。
MiMo MoE · 309.8B CAPACITY15B active path · total capacity ≠ active compute
DEPTH48main layers
HIDDEN4,09664 heads
ATTN5 : 1SWA / full
WINDOW128+ sink
EXPERTS256top-8
CONTEXT256Ktokens