official FunAudioLLM/SenseVoice repository and SenseVoiceSmall config.yaml
FUNASR / NON-AUTOREGRESSIVE RICH ASR / 2024
SenseVoice-Small
TOTAL≈234Mparameters
ACTIVE≈234Mruntime path
DEPTH50 + 20repeated blocks
HIDDEN512main trunk
CONTEXTlong audiostream / sequence
CHECKPOINT≈930MB1 shards · FunASR Model License
INPUT30SECONDS
TRANSFORMER STACK50L SAN-M + 20L TPencoder-only recognition with temporal predictor
Encoder Residual
++
OUTPUT256TOKENS
wraps every repeated block
SAN-M SELF-ATTENTION · RICH TRANSCRIPTION TOKENS同一次非自回归识别输出语言、情绪和环境事件标记。
Temporal Predictor不需要逐 token decoder,输出延迟主要由 encoder forward 决定。
ENCODER50LSAN-M
TP20Lpredictor
HIDDEN512output
HEADS4attention
FFN2,048dim
PARAMS≈234Msmall