vllm.model_executor.models.mimo_v2_mtp
¶
Inference-only MiMo-V2 MTP (Multi-Token Prediction) draft model.
Supports both MiMo-V2-Pro and MiMo-V2-Flash checkpoints.
Checkpoint weight layout (model.mtp.layers.{idx}.): enorm - RMSNorm for token embeddings hnorm - RMSNorm for previous hidden states eh_proj - ReplicatedLinear(hidden2 -> hidden) input_layernorm - pre-attention RMSNorm self_attn. - attention weights; format differs by variant: Pro: fused qkv_proj [Q;K;V] concatenated Flash: separate q_proj, k_proj, v_proj pre_mlp_layernorm - post-attention / pre-MLP RMSNorm mlp. - dense MLP (gate_proj / up_proj / down_proj) final_layernorm - norm applied before logit computation
Classes:
-
MiMoV2MTPLayer–Single MTP predictor layer for MiMo-V2 (Pro and Flash).
MiMoV2MTPLayer
¶
Bases: Module
Single MTP predictor layer for MiMo-V2 (Pro and Flash).
Mirrors the single-layer MiMo-V2 nextn reference implementation.
Source code in vllm/model_executor/models/mimo_v2_mtp.py
_MiMoV2MTPLayers
¶
Bases: Module
Thin wrapper so parameter paths match checkpoint: model.mtp.layers.*.