vllm.model_executor.layers.fused_moe.routed_experts_capturer
¶
Classes:
-
RoutedExpertsCapturer–Worker-side capturer for routed experts, lives on GPU.
-
RoutedExpertsSink–Layer-owned buffer and callback for the expert ids a monolithic kernel
Functions:
-
bind_routed_experts_capturer–Attach capture callbacks to the target model's MoE routers.
RoutedExpertsCapturer
¶
Worker-side capturer for routed experts, lives on GPU.
Layer-level hooks call :meth:capture inside the forward pass. Routing
rows owned by this DP rank are written into a preallocated device buffer.
The device buffer uses int32. Stable snapshots use the narrowest dtype
that can represent every logical expert ID.
Invariants
- One instance per worker; shape is fixed at init and covers the
worst-case step (
max_num_batched_tokenstokens). - Every routed layer overwrites the current step's token rows.
Methods:
-
capture–Capture expert routing decisions for a specific layer.
-
snapshot_routing_data–Return a stable snapshot of the current routing data.
Source code in vllm/model_executor/layers/fused_moe/routed_experts_capturer.py
44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 | |
capture(layer_id, topk_ids)
¶
Capture expert routing decisions for a specific layer.
Under data parallelism, topk_ids may have four different batch
layouts depending on where the DP combine happens and whether
Expert Parallelism (EP) or Sequence Parallelism (SP) is active for the
MoE layer:
- n == total (naive dispatch): all DP ranks' tokens are
concatenated before routing; we slice out this rank's span
using the cumulative per-rank counts.
- n == token_num_per_dp (modular-kernel path): DP combine
happens inside quant_method.apply; select_experts only
ever sees this rank's tokens, so we take the whole tensor.
- n == sum(dp_metadata.local_sizes) (naive DP+EP dispatch):
sequence-parallel shards from every DP rank are gathered through
the flattened EP group. The shard sizes include CUDA-graph / SP
padding, so we use them to locate this DP rank's unpadded rows.
- n == ceil(token_num_per_dp / tp_size) (SP + modular-kernel
path): tokens were split along dim=0 across the TP group by
_sequence_parallel_context
(moe_runner_base.py:_sequence_parallel_context), so each
TP rank only sees its shard. We all-gather along dim=0 to
reconstruct this DP rank's full routing tensor. SP pads with
ceil-div (see _compute_sp_num_tokens in
forward_context.py), so the gathered tensor may contain a
few trailing padding rows which are trimmed by the downstream
[:token_num_per_dp] slice.
Parameters:
-
(layer_id¶int) –The layer index.
-
(topk_ids¶Tensor) –Tensor of shape (batch_size, num_routed_experts).
Source code in vllm/model_executor/layers/fused_moe/routed_experts_capturer.py
87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 | |
snapshot_routing_data(num_tokens)
¶
Return a stable snapshot of the current routing data.
RoutedExpertsSink
¶
Layer-owned buffer and callback for the expert ids a monolithic kernel routes to; it outlives kernel rebuilds on weight reload.
Source code in vllm/model_executor/layers/fused_moe/routed_experts_capturer.py
bind_routed_experts_capturer(model, capturer)
¶
Attach capture callbacks to the target model's MoE routers.