vllm.model_executor.layers.fused_moe.prepare_finalize.moonep
¶
MoonEP (https://github.com/MoonshotAI/MoonEP) prepare/finalize.
BF16 correctness-first proof of concept on top of vLLM's modular kernel interface.
MoonEP differs from DeepEP-style backends in two ways that shape this integration:
dispatchreturns tokens already grouped by expert row (a fixed[NvS, H]layout withNvS = S x Kreal slots plus padding) together with acu_seqlens[E+B]segment table and an opaqueplan. There is no per-token topk id tensor after dispatch; the expert compute must be a grouped GEMM overcu_seqlenssegments.- Rows
[E, E+B)of the weight/segment space are dynamic redundant-expert prefetch slots.plan.experts_to_copynames the source expert of each slot andBuffer.prefetch_weightmust run between dispatch and expert compute.
PoC limitations:
- BF16 / unquantized only, eager only.
- Expert weights are replicated in global-expert order on every rank
(memory-heavy). Production Kimi-K3 serving requires sharded
symmetric-memory expert ownership, where rows [0, E) physically alias
each home rank's parameter memory.
- Route weights are applied inside the expert compute and MoonEP's
combine performs the K-sum, so finalize requires
TopKWeightAndReduceNoOP.
Classes:
-
MoonEPBufferPool–Lazily creates
moonep.Bufferinstances at power-of-two capacities. -
MoonEPExpertWeightLayout–Contiguous BF16 expert weights in MoonEP
[E+B, ...]layout. -
MoonEPPrepareAndFinalize–Prepare/Finalize using MoonEP balanced dispatch/combine.
Functions:
-
gather_moonep_weight_layout–Build the replicated
[E+B, ...]layout from this rank's local experts. -
make_moonep_weight_layout–Build the replicated
[E+B, ...]PoC weight layout.
MoonEPBufferPool
¶
Lazily creates moonep.Buffer instances at power-of-two capacities.
Buffer fixes its token capacity S at construction, so a single
full-capacity buffer would pad every dispatch to
max_num_batched_tokens. The pool instead serves the smallest
power-of-two capacity that fits the step's token count, so decode-sized
batches dispatch a few hundred slots rather than the full capacity.
Selection must be identical on every EP rank (dispatch is collective and
Buffer construction is itself a collective): callers derive the token
count from dp_metadata (the max across DP ranks) so all ranks create
and pick the same buffer at the same step.
Methods:
-
get–Return
(capacity, buffer)for the given step token count.
Source code in vllm/model_executor/layers/fused_moe/prepare_finalize/moonep.py
get(num_tokens)
¶
Return (capacity, buffer) for the given step token count.
Source code in vllm/model_executor/layers/fused_moe/prepare_finalize/moonep.py
MoonEPExpertWeightLayout
¶
Bases: NamedTuple
Contiguous BF16 expert weights in MoonEP [E+B, ...] layout.
Rows [0, E) hold expert weights in global expert order; rows
[E, E+B) are mutable prefetch slots filled by
Buffer.prefetch_weight.
The prefetch slots never need re-zeroing between calls: the planner
marks unused slots with -1 in plan.experts_to_copy (skipped by
prefetch_weight) and gives them an empty cu_seqlens segment, so
stale slot contents are never read; used slots are fully overwritten.
Source code in vllm/model_executor/layers/fused_moe/prepare_finalize/moonep.py
MoonEPPrepareAndFinalize
¶
Bases: FusedMoEPrepareAndFinalizeModular
Prepare/Finalize using MoonEP balanced dispatch/combine.
prepare pads the batch to the selected buffer's static token
capacity, dispatches, runs prefetch_weight for the planned redundant
experts, and stashes the plan for finalize (the same pattern
DeepEP-HT uses for its handle). Downstream expert compute must consume
the expert-grouped [NvS, H] layout via cu_seqlens.
Attributes:
-
num_dispatched_slots(int) –NvSof the buffer selected by the current step's prepare().
Source code in vllm/model_executor/layers/fused_moe/prepare_finalize/moonep.py
206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 | |
num_dispatched_slots
property
¶
NvS of the buffer selected by the current step's prepare().
gather_moonep_weight_layout(w13_local, w2_local, num_global_experts, num_prefetch_slots)
¶
Build the replicated [E+B, ...] layout from this rank's local experts.
PoC bridge: each EP rank loads only its own experts (linear placement),
so all-gather them once at load time into global expert order on every
rank. Production MoonEP instead maps rows [0, E) onto each home
rank's parameter memory via symmetric memory (RFC #52095 item 5).
Source code in vllm/model_executor/layers/fused_moe/prepare_finalize/moonep.py
make_moonep_weight_layout(w13_weight, w2_weight, num_prefetch_slots)
¶
Build the replicated [E+B, ...] PoC weight layout.
w13_weight must be [E, 2I, H] (gate rows first) and w2_weight
[E, H, I], both BF16 in global expert order.