vllm.model_executor.layers.fused_moe.experts
¶
Modules:
-
aiter_mxfp4_w4a16_moe– -
aiter_mxfp4_w4a8_moe– -
aiter_mxfp8_moe–MXFP8 (1x32 block, E8M0) MoE via AITER's FlyDSL two-stage grouped GEMM
-
batched_deep_gemm_moe– -
cpu_int4_moe–CPU INT4 W4A8 dynamic quantized fused MoE experts.
-
cpu_moe–CPU fused MoE experts.
-
cutlass_moe–CUTLASS based Fused MoE kernels.
-
deep_gemm_moe– -
fallback– -
flashinfer_b12x_moe– -
flashinfer_cutedsl_batched_moe– -
flashinfer_cutedsl_moe– -
flashinfer_cutlass_moe– -
flashinfer_moe_ep–FlashInfer MoE-EP megakernels as a modular-kernel experts implementation.
-
fused_batched_moe–Fused batched MoE kernel.
-
fused_humming_moe–Fused MoE utilities for Humming.
-
gpt_oss_triton_kernels_moe– -
int4_emulation_moe–Int4 weight-only quantization emulation for MoE.
-
lora_context– -
lora_experts_mixin– -
marlin_moe–Fused MoE utilities for GPTQ.
-
moonep_experts–MoonEP experts: grouped GEMM over MoonEP's expert-grouped activations.
-
mxfp8_emulation_moe–MXFP8 (1x32 block, E8M0 scale) MoE experts on Triton.
-
mxfp8_native_moe–Native MXFP8 (1x32 block, E8M0 scale) MoE for AMD CDNA4 (gfx950) via Triton
-
nvfp4_emulation_moe–NVFP4 quantization emulation for MoE.
-
ocp_mx_emulation_moe–OCP MX quantization emulation for MoE.
-
rdna3_moe–Fused MoE W4A16 experts on the RDNA3 (gfx1100) HIP kernel.
-
rocm_aiter_moe– -
triton_cutlass_moe– -
triton_deep_gemm_moe– -
triton_moe–Triton-based MoE expert implementations.
-
trtllm_bf16_moe– -
trtllm_fp8_moe– -
trtllm_lora_moe–LoRA-aware FlashInfer TRT-LLM MoE experts (BF16).
-
trtllm_mxfp4_moe– -
trtllm_mxint4_moe– -
trtllm_nvfp4_moe– -
xpu_moe– -
zentorch_moe–Zentorch fused MoE experts for AMD Zen CPUs.