vllm.triton_utils.dispatcher
¶
Registry for overriding Triton kernels with platform implementations.
Non-CUDA platforms can replace Triton kernels defined in vLLM core without modifying core code:
from vllm.triton_utils.dispatcher import register_kernels
register_kernels({
"vllm.v1.sample.rejection_sampler.expand_kernel": my_expand_impl,
"vllm.v1.worker.mamba_utils.batch_memcpy_kernel": my_memcpy_impl,
})
Classes:
-
KernelOverride–Stands in for a Triton kernel and routes launches to
impl.
Functions:
-
register_kernels–Register platform-specific implementations for Triton kernels.
KernelOverride
¶
Stands in for a Triton kernel and routes launches to impl.
The JIT warmup infrastructure launches kernels as
kernel[grid](**kwargs) where kwargs are keyed by the original
kernel's argument names. This wrapper forwards those launch arguments
to the platform implementation: by keyword when the implementation's
parameter names match the kernel's, otherwise positionally in the
kernel's parameter order.
Source code in vllm/triton_utils/dispatcher.py
_kernel_arg_names(kernel)
¶
Return the argument names of a Triton kernel or its fallback.
Source code in vllm/triton_utils/dispatcher.py
_rebind_kernels(overrides)
¶
Point every reference to each original kernel at its wrapper.
A kernel object can be captured in two kinds of places besides its defining module, and both must be patched for the overrides to take effect everywhere:
- plain attributes of other modules (
from mod import kernelcopies), rebound to the wrapper; - JIT warmup owners, whose
kernelattribute is the kernel they launch; swapped, and their cached_kernel_arg_namesinvalidated so the launch binding re-derives from the wrapper.
One pass over sys.modules handles all kernels in overrides.
Source code in vllm/triton_utils/dispatcher.py
_resolve_kernel(name)
¶
Return the (host object, attribute) holding the kernel name.
Source code in vllm/triton_utils/dispatcher.py
register_kernels(overrides)
¶
Register platform-specific implementations for Triton kernels.
Parameters:
-
(overrides¶Mapping[str, Callable[..., Any]]) –Mapping from fully qualified kernel name to the platform implementation. Kernel names look like "vllm.v1.sample.rejection_sampler.expand_kernel"; for kernels owned by a JIT warmup class, include the attribute path, e.g. "vllm.v1.worker.block_table.ComputeSlotMappingKernel.kernel". Implementations are invoked like the kernels themselves, with the launch arguments in the original kernel's parameter order (or by keyword when their parameter names match the kernel's).