vllm.model_executor.kernels.linear.scaled_mm.zentorch
¶
Zentorch dynamic-symmetric W8A8 int8 linear kernel for AMD Zen CPUs.
Selected by choose_scaled_mm_linear_kernel ahead of the generic
oneDNN-backed CPUInt8ScaledMMLinearKernel. When is_supported or
can_implement rejects a layer, the selector falls through to the next
kernel in _POSSIBLE_INT8_KERNELS[PlatformEnum.CPU].
Classes:
ZentorchInt8ScaledMMLinearKernel
¶
Bases: Int8ScaledMMLinearKernel
Methods:
-
process_weights_after_loading–Prepare weights for
zentorch_dynamic_qlinear.
Source code in vllm/model_executor/kernels/linear/scaled_mm/zentorch.py
process_weights_after_loading(layer)
¶
Prepare weights for zentorch_dynamic_qlinear.
Keeps weight in [N, K] layout (int8, contiguous) and converts the
per-channel weight scale to bf16 with shape (N,).