vllm.v1.attention.backends.mla.prefill.zen_cpu_sdpa
¶
CPU MLA prefill backend using zentorch SDPA on AMD Zen CPUs.
Subclasses the generic CPU SDPA backend and replaces the per-request attention
with zentorch_sdpa, which fuses the scale, causal mask, softmax and both
matmuls of the base implementation into a single call.
zentorch_sdpa does not return a usable log-sum-exp, so requests that need
one (context chunks, prefix/suffix merge) are delegated to the base backend.
Availability is reported through :meth:is_available, so the prefill selector
gates this backend once at selection time rather than per attention call.
Classes:
-
ZenCPUSDPAMLAPrefillBackend–MLA prefill backend for AMD Zen CPUs backed by
zentorch_sdpa.
ZenCPUSDPAMLAPrefillBackend
¶
Bases: CPUSDPAMLAPrefillBackend
MLA prefill backend for AMD Zen CPUs backed by zentorch_sdpa.
Source code in vllm/v1/attention/backends/mla/prefill/zen_cpu_sdpa.py
_sdpa_layout(q_r, k_r, v_r)
staticmethod
¶
Convert one request's [S, NH, D] tensors to the [1, NH, S, D] layout.
MLA has v_head_dim < qk_head_dim, which the fused kernel does not
accept, so V is padded up to the query head dim. The original V head
dim is returned so the caller can slice the padding off afterwards.