vllm.v1.attention.backends
¶
Modules:
-
b12x–b12x paged causal attention backend for SM12x.
-
composite–Reusable attention backend composition and routing policies.
-
cpu_attn– -
fa_utils– -
flash_attn–Attention layer with FlashAttention.
-
flash_attn_diffkv–Attention layer with FlashAttention.
-
flashinfer–Attention layer with FlashInfer.
-
flex_attention–Attention layer with FlexAttention.
-
gdn_attn–Backend for GatedDeltaNet attention.
-
hpc_attn–HPC Attention Backend.
-
mamba2_attn– -
mamba_attn– -
mla– -
recoverssm_metadata– -
registry–Attention backend registry.
-
rocm_aiter_fa–Attention layer with AiterFlashAttention.
-
rocm_aiter_unified_attn–Attention layer with PagedAttention and Triton prefix prefill.
-
rocm_attn–Attention layer with PagedAttention and Triton prefix prefill.
-
short_conv_attn– -
triton_attn–High-Performance Triton-only Attention layer.
-
triton_attn_diffkv–Triton attention backend with different K/V head dimensions (DiffKV).
-
triton_flash_attn–Triton image-mask and FlashAttention causal attention composite.
-
triton_flashinfer–Triton image-mask and FlashInfer causal attention composite.
-
turboquant_attn–TurboQuant attention backend for vLLM.
-
utils– -
zentorch_sdpa–Encoder attention through the zentorch SDPA kernel on Zen CPUs.