vllm.v1.attention.ops.vit_attn_wrappers
¶
This file contains ops for ViT attention to be compatible with torch.compile
as there are operations here not supported by torch.compile (for instance,
.item() in flash attention)
Using these ops and wrapping vision blocks with torch.compile can speed up
throughput in vision models by ~5% relative on H100, and improve token
latencies by ~7% (see qwen2_5_vl for example usage)
To use these ops, you must have a recent version of PyTorch installed (>= 2.4.0)
Functions:
-
apply_sdpa–Input shape:
apply_sdpa(q, k, v, scale=None, enable_gqa=False)
¶
Input shape: (batch_size x seq_len x num_heads x head_size)