vllm.v1.attention.ops.triton_prefill_attention
¶
Memory-efficient attention for prefill. It supports page size = 1.
Functions:
-
context_attention_fwd–q, k, v: [b * s, head, head_dim]
_prefer_narrow_kv_tile()
¶
RDNA3/RDNA4 prefer a narrower KV tile than the shared default.
Source code in vllm/v1/attention/ops/triton_prefill_attention.py
context_attention_fwd(q, k, v, o, b_start_loc, b_seq_len, max_input_len, is_causal=True, softmax_scale=None, sliding_window_q=None, sliding_window_k=None, sinks=None, sinks_bias_key0=False)
¶
q, k, v: [b * s, head, head_dim] b_start_loc: [b] b_seq_len: [b] out: [b * s, head, head_dim]