vllm.model_executor.layers.rotary_embedding.packed_qk_rope
¶
In-place interleaved RoPE on the Q/K slices of a packed QKV buffer.
Functions:
-
packed_qk_rope_–Rotate Q and K in place inside a packed QKV buffer.
packed_qk_rope_(xqkv, freqs_cis)
¶
Rotate Q and K in place inside a packed QKV buffer.
Parameters:
-
(xqkv¶Tensor) –contiguous (seqlen, 3, nheads, headdim); only the Q and K slices are rotated.
-
(freqs_cis¶Tensor) –contiguous (seqlen, headdim // 2) complex64 rotary freqs, read through its interleaved re/im fp32 view.
Matches ApplyRotaryEmb(enable_fp32_compute=True) applied to Q and K
separately to within one ULP, but in one kernel with ~5x less memory
traffic. Both compute in fp32; the gap is which product the compiler
keeps exact inside the fused multiply-add, which differs per backend.
Requires triton: callers must check HAS_TRITON and use the unfused path
otherwise. Contract violations raise.