vllm.model_executor.models.k2_horizon
¶
Inference-only K2Horizon model compatible with HuggingFace weights.
Functions:
-
is_rope_weights_folding_supported–Whether the partial-RoPE weight-fold path is safe for this attention.
_rope_weight_perm(head_dim, rope_head_dim)
¶
Per-head gather index that folds K2Horizon's partial-RoPE channel reordering.
The K2Horizon checkpoint stores each head's channels in a GPT-J-interleaved
convention with rope/nope channels interleaved. Applying this fixed
(position-independent) permutation to the q/k weight rows lets the runtime
use vLLM's native NeoX partial-RoPE path (head_size = head_dim,
rotary_dim = rope_head_dim) with no per-forward permutes/all-gather.
Source code in vllm/model_executor/models/k2_horizon.py
is_rope_weights_folding_supported(qk_proj, dual_chunk_attention_config)
¶
Whether the partial-RoPE weight-fold path is safe for this attention.
Folding permutes the q/k projection's output channels at load time, which only produces correct numerics when:
- the q/k projection is dense/unquantized and
- dual-chunk attention is disabled -- the P.RoPE equivalence has not been validated with dual-chunk attention.
When either condition fails we fall back to the runtime channel-permutation RoPE path instead of folding.