Skip to content

vllm.model_executor.models.k2_horizon

Inference-only K2Horizon model compatible with HuggingFace weights.

Functions:

_rope_weight_perm(head_dim, rope_head_dim)

Per-head gather index that folds K2Horizon's partial-RoPE channel reordering.

The K2Horizon checkpoint stores each head's channels in a GPT-J-interleaved convention with rope/nope channels interleaved. Applying this fixed (position-independent) permutation to the q/k weight rows lets the runtime use vLLM's native NeoX partial-RoPE path (head_size = head_dim, rotary_dim = rope_head_dim) with no per-forward permutes/all-gather.

Source code in vllm/model_executor/models/k2_horizon.py
def _rope_weight_perm(head_dim: int, rope_head_dim: int) -> torch.Tensor:
    """Per-head gather index that folds K2Horizon's partial-RoPE channel reordering.

    The K2Horizon checkpoint stores each head's channels in a GPT-J-interleaved
    convention with rope/nope channels interleaved. Applying this fixed
    (position-independent) permutation to the q/k weight rows lets the runtime
    use vLLM's *native* NeoX partial-RoPE path (``head_size = head_dim``,
    ``rotary_dim = rope_head_dim``) with no per-forward permutes/all-gather.
    """
    D, R = head_dim, rope_head_dim
    assert R % 2 == 0 and R <= D, (
        f"partial NeoX RoPE requires an even rope_head_dim <= head_dim, "
        f"got R={R}, D={D}"
    )
    h = D // 2
    idx = torch.cat(
        [
            torch.arange(0, R // 2),  # rope first-half channels
            torch.arange(h, h + R // 2),  # rope second-half channels
            torch.arange(R // 2, h),  # nope remainder (first half)
            torch.arange(h + R // 2, D),  # nope remainder (second half)
        ]
    )
    return idx

is_rope_weights_folding_supported(qk_proj, dual_chunk_attention_config)

Whether the partial-RoPE weight-fold path is safe for this attention.

Folding permutes the q/k projection's output channels at load time, which only produces correct numerics when:

  • the q/k projection is dense/unquantized and
  • dual-chunk attention is disabled -- the P.RoPE equivalence has not been validated with dual-chunk attention.

When either condition fails we fall back to the runtime channel-permutation RoPE path instead of folding.

Source code in vllm/model_executor/models/k2_horizon.py
def is_rope_weights_folding_supported(
    qk_proj: nn.Module,
    dual_chunk_attention_config: dict[str, Any] | None,
) -> bool:
    """Whether the partial-RoPE weight-fold path is safe for this attention.

    Folding permutes the q/k projection's output channels at load time, which
    only produces correct numerics when:

    * the q/k projection is dense/unquantized and
    * dual-chunk attention is disabled -- the P.RoPE equivalence has not been
      validated with dual-chunk attention.

    When either condition fails we fall back to the runtime channel-permutation
    RoPE path instead of folding.
    """
    is_dense_qk_proj = isinstance(
        getattr(qk_proj, "quant_method", None), UnquantizedLinearMethod
    )
    return is_dense_qk_proj and dual_chunk_attention_config is None