Skip to content

vllm.models.hy_v4.nvidia

Modules:

  • attention –

    MLA attention and lightning indexer for HY V4 (NVIDIA).

  • flashmla_sparse –

    Sink-capable FlashMLA sparse backend for HY V4 (NVIDIA).

  • hc –

    iHC (independent Hyper-Connections) layers for HY V4 (NVIDIA).

  • model –

    Inference-only HY V4 model compatible with HuggingFace weights (NVIDIA).

  • moe –

    Dense FFN and MoE blocks for HY V4 (NVIDIA).

  • mtp –

    Multi-token prediction (MTP) head for HY V4 (NVIDIA).

  • triton_ihc –

    Triton iHC pre/post kernels for HY V4.