Skip to content

vllm.models.qwen4_exp.nvidia.low_latency_gemm

Qwen4Exp decode GEMM selection on Hopper and Blackwell.

Dispatch follows Kimi-K3 and uses the local (N, K) shape and token count. Plans contain measured CUDA graph capture sizes; other token counts use the standard linear implementation.