vllm.models.deepseek_v41.nvidia.ops.o_proj
¶
DSV4.1 output projection with a small-batch SM100/SM103 fusion.
Functions:
-
dsv41_o_proj–deep_gemm_fp8_o_projwith WO-A fused for small SM100/SM103 batches. -
register_dsv41_o_proj_warmup–Warm the fused WO-A token counts
dsv41_o_projcan dispatch.
_can_fuse_wo_a(layer)
¶
Whether the attention layer matches the fused WO-A kernel's layout.
Source code in vllm/models/deepseek_v41/nvidia/ops/o_proj.py
dsv41_o_proj(layer, attn_out, positions)
¶
deep_gemm_fp8_o_proj with WO-A fused for small SM100/SM103 batches.
Source code in vllm/models/deepseek_v41/nvidia/ops/o_proj.py
register_dsv41_o_proj_warmup(layer)
¶
Warm the fused WO-A token counts dsv41_o_proj can dispatch.