Skip to content

vllm.model_executor.kernels.linear.mxfp8

Modules:

Classes:

Mxfp8LinearLayerConfig dataclass

Configuration for an MXFP8 linear layer.

All MXFP8 layers share the same structure: FP8-E4M3 weights with uint8 (E8M0) per-block scales at block size 32.

Attributes:

  • weight_shape (tuple[int, int]) –

    The layer's (out_features, in_features), i.e. (N, K).

Source code in vllm/model_executor/kernels/linear/mxfp8/Mxfp8LinearKernel.py
@dataclass
class Mxfp8LinearLayerConfig:
    """Configuration for an MXFP8 linear layer.

    All MXFP8 layers share the same structure: FP8-E4M3 weights with
    uint8 (E8M0) per-block scales at block size 32.

    Attributes:
        weight_shape: The layer's `(out_features, in_features)`, i.e. `(N, K)`.

    """

    weight_shape: tuple[int, int]
    bmm_batch_size: int | None = None