Skip to content

vllm.model_executor.layers.quantization.utils.humming

Humming quantization integration.

Modules:

  • activation –

    MoE activation expressions for Humming input processing.

  • linear –

    Prepare and execute Humming linear layers.

  • moe –

    Configure, prepare weights for, and assemble Humming MoE kernels.

  • schema –

    Map Humming schemas and handle shared checkpoint quantization settings.

Functions:

convert_linear_layer_to_humming_standard(layer, name_map)

Rename/reshape a linear layer's quantized params (the canonical MPLinear layout: weight_packed int32 + weight_scale) into the parameter names and layout humming's weight schema expects (weight / weight_scale).

Source code in vllm/model_executor/layers/quantization/utils/humming/linear.py
def convert_linear_layer_to_humming_standard(
    layer: LinearBase, name_map: dict[str, str]
):
    """Rename/reshape a linear layer's quantized params (the canonical MPLinear
    layout: ``weight_packed`` int32 + ``weight_scale``) into the parameter names
    and layout humming's weight schema expects (``weight`` / ``weight_scale``)."""
    for name, checkpoint_name in name_map.items():
        tensor = getattr(layer, checkpoint_name)
        delattr(layer, checkpoint_name)

        if name == "weight":
            input_dim = getattr(tensor, "input_dim", 1)
            output_dim = getattr(tensor, "output_dim", 0)

            if input_dim == 0 and output_dim == 1:
                tensor = tensor.transpose(1, 0).contiguous()
            else:
                assert output_dim == 0 and input_dim == 1

            tensor = tensor.view(tensor.size(0), -1).view(torch.int32)
        elif name in ["weight_scale", "zero_point"]:
            if getattr(tensor, "output_dim", 0) == 1:
                tensor = tensor.transpose(0, 1).contiguous()
            if tensor.ndim == 1:
                tensor = tensor.unsqueeze(1)

            tensor = tensor.view(torch.int32) if name == "zero_point" else tensor

        if isinstance(tensor, torch.nn.Parameter):
            param = tensor
        else:
            param = torch.nn.Parameter(tensor, requires_grad=False)

        setattr(layer, name, param)

convert_to_humming_moe_kernel_format(layer, quant_config=None, sublayer_configs=None, weight_schema=None, input_schema=None, force_weight_schema=None, allow_input_schema_fallback=True)

Convert MoE weights from checkpoint format to Humming kernel format.

This function processes weights for each sublayer (w13, w2) by: 1. Converting from checkpoint format to humming format if needed 2. Force requanting if a different quantization schema is specified 3. Preparing layer metadata for the Humming kernel 4. Transforming weights for inference

Parameters:

  • layer

    (RoutedExperts) –

    The RoutedExperts layer containing weights to process

  • quant_config

    (dict | None, default: None ) –

    Optional quantization config dict. Required if weight_schema or input_schema are None. Used to build schemas via BaseWeightSchema.from_config().

  • sublayer_configs

    (dict[str, Any] | None, default: None ) –

    Optional configuration dict for each sublayer (w13, w2). Each config must have "shape_n" and "shape_k" keys. If None, configs are built from layer.moe_config properties.

  • weight_schema

    (Any | None, default: None ) –

    Optional initial weight quantization schema. If None, built from quant_config.

  • input_schema

    (Any | None, default: None ) –

    Optional initial input quantization schema. If None, built from quant_config or env vars.

  • force_weight_schema

    (Any | None, default: None ) –

    Optional schema to force requantization to

  • allow_input_schema_fallback

    (bool, default: True ) –

    Whether incompatible input schemas may be replaced.

Side effects
  • Modifies layer parameters in place
  • Sets layer.weight_schemas and layer.input_schemas
  • Sets layer.humming_configs for quant config construction
Source code in vllm/model_executor/layers/quantization/utils/humming/moe.py
def convert_to_humming_moe_kernel_format(
    layer: "RoutedExperts",
    quant_config: dict | None = None,
    sublayer_configs: dict[str, Any] | None = None,
    weight_schema: Any | None = None,
    input_schema: Any | None = None,
    force_weight_schema: Any | None = None,
    allow_input_schema_fallback: bool = True,
) -> dict[str, "LayerConfig"]:
    """Convert MoE weights from checkpoint format to Humming kernel format.

    This function processes weights for each sublayer (w13, w2) by:
    1. Converting from checkpoint format to humming format if needed
    2. Force requanting if a different quantization schema is specified
    3. Preparing layer metadata for the Humming kernel
    4. Transforming weights for inference

    Args:
        layer: The RoutedExperts layer containing weights to process
        quant_config: Optional quantization config dict. Required if weight_schema
                     or input_schema are None. Used to build schemas via
                     BaseWeightSchema.from_config().
        sublayer_configs: Optional configuration dict for each sublayer (w13, w2).
                         Each config must have "shape_n" and "shape_k" keys.
                         If None, configs are built from layer.moe_config properties.
        weight_schema: Optional initial weight quantization schema.
                      If None, built from quant_config.
        input_schema: Optional initial input quantization schema.
                     If None, built from quant_config or env vars.
        force_weight_schema: Optional schema to force requantization to
        allow_input_schema_fallback: Whether incompatible input schemas may be replaced.

    Side effects:
        - Modifies layer parameters in place
        - Sets layer.weight_schemas and layer.input_schemas
        - Sets layer.humming_configs for quant config construction

    """
    # Build schemas from quant_config if not provided
    has_bias = layer.moe_config.has_bias
    num_experts = layer.moe_config.num_local_experts
    param_dtype = layer.params_dtype

    if weight_schema is None or input_schema is None:
        if quant_config is None:
            raise ValueError(
                "Must provide either weight_schema/input_schema or quant_config"
            )

        from vllm.utils.humming import BaseWeightSchema, HummingInputSchema

        if weight_schema is None:
            weight_schema = BaseWeightSchema.from_config(quant_config)

        if input_schema is None:
            input_quant_config = (envs.VLLM_HUMMING_INPUT_QUANT_CONFIG or {}).copy()
            if humming_is_layer_skipped(input_quant_config, layer.layer_name):
                input_schema = HummingInputSchema()
            else:
                # TODO: read input_quant_config from quant_config
                input_quant_config = humming_schema.resolve_humming_layer_config(
                    input_quant_config, layer.layer_name
                )
                allow_input_schema_fallback = input_quant_config.pop(
                    "allow_fallback", False
                )
                input_schema = HummingInputSchema.from_config(input_quant_config)

    # Build sublayer configs from layer properties if not provided
    if sublayer_configs is None:
        is_gated = layer.moe_config.activation.is_gated
        intermediate_size = layer.moe_config.intermediate_size_per_partition
        sublayer_configs = {
            "w13": {
                "shape_n": intermediate_size * (2 if is_gated else 1),
                "shape_k": layer.moe_config.hidden_dim,
            },
            "w2": {
                "shape_n": layer.moe_config.hidden_dim,
                "shape_k": intermediate_size,
            },
        }

    layer.weight_schemas = {}
    layer.input_schemas = {}
    humming_configs = {}

    for sublayer_name, configs in sublayer_configs.items():
        final_weight_schema, final_input_schema, humming_config = (
            _process_single_sublayer(
                layer=layer,
                sublayer_name=sublayer_name,
                shape_n=configs["shape_n"],
                shape_k=configs["shape_k"],
                weight_schema=weight_schema,
                input_schema=input_schema,
                has_bias=has_bias,
                num_experts=num_experts,
                param_dtype=param_dtype,
                force_weight_schema=force_weight_schema,
                allow_input_schema_fallback=allow_input_schema_fallback,
            )
        )

        layer.weight_schemas[sublayer_name] = final_weight_schema
        layer.input_schemas[sublayer_name] = final_input_schema
        humming_configs[sublayer_name] = humming_config

    layer.humming_configs = humming_configs
    return humming_configs

select_humming_moe_experts(config, weight_key, activation_key)

Select the primary Humming MoE Experts class Note: Shape-specific fallbacks may still occur at runtime.

Source code in vllm/model_executor/layers/quantization/utils/humming/moe.py
def select_humming_moe_experts(
    config: FusedMoEConfig,
    weight_key: QuantKey | None,
    activation_key: QuantKey | None,
) -> type[mk.FusedMoEExperts] | None:
    """Select the primary Humming MoE Experts class
    Note: Shape-specific fallbacks may still occur at runtime.
    """
    if not has_humming():
        return None

    # NOTE: the kernels are selected in the following order.
    AVAILABLE_EXPERTS: list[type[mk.FusedMoEExperts]] = [
        BatchedHummingGroupedExperts,
        HummingGroupedExperts,
        HummingIndexedExperts,
    ]

    # NOTE(rob): We need to peak into the P/F selection to determine
    # if we are using the batched or standard expert format, which
    # if not ideal. Once we unify TP + DP/EP, we can select P/F first.
    activation_format = (
        mk.FusedMoEActivationFormat.BatchedExperts
        if config.moe_parallel_config.use_batched_activation_format
        else mk.FusedMoEActivationFormat.Standard
    )

    def _make_log_backend(experts_cls: type[mk.FusedMoEExperts]):
        return f"Using {experts_cls.__name__} Humming MoE backend."

    def _make_log_unsupported(
        experts_cls: type[mk.FusedMoEExperts], reason: str | None
    ) -> str:
        if reason:
            return (
                f"Humming MoE experts {experts_cls.__name__} does not support the "
                f"deployment configuration since {reason}."
            )
        else:
            return (
                f"Humming MoE experts '{experts_cls.__name__}' does not support the "
                "deployment configuration."
            )

    for k_cls in AVAILABLE_EXPERTS:
        supported, reason = k_cls.is_supported_config(
            k_cls,
            config,
            weight_key,
            activation_key,
            activation_format,
        )
        if supported:
            logger.info_once(_make_log_backend(k_cls))
            return k_cls
        else:
            logger.debug_once(_make_log_unsupported(k_cls, reason))

    return None