Skip to content

vllm.model_executor.layers.quantization.utils.humming.schema

Map Humming schemas and handle shared checkpoint quantization settings.

_group_shape(group_size, group_size_n=0)

Map humming group sizes to QuantKey GroupShape.

group_size: elements per group along K (col); 0 means full dimension. group_size_n: elements per group along N (row); 0 means 1 (per-row).

GroupShape convention: row = N dim, col = K dim.

Source code in vllm/model_executor/layers/quantization/utils/humming/schema.py
def _group_shape(group_size: int, group_size_n: int = 0) -> GroupShape:
    """Map humming group sizes to QuantKey GroupShape.

    group_size:   elements per group along K (col); 0 means full dimension.
    group_size_n: elements per group along N (row); 0 means 1 (per-row).

    GroupShape convention: row = N dim, col = K dim.
    """
    if group_size == 0 and group_size_n == 0:
        return GroupShape.PER_CHANNEL

    row = group_size_n if group_size_n > 0 else 1
    col = group_size if group_size > 0 else -1
    return GroupShape(row=row, col=col)