Skip to content

vllm.model_executor.model_loader.weight_cache.utils

Helpers shared by the weight cache daemon, the IPC loader and the engine.

These live outside protocol so callers that only need to know whether a draft is cached, or how a daemon group is named, do not have to import the wire format.

Functions:

format_daemon_role(is_draft)

Name of a daemon group: the target model or the draft.

Source code in vllm/model_executor/model_loader/weight_cache/utils.py
def format_daemon_role(is_draft: bool) -> str:
    """Name of a daemon group: the target model or the draft."""
    return "draft" if is_draft else "target"

format_socket_role_suffix(is_draft)

Socket-name suffix keeping the draft group distinct from the target.

Source code in vllm/model_executor/model_loader/weight_cache/utils.py
def format_socket_role_suffix(is_draft: bool) -> str:
    """Socket-name suffix keeping the draft group distinct from the target."""
    return "_draft" if is_draft else ""

is_draft_model_cacheable(speculative_config)

Whether the daemon serves the speculative draft as a separate role.

Source code in vllm/model_executor/model_loader/weight_cache/utils.py
def is_draft_model_cacheable(speculative_config: SpeculativeConfig | None) -> bool:
    """Whether the daemon serves the speculative draft as a separate role."""
    return (
        speculative_config is not None
        and speculative_config.method in WEIGHT_CACHE_DRAFT_METHODS
        and speculative_config.draft_model_config is not None
    )