vllm.v1.worker.gpu.spec_decode.utils
¶
Functions:
-
get_parallel_drafting_token_id–Resolve the mask token id used for parallel drafting slots.
-
get_pp_safe_draft_load_config–Avoid collectives that include PP ranks without a draft model.
get_parallel_drafting_token_id(hf_config)
¶
Resolve the mask token id used for parallel drafting slots.
Checks (in order): dflash_config.mask_token_id, top-level mask_token_id,
dspark_noise_token_id, pard_token, ptd_token_id. Raises ValueError if
none are present.
Source code in vllm/v1/worker/gpu/spec_decode/utils.py
get_pp_safe_draft_load_config(load_config)
¶
Avoid collectives that include PP ranks without a draft model.