vllm.v1.worker.gpu.spec_decode.eagle.utils
¶
Functions:
-
get_target_lm_head–The target's lm_head — from get_language_model() for
-
maybe_share_target_embed–Share the target input embedding with the drafter when needed.
_should_share(eagle, flag, draft, target)
¶
Share unless the draft declares its own copy that differs from the target.
A draft that declares its own copy but has no top-level one (e.g. MTP heads stored per layer) keeps it.
Source code in vllm/v1/worker/gpu/spec_decode/eagle/utils.py
get_target_lm_head(target_model, target_language_model)
¶
The target's lm_head — from get_language_model() for *ForConditionalGeneration targets, else the top-level module.
Source code in vllm/v1/worker/gpu/spec_decode/eagle/utils.py
maybe_share_target_embed(draft_model, draft_inner, target_inner)
¶
Share the target input embedding with the drafter when needed.