vllm.v1.worker.gpu.spec_decode.rejection_sampler
¶
Functions:
-
gather_draft_sampled–Gather the input token and position of each logits row.
-
get_max_chunk_logits–Largest number of logits rows one verification chunk may hold.
_iter_request_chunks(cu_num_logits, max_chunk_logits)
¶
Yield maximally packed request ranges without splitting requests.
Source code in vllm/v1/worker/gpu/spec_decode/rejection_sampler.py
gather_draft_sampled(input_ids, positions, logits_indices, expanded_idx_mapping, expanded_local_pos, prefill_len)
¶
Gather the input token and position of each logits row.
Draft rows of requests that have not yet sampled past their prefill are set to -1 so that the rejection kernels reject them.
Source code in vllm/v1/worker/gpu/spec_decode/rejection_sampler.py
get_max_chunk_logits(vocab_size)
¶
Largest number of logits rows one verification chunk may hold.