vllm.v1.worker.gpu.sample.logits_processor
¶
Custom logits processors for the V2 model runner.
Kept import-light: the frontend process imports this package to validate per-request params, so nothing here may pull in model-runner side modules (torch, triton, worker state) at import time.
Modules:
-
interface–V2 logits processor interface.
-
loader–Loading of custom logits processor classes for the V2 model runner.
Classes:
-
LogitsContext–The current step's batch layout, passed to every
apply()call. -
LogitsProcRequestState–State associated with active requests, shared with logits processors.
-
LogitsProcessor–Custom logits processor for Model Runner V2.
Functions:
-
build_custom_logits_processors–Load and instantiate custom logits processors, entrypoint plugins first.
-
build_custom_logits_processors_params_validator–Load custom processor classes once and return a params validator.
LogitsContext
dataclass
¶
The current step's batch layout, passed to every apply() call.
A row is a logits row, not a request: rows are reordered every step, and
under speculative decoding a request owns one row per draft token.
Committed tokens live in req_states.all_token_ids (valid up to
total_len); this step's draft tokens are only in input_ids.
Source code in vllm/v1/worker/gpu/sample/logits_processor/interface.py
LogitsProcRequestState
dataclass
¶
State associated with active requests, shared with logits processors.
Wraps the model runner's per-slot buffers, which are mutated in place, so reads always see current values. Processors must treat every field as read-only.
Source code in vllm/v1/worker/gpu/sample/logits_processor/interface.py
LogitsProcessor
¶
Bases: ABC
Custom logits processor for Model Runner V2.
Per-request state is keyed by the request slot index; slots are recycled
through a free list, so per-slot state must be fully (re)initialized in
add_request().
apply() runs after the built-in bias, penalty, bad-words and grammar
stages and before temperature, min_p and top-k/top-p, so it sees unscaled
logits and must not re-inflate grammar-masked tokens. Thinking-budget
forcing runs after apply() and wins over its edits for requests whose
budget is exhausted.
State that is constant for a request belongs in __init__() or
add_request(); apply() receives only what changes per step.
Methods:
-
__init__–Capture what stays constant for the processor's lifetime.
-
add_request–Initialize per-slot state for a request entering the batch.
-
apply–Modify logits in place or return a new tensor. In-place modification
-
apply_staged_writes–Flush any host-side writes staged by
add_request()to the device. -
validate_params–Raise
ValueErrorfor invalid per-request arguments.
Source code in vllm/v1/worker/gpu/sample/logits_processor/interface.py
__init__(vllm_config, req_states)
¶
Capture what stays constant for the processor's lifetime.
req_states exposes the on-device token history and batch constants a
processor may read. Treat it as read-only.
Source code in vllm/v1/worker/gpu/sample/logits_processor/interface.py
add_request(req_idx, sampling_params)
¶
Initialize per-slot state for a request entering the batch.
The slot may hold a previous occupant's state; overwrite or neutralize all of it here.
Returns whether this processor modifies logits for the request.
Source code in vllm/v1/worker/gpu/sample/logits_processor/interface.py
apply(logits, ctx)
abstractmethod
¶
Modify logits in place or return a new tensor. In-place modification is preferred for efficiency.
apply() is called once for the whole batch, including rows of
requests this processor declined in add_request(), so filter rows
via ctx.expanded_idx_mapping.
Parameters:
-
(logits¶Tensor) –[num_logits_rows, vocab_size] float32 tensor.
-
(ctx¶LogitsContext) –this step's batch layout.
Source code in vllm/v1/worker/gpu/sample/logits_processor/interface.py
apply_staged_writes()
¶
Flush any host-side writes staged by add_request() to the device.
Called once per step before the forward pass, after the model runner
has flushed req_states, so a processor that stages writes here can
read the request's tokens back on device.
Source code in vllm/v1/worker/gpu/sample/logits_processor/interface.py
validate_params(sampling_params)
classmethod
¶
Raise ValueError for invalid per-request arguments.
Runs at request admission, so invalid arguments fail the request with an error instead of reaching the sampler.
Source code in vllm/v1/worker/gpu/sample/logits_processor/interface.py
build_custom_logits_processors(vllm_config, req_states, is_pooling_model, custom_logitsprocs=())
¶
Load and instantiate custom logits processors, entrypoint plugins first.
Raises:
-
ValueError–if a pooling model specifies custom processors, or a loaded class does not implement the V2 interface.
-
RuntimeError–if an FQCN fails to import.
Source code in vllm/v1/worker/gpu/sample/logits_processor/loader.py
build_custom_logits_processors_params_validator(custom_logitsprocs)
¶
Load custom processor classes once and return a params validator.
Called from the frontend at startup. The returned callable runs each
processor's validate_params at request admission.
Raises:
-
ValueError–if a loaded class does not implement the V2 interface.
-
RuntimeError–if an FQCN fails to import.