Skip to content

vllm.entrypoints.scale_out.token_in_token_out.protocol

Classes:

Functions:

DerenderChatRequest

Bases: BaseModel

Request for the /v1/chat/completions/derender endpoint (non-streaming).

Wraps a complete GenerateResponse and caller supplied metadata needed to produce a fully formed ChatCompletionResponse without a GPU.

Attributes:

  • chat_request (ChatCompletionRequest | None) –

    The original (post-adjust_request) ChatCompletionRequest from /render.

  • generate_response (GenerateTokensResponse) –

    The complete token-in / token-out engine response to derender.

  • model (str | None) –

    Served model name. Defaults to the server's served model name.

  • prompt_token_ids (list[NonNegativeInt] | None) –

    Prompt token IDs (GenerateRequest.token_ids from /render). Seeds

  • prompt_tokens (int | None) –

    Prompt token count for usage; defaults to 0 if omitted.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class DerenderChatRequest(BaseModel):
    """Request for the /v1/chat/completions/derender endpoint (non-streaming).

    Wraps a complete GenerateResponse and caller supplied metadata needed to
    produce a fully formed ChatCompletionResponse without a GPU.
    """

    # --8<-- [start:derender-chat-request]
    stream: Literal[False] = False

    model: str | None = None
    """Served model name. Defaults to the server's served model name."""

    generate_response: GenerateTokensResponse
    """The complete token-in / token-out engine response to derender.

    Only `output_mode="tokens"` responses are accepted. Other modes are
    already detokenized and re-decoding their `token_ids` would undo the
    engine's stop string truncation.
    """

    prompt_tokens: int | None = None
    """Prompt token count for usage; defaults to 0 if omitted.

    GenerateResponse carries only output tokens; the caller already has
    len(GenerateRequest.token_ids) from the render step.
    """

    prompt_token_ids: list[NonNegativeInt] | None = None
    """Prompt token IDs (`GenerateRequest.token_ids` from /render). Seeds
    detokenization from the prompt tail so the first output token keeps its
    leading space on SentencePiece tokenizers. Falls back to
    `generate_response.prompt_token_ids`, then to unseeded decoding.

    Only the last few IDs are read, so a suffix of the prompt is enough.
    """

    chat_request: ChatCompletionRequest | None = None
    """The original (post-adjust_request) ChatCompletionRequest from /render.

    Required by the parsing so that tool/reasoning parsers can receive the full
    request context they expect (request.tools, request.tool_choice,
    request._grammar_from_parser, etc.).
    """

chat_request = None class-attribute instance-attribute

The original (post-adjust_request) ChatCompletionRequest from /render.

Required by the parsing so that tool/reasoning parsers can receive the full request context they expect (request.tools, request.tool_choice, request._grammar_from_parser, etc.).

generate_response instance-attribute

The complete token-in / token-out engine response to derender.

Only output_mode="tokens" responses are accepted. Other modes are already detokenized and re-decoding their token_ids would undo the engine's stop string truncation.

model = None class-attribute instance-attribute

Served model name. Defaults to the server's served model name.

prompt_token_ids = None class-attribute instance-attribute

Prompt token IDs (GenerateRequest.token_ids from /render). Seeds detokenization from the prompt tail so the first output token keeps its leading space on SentencePiece tokenizers. Falls back to generate_response.prompt_token_ids, then to unseeded decoding.

Only the last few IDs are read, so a suffix of the prompt is enough.

prompt_tokens = None class-attribute instance-attribute

Prompt token count for usage; defaults to 0 if omitted.

GenerateResponse carries only output tokens; the caller already has len(GenerateRequest.token_ids) from the render step.

DerenderChatStreamRequest

Bases: BaseModel

One chunk streaming derender request for /v1/chat/completions/derender.

The client sends one request per SSE chunk received from /inference/v1/generate. Each request carries the generate chunk plus the stream_state returned by the previous call (None on the first call). The response contains the derendered chunk and the updated state to be passed to the next call.

This implements stateless no server side session. All mutable state lives in the client carried stream_state.

Attributes:

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class DerenderChatStreamRequest(BaseModel):
    """One chunk streaming derender request for /v1/chat/completions/derender.

    The client sends one request per SSE chunk received from
    `/inference/v1/generate`.  Each request carries the generate chunk
    plus the `stream_state` returned by the previous call (`None` on the
    first call).  The response contains the derendered chunk and the updated
    state to be passed to the next call.

    This implements stateless no server side session. All mutable state lives in
    the client carried `stream_state`.
    """

    # --8<-- [start:derender-chat-stream-request]
    stream: Literal[True]

    model: str | None = None
    generate_chunk: GenerateTokensStreamResponse
    """One `output_mode="tokens"` SSE chunk from `/inference/v1/generate`
    (`stream=True`)."""

    stream_state: DerenderStreamState | None = None
    """Client carried detok state from the previous call. `None` on first."""

    prompt_tokens: int | None = None
    """Prompt token count for usage. Forwarded from the render step."""

    prompt_token_ids: list[NonNegativeInt] | None = None
    """Prompt token IDs. Required by the parser path's `parse_delta` to
    settle its initial reasoning state (e.g. chat templates that pre-open
    `<think>`). `prompt_tokens` is a usage count and cannot serve this
    purpose. Sourced from `GenerateRequest.token_ids` at the render step.

    Rejected with a 400 (by `ServingDerender`) when a tool or reasoning
    parser is configured and this is omitted. Without it, `parse_delta`
    cannot tell whether the prompt left reasoning open and would silently
    misclassify reasoning content as plain content. On all paths it also
    seeds detokenization on the first chunk, falling back to
    `generate_chunk.prompt_token_ids` when omitted (see
    `DerenderChatRequest.prompt_token_ids`).

    With a parser configured, send the full list. Parsers look back to the
    last reasoning marker which can be anywhere in the prompt, so a suffix
    can flip the initial reasoning state. Without a parser, a suffix is
    enough.
    """

    chat_request: ChatCompletionRequest | None = None
    """The original (post adjust_request) ChatCompletionRequest from /render."""

chat_request = None class-attribute instance-attribute

The original (post adjust_request) ChatCompletionRequest from /render.

generate_chunk instance-attribute

One output_mode="tokens" SSE chunk from /inference/v1/generate (stream=True).

prompt_token_ids = None class-attribute instance-attribute

Prompt token IDs. Required by the parser path's parse_delta to settle its initial reasoning state (e.g. chat templates that pre-open <think>). prompt_tokens is a usage count and cannot serve this purpose. Sourced from GenerateRequest.token_ids at the render step.

Rejected with a 400 (by ServingDerender) when a tool or reasoning parser is configured and this is omitted. Without it, parse_delta cannot tell whether the prompt left reasoning open and would silently misclassify reasoning content as plain content. On all paths it also seeds detokenization on the first chunk, falling back to generate_chunk.prompt_token_ids when omitted (see DerenderChatRequest.prompt_token_ids).

With a parser configured, send the full list. Parsers look back to the last reasoning marker which can be anywhere in the prompt, so a suffix can flip the initial reasoning state. Without a parser, a suffix is enough.

prompt_tokens = None class-attribute instance-attribute

Prompt token count for usage. Forwarded from the render step.

stream_state = None class-attribute instance-attribute

Client carried detok state from the previous call. None on first.

DerenderChatStreamResponse

Bases: BaseModel

Response for one streaming chat derender chunk.

Pairs the derendered SSE chunk with the updated client carried state to pass to the next call.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class DerenderChatStreamResponse(BaseModel):
    """Response for one streaming chat derender chunk.

    Pairs the derendered SSE chunk with the updated client carried state to
    pass to the next call.
    """

    chunk: ChatCompletionStreamResponse
    stream_state: DerenderStreamState

DerenderCompletionRequest

Bases: BaseModel

Request for the /v1/completions/derender endpoint (non-streaming).

Parallel to DerenderChatRequest but handles the multi-prompt completions case: one GenerateResponse per prompt, mirroring the list[GenerateRequest] returned by /v1/completions/render.

Attributes:

  • completion_request (CompletionRequest | None) –

    The original (post-adjust_request) CompletionRequest from /render.

  • generate_responses (list[GenerateTokensResponse]) –

    One output_mode="tokens" response per prompt, parallel to the

  • model (str | None) –

    Served model name. Defaults to the server's served model name.

  • prompt_token_ids (list[list[NonNegativeInt] | None] | None) –

    One prompt token ID list per response, used to seed detokenization.

  • prompt_tokens (list[int] | None) –

    One prompt token count per response; each defaults to 0 if omitted.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class DerenderCompletionRequest(BaseModel):
    """Request for the /v1/completions/derender endpoint (non-streaming).

    Parallel to DerenderChatRequest but handles the multi-prompt completions
    case: one GenerateResponse per prompt, mirroring the list[GenerateRequest]
    returned by /v1/completions/render.
    """

    # --8<-- [start:derender-completion-request]
    stream: Literal[False] = False

    model: str | None = None
    """Served model name. Defaults to the server's served model name."""

    generate_responses: list[GenerateTokensResponse]
    """One `output_mode="tokens"` response per prompt, parallel to the
    list[GenerateRequest] returned by /v1/completions/render."""

    prompt_tokens: list[int] | None = None
    """One prompt token count per response; each defaults to 0 if omitted.

    If provided, len(prompt_tokens) must equal len(generate_responses).
    """

    prompt_token_ids: list[list[NonNegativeInt] | None] | None = None
    """One prompt token ID list per response, used to seed detokenization.
    See `DerenderChatRequest.prompt_token_ids`.

    If provided, len(prompt_token_ids) must equal len(generate_responses).
    A `None` entry falls back to `generate_responses[i].prompt_token_ids`.
    """

    completion_request: CompletionRequest | None = None
    """The original (post-adjust_request) CompletionRequest from /render.

    Mirrors chat_request on DerenderChatRequest. Required by the parsing
    so parsers receive the full request context.
    """
    # --8<-- [end:derender-completion-request]

    @model_validator(mode="after")
    def _validate_prompt_tokens_length(self) -> "DerenderCompletionRequest":
        if self.prompt_tokens is not None and len(self.prompt_tokens) != len(
            self.generate_responses
        ):
            raise ValueError(
                f"prompt_tokens length ({len(self.prompt_tokens)}) must equal "
                f"generate_responses length ({len(self.generate_responses)})"
            )
        if self.prompt_token_ids is not None and len(self.prompt_token_ids) != len(
            self.generate_responses
        ):
            raise ValueError(
                f"prompt_token_ids length ({len(self.prompt_token_ids)}) must "
                f"equal generate_responses length ({len(self.generate_responses)})"
            )
        return self

completion_request = None class-attribute instance-attribute

The original (post-adjust_request) CompletionRequest from /render.

Mirrors chat_request on DerenderChatRequest. Required by the parsing so parsers receive the full request context.

generate_responses instance-attribute

One output_mode="tokens" response per prompt, parallel to the list[GenerateRequest] returned by /v1/completions/render.

model = None class-attribute instance-attribute

Served model name. Defaults to the server's served model name.

prompt_token_ids = None class-attribute instance-attribute

One prompt token ID list per response, used to seed detokenization. See DerenderChatRequest.prompt_token_ids.

If provided, len(prompt_token_ids) must equal len(generate_responses). A None entry falls back to generate_responses[i].prompt_token_ids.

prompt_tokens = None class-attribute instance-attribute

One prompt token count per response; each defaults to 0 if omitted.

If provided, len(prompt_tokens) must equal len(generate_responses).

DerenderCompletionStreamRequest

Bases: BaseModel

One chunk streaming derender request for /v1/completions/derender.

Parallel to DerenderChatStreamRequest for the completions endpoint. Each call processes one SSE chunk (one output sequence's delta) and returns the derendered chunk plus updated state.

Attributes:

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class DerenderCompletionStreamRequest(BaseModel):
    """One chunk streaming derender request for /v1/completions/derender.

    Parallel to `DerenderChatStreamRequest` for the completions endpoint.
    Each call processes one SSE chunk (one output sequence's delta) and
    returns the derendered chunk plus updated state.
    """

    # --8<-- [start:derender-completion-stream-request]
    stream: Literal[True]

    model: str | None = None
    generate_chunk: GenerateTokensStreamResponse
    """One `output_mode="tokens"` SSE chunk from `/inference/v1/generate`."""

    stream_state: DerenderStreamState | None = None
    """Client-carried detok state. `None` on the first call."""

    prompt_tokens: int | None = None
    """Prompt token count for usage."""

    prompt_token_ids: list[NonNegativeInt] | None = None
    """Prompt token IDs, used on the first chunk to seed detokenization.
    Falls back to `generate_chunk.prompt_token_ids`. See
    `DerenderChatRequest.prompt_token_ids`.
    """

    completion_request: CompletionRequest | None = None
    """The original (post adjust_request) CompletionRequest from /render."""

completion_request = None class-attribute instance-attribute

The original (post adjust_request) CompletionRequest from /render.

generate_chunk instance-attribute

One output_mode="tokens" SSE chunk from /inference/v1/generate.

prompt_token_ids = None class-attribute instance-attribute

Prompt token IDs, used on the first chunk to seed detokenization. Falls back to generate_chunk.prompt_token_ids. See DerenderChatRequest.prompt_token_ids.

prompt_tokens = None class-attribute instance-attribute

Prompt token count for usage.

stream_state = None class-attribute instance-attribute

Client-carried detok state. None on the first call.

DerenderCompletionStreamResponse

Bases: BaseModel

Response for one streaming completions derender chunk.

Parallel to DerenderChatStreamResponse for the completions endpoint.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class DerenderCompletionStreamResponse(BaseModel):
    """Response for one streaming completions derender chunk.

    Parallel to `DerenderChatStreamResponse` for the completions endpoint.
    """

    chunk: CompletionStreamResponse
    stream_state: DerenderStreamState

DerenderStreamState

Bases: BaseModel

Per sequence state for stateless streaming derender.

The client carries this between successive per chunk HTTP calls to the streaming derender endpoint. All fields are plain JSON serializable data. No opaque tokenizer or parser internals are stored here.

Two separate sets of fields support two different streaming modes:

  • For plain streaming (no parser configured including the completions path): prev_tokens, prefix_offset and read_offset maintain a bounded incremental decoding window. This requires O(window) transport and O(delta) computation per chunk.
  • For parser enabled chat streaming: output_token_ids, output_chunk_lens, tools_streamed and last_tool_call_ids are used to replay parse_delta() from scratch on every chunk because parser state cannot be serialized. This incurs O(n) transport per chunk (O(n²) per generation) and O(n²) parse_delta() calls per generation. Since many parsers re-scan the entire accumulated text on each invocation, parse_delta() itself is O(n) yielding a true worst case compute cost of O(n³) per generation. No caching is performed. Work is bounded by max_model_len. See OnlineDerenderer._derender_chat_stream_parsed.

The detokenization strategy carries the incremental decode offsets directly rather than re-sending the whole token history each chunk. detokenize_incrementally only ever reads the trailing token window prev_tokens[prefix_offset:], so we carry just that tail plus the two offsets. Each chunk resumes exactly where the last one stopped, including any partially processed multi-byte character (tracked by read_offset), then trims and rebases the window so it never grows with generation length.

Performance: - Compute per chunk is O(delta). One detokenize_incrementally call per new token, independent of how many tokens preceded it. - Transport per chunk is O(window). The carried tail is bounded by the incremental detokenization offset, so cumulative bytes over the wire are O(n) rather than the O(n^2) a full history round trip would incur.

Attributes:

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class DerenderStreamState(BaseModel):
    """Per sequence state for stateless streaming derender.

    The client carries this between successive per chunk HTTP calls to the
    streaming derender endpoint. All fields are plain JSON serializable data.
    No opaque tokenizer or parser internals are stored here.

    Two separate sets of fields support two different streaming modes:

    - For plain streaming (no parser configured including the completions path):
      `prev_tokens`, `prefix_offset` and `read_offset` maintain a bounded incremental
      decoding window. This requires O(window) transport and O(delta) computation
      per chunk.
    - For parser enabled chat streaming: `output_token_ids`, `output_chunk_lens`,
      `tools_streamed` and `last_tool_call_ids` are used to replay `parse_delta()`
      from scratch on every chunk because parser state cannot be serialized. This
      incurs O(n) transport per chunk (O(n²) per generation) and O(n²)
      `parse_delta()` calls per generation. Since many parsers re-scan the
      entire accumulated text on each invocation, `parse_delta()` itself is
      O(n) yielding a true worst case compute cost of O(n³) per generation.
      No caching is performed. Work is bounded by `max_model_len`.
      See `OnlineDerenderer._derender_chat_stream_parsed`.

    The detokenization strategy carries the incremental decode offsets
    directly rather than re-sending the whole token history each chunk.
    `detokenize_incrementally` only ever reads the trailing token window
    `prev_tokens[prefix_offset:]`, so we carry just that tail plus the two
    offsets. Each chunk resumes exactly where the last one stopped, including
    any partially processed multi-byte character (tracked by `read_offset`),
    then trims and rebases the window so it never grows with generation length.

    Performance:
    - Compute per chunk is O(delta). One `detokenize_incrementally` call per
      new token, independent of how many tokens preceded it.
    - Transport per chunk is O(window). The carried tail is bounded by the
      incremental detokenization offset, so cumulative bytes over the wire are
      O(n) rather than the O(n^2) a full history round trip would incur.
    """

    prev_tokens: list[str] = Field(default_factory=list)
    """Trailing decode window. Token strings from `prefix_offset` onward.

    Bounded, trimmed and rebased each chunk to the tail
    `detokenize_incrementally` still reads, so it does not grow with the
    number of chunks.
    """

    prefix_offset: int = Field(default=0, ge=0)
    """Prefix offset into `prev_tokens` for incremental detokenization."""

    read_offset: int = Field(default=0, ge=0)
    """Read offset into `prev_tokens` for incremental detokenization."""

    @field_validator("prev_tokens")
    @classmethod
    def _bound_prev_tokens(cls, v: list[str]) -> list[str]:
        # INITIAL_INCREMENTAL_DETOKENIZATION_OFFSET is small (5) and the trimmed
        # window is O(offset). A generous limit rejects unusually large or malformed
        # payloads without restricting legitimate multi-byte sequences.
        limit = 1024
        if len(v) > limit:
            raise ValueError(f"prev_tokens length ({len(v)}) exceeds maximum ({limit})")
        return v

    role_sent: bool = False
    """True once the initial `role: "assistant"` delta has been emitted.

    Prevents re-emitting the role on subsequent chunks even when the detok
    window is transiently empty (e.g. usage only final chunk).
    """

    logprob_context_token_ids: list[int] = Field(default_factory=list, max_length=4)
    """Trailing sampled token IDs carried across chunks so byte-fallback
    (U+FFFD) correction during logprob placeholder resolution has context
    at chunk boundaries. Bounded to the 4-token window that
    ``_correct_decoded_token`` reads."""

    logprob_text_offset: int = Field(default=0, ge=0)
    """Cumulative emitted text length. Seeds ``text_offset`` for completion
    streaming logprobs so offsets stay absolute across chunks, mirroring
    ``initial_text_offset`` in the generate streaming path."""

    output_token_ids: list[int] = Field(default_factory=list)
    """All output tokens seen so far. Parser path only.

    Replay buffer: each chunk rebuilds a fresh parser and replays every
    token in here through `parse_delta` (discarding the result) before
    processing the current chunk's tokens since parser internal state
    cannot be serialized into this stateless model. Unavoidably O(n)
    bounded by `max_model_len` (enforced server side, not by a field
    validator here since the bound is model dependent).
    """

    output_chunk_lens: list[Annotated[int, Field(gt=0)]] = Field(default_factory=list)
    """Token count of each chunk in `output_token_ids`. Parser path only.

    Replay uses these to reproduce the original `parse_delta` call
    boundaries. Must sum to `len(output_token_ids)`.
    """

    @model_validator(mode="after")
    def _validate_output_chunk_lens(self) -> "DerenderStreamState":
        total = sum(self.output_chunk_lens)
        if total != len(self.output_token_ids):
            raise ValueError(
                f"output_chunk_lens must sum to len(output_token_ids) "
                f"(got sum={total}, len(output_token_ids)="
                f"{len(self.output_token_ids)})"
            )
        return self

    tools_streamed: bool = False
    """True once a tool call delta has been emitted. Parser path only.

    Drives the `finish_reason` -> `"tool_calls"` rewrite on the final
    chunk mirroring the generate streaming path.
    """

    last_tool_call_ids: list[str] = Field(default_factory=list)
    """Stable tool call IDs, assigned once when each call first appears.

    Indexed by tool call index. Parser path only. Prevents ID regeneration
    when replay reprocesses a tool call that already has a pinned ID.
    """

last_tool_call_ids = Field(default_factory=list) class-attribute instance-attribute

Stable tool call IDs, assigned once when each call first appears.

Indexed by tool call index. Parser path only. Prevents ID regeneration when replay reprocesses a tool call that already has a pinned ID.

logprob_context_token_ids = Field(default_factory=list, max_length=4) class-attribute instance-attribute

Trailing sampled token IDs carried across chunks so byte-fallback (U+FFFD) correction during logprob placeholder resolution has context at chunk boundaries. Bounded to the 4-token window that _correct_decoded_token reads.

logprob_text_offset = Field(default=0, ge=0) class-attribute instance-attribute

Cumulative emitted text length. Seeds text_offset for completion streaming logprobs so offsets stay absolute across chunks, mirroring initial_text_offset in the generate streaming path.

output_chunk_lens = Field(default_factory=list) class-attribute instance-attribute

Token count of each chunk in output_token_ids. Parser path only.

Replay uses these to reproduce the original parse_delta call boundaries. Must sum to len(output_token_ids).

output_token_ids = Field(default_factory=list) class-attribute instance-attribute

All output tokens seen so far. Parser path only.

Replay buffer: each chunk rebuilds a fresh parser and replays every token in here through parse_delta (discarding the result) before processing the current chunk's tokens since parser internal state cannot be serialized into this stateless model. Unavoidably O(n) bounded by max_model_len (enforced server side, not by a field validator here since the bound is model dependent).

prefix_offset = Field(default=0, ge=0) class-attribute instance-attribute

Prefix offset into prev_tokens for incremental detokenization.

prev_tokens = Field(default_factory=list) class-attribute instance-attribute

Trailing decode window. Token strings from prefix_offset onward.

Bounded, trimmed and rebased each chunk to the tail detokenize_incrementally still reads, so it does not grow with the number of chunks.

read_offset = Field(default=0, ge=0) class-attribute instance-attribute

Read offset into prev_tokens for incremental detokenization.

role_sent = False class-attribute instance-attribute

True once the initial role: "assistant" delta has been emitted.

Prevents re-emitting the role on subsequent chunks even when the detok window is transiently empty (e.g. usage only final chunk).

tools_streamed = False class-attribute instance-attribute

True once a tool call delta has been emitted. Parser path only.

Drives the finish_reason -> "tool_calls" rewrite on the final chunk mirroring the generate streaming path.

GenerateChoiceBase

Bases: BaseModel

Fields shared by every output_mode of a non-streaming choice.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class GenerateChoiceBase(BaseModel):
    """Fields shared by every `output_mode` of a non-streaming choice."""

    index: int
    # per OpenAI spec this is the default
    finish_reason: str | None = "stop"
    token_ids: list[int] | None = None
    # Per-token expert routing decisions, base64-encoded `.npy` bytes
    # (numpy serialization). Shape after decode:
    #   (num_tokens - 1, num_layers, num_experts_per_tok) dtype uint8/uint16/int32
    # `num_tokens - 1` because the last sampled token has not been
    # forwarded yet and therefore has no routing data.
    # Decode:
    #   np.load(io.BytesIO(base64.b64decode(s)))
    # `None` if (a) the request was aborted before any forward pass,
    # or (b) `enable_return_routed_experts` is off server-side.
    routed_experts: str | None = None
    sampling_mask: list[list[int]] | None = None

    @field_validator("token_ids")
    @classmethod
    def validate_token_ids(cls, v: list[int] | None) -> list[int] | None:
        if v is not None and any(t < 0 for t in v):
            raise ValueError("token_ids must not contain negative values")
        return v

GenerateRequest

Bases: BaseModel

Methods:

Attributes:

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class GenerateRequest(BaseModel):
    request_id: str = Field(
        default_factory=lambda: f"{random_uuid()}",
        description=(
            "The request_id related to this request. If the caller does "
            "not set it, a random_uuid will be generated. This id is used "
            "through out the inference process and return in response."
        ),
    )
    token_ids: list[int] = Field(min_length=1)
    """The token ids to generate text from."""

    @field_validator("token_ids")
    @classmethod
    def validate_token_ids(cls, v: list[int]) -> list[int]:
        if any(t < 0 for t in v):
            raise ValueError("token_ids must not contain negative values")
        return v

    token_offsets: list[tuple[int, int]] | None = None
    """Char-level (start, end) offsets per token, relative to the
    tokenized source string. Present only when the request set
    `return_token_offsets=True` and the renderer was able to compute
    them (Fast tokenizer, text input, no multimodal data). List length
    equals `token_ids` length when present. None otherwise."""

    features: MultiModalFeatures | None = None
    """Multimodal hashes and placeholder positions (populated for MM inputs)."""

    content_parts: list[dict[str, Any]] | None = None
    """Raw multimodal input; server resolves media. Mutually exclusive
    with `features`."""

    @model_validator(mode="after")
    def _check_mm_fields_exclusive(self) -> "GenerateRequest":
        if self.content_parts and self.features:
            raise ValueError("content_parts and features are mutually exclusive")
        return self

    @model_validator(mode="after")
    def _require_ec_for_metadata_only(self) -> "GenerateRequest":
        features = self.features
        if features is None:
            return self
        if not _has_serialized_mm_items(features.mm_metadata):
            return self
        if self.ec_transfer_params:
            return self
        kwargs_data = features.kwargs_data or {}
        for modality, metadata_items in (features.mm_metadata or {}).items():
            kwargs_items = kwargs_data.get(modality, [None] * len(metadata_items))
            if any(
                metadata is not None and kwargs is None
                for metadata, kwargs in zip(metadata_items, kwargs_items, strict=True)
            ):
                raise ValueError(
                    "metadata-only multimodal items require ec_transfer_params"
                )
        return self

    sampling_params: SamplingParams
    """The sampling parameters for the model."""

    model: str | None = None

    return_token_ids: bool | None = Field(
        default=None,
        description=(
            "If true, return the final prompt token IDs after multimodal "
            "placeholder expansion, together with multimodal placeholder ranges. "
            "In streaming mode, this metadata is included only in the first chunk."
        ),
    )

    output_mode: OutputMode = Field(
        default="tokens",
        description=(
            "Response level. 'tokens' returns token IDs only. 'text' also "
            "returns the detokenized text and logprobs with decoded token "
            "strings. 'text' needs a server that loads a tokenizer and "
            "sampling_params.detokenize to be true."
        ),
    )

    stream: bool | None = False
    stream_options: StreamOptions | None = None
    cache_salt: str | None = Field(
        default=None,
        min_length=1,
        max_length=1024,
        description=(
            "If specified, the prefix cache will be salted with the provided "
            "string to prevent an attacker to guess prompts in multi-user "
            "environments. The salt should be random, protected from "
            "access by 3rd parties, and long enough to be "
            "unpredictable (e.g., 43 characters base64-encoded, corresponding "
            "to 256 bit)."
        ),
    )
    priority: int = Field(
        default=0,
        ge=-(2**63),
        le=2**63 - 1,
        description=(
            "The priority of the request (lower means earlier handling; "
            "default: 0). Any priority other than 0 will raise an error "
            "if the served model does not use priority scheduling."
        ),
    )
    kv_transfer_params: dict[str, Any] | None = Field(
        default=None,
        description="KVTransfer parameters used for disaggregated serving.",
    )
    ec_transfer_params: dict[str, Any] | None = Field(
        default=None,
        description=(
            "ECTransfer parameters used for encoder-cache disaggregated serving."
        ),
    )

    # Tracks which keys the caller explicitly set inside `sampling_params`
    # when the request was parsed from a JSON body. Lets the server tell
    # "client said max_tokens=16" from "client said nothing → dataclass
    # default 16" so it can apply server-side defaulting only in the latter
    # case. `None` means the request was constructed with a pre-built
    # `SamplingParams` instance (e.g. from internal callers that have
    # already resolved values), in which case all fields are considered set.
    _sampling_params_provided_keys: set[str] | None = PrivateAttr(default=None)
    _response_mm_placeholders: dict[str, list[PlaceholderRangeInfo]] | None = (
        PrivateAttr(default=None)
    )

    @model_validator(mode="wrap")
    @classmethod
    def _capture_sampling_params_provided_keys(cls, data: Any, handler):
        provided: set[str] | None = None
        if isinstance(data, dict):
            sp = data.get("sampling_params")
            if isinstance(sp, dict):
                provided = set(sp.keys())
        instance = handler(data)
        instance._sampling_params_provided_keys = provided
        return instance

    @model_validator(mode="before")
    @classmethod
    def _validate_cache_salt(cls, data: Any) -> Any:
        if isinstance(data, dict):
            validate_cache_salt(data.get("cache_salt"))
        return data

    @model_validator(mode="after")
    def _check_output_mode_detokenizes(self) -> "GenerateRequest":
        if self.output_mode != "tokens" and not self.sampling_params.detokenize:
            raise ValueError(
                f"output_mode={self.output_mode!r} requires "
                "sampling_params.detokenize to be true"
            )
        return self

    @model_validator(mode="after")
    def _validate_multimodal_feature_bounds(self) -> "GenerateRequest":
        if self.features is None:
            return self

        prompt_len = len(self.token_ids)
        for ranges in self.features.mm_placeholders.values():
            for placeholder in ranges:
                if placeholder.offset + placeholder.length > prompt_len:
                    raise ValueError(
                        "mm_placeholders must remain within the token_ids sequence"
                    )
        return self

    def is_sampling_param_provided(self, name: str) -> bool:
        """Whether the caller explicitly set `sampling_params.<name>`.

        For requests parsed from a JSON body, this reflects the raw input
        dict. For requests constructed with a pre-built `SamplingParams`
        instance, all fields are considered provided so server-side defaults
        do not clobber values already resolved upstream.
        """
        if self._sampling_params_provided_keys is None:
            return True
        return name in self._sampling_params_provided_keys

    def build_tok_params(self, model_config: ModelConfig) -> TokenizeParams:
        return TokenizeParams(
            max_total_tokens=None,
            max_output_tokens=0,
        )

content_parts = None class-attribute instance-attribute

Raw multimodal input; server resolves media. Mutually exclusive with features.

features = None class-attribute instance-attribute

Multimodal hashes and placeholder positions (populated for MM inputs).

sampling_params instance-attribute

The sampling parameters for the model.

token_ids = Field(min_length=1) class-attribute instance-attribute

The token ids to generate text from.

token_offsets = None class-attribute instance-attribute

Char-level (start, end) offsets per token, relative to the tokenized source string. Present only when the request set return_token_offsets=True and the renderer was able to compute them (Fast tokenizer, text input, no multimodal data). List length equals token_ids length when present. None otherwise.

is_sampling_param_provided(name)

Whether the caller explicitly set sampling_params.<name>.

For requests parsed from a JSON body, this reflects the raw input dict. For requests constructed with a pre-built SamplingParams instance, all fields are considered provided so server-side defaults do not clobber values already resolved upstream.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
def is_sampling_param_provided(self, name: str) -> bool:
    """Whether the caller explicitly set `sampling_params.<name>`.

    For requests parsed from a JSON body, this reflects the raw input
    dict. For requests constructed with a pre-built `SamplingParams`
    instance, all fields are considered provided so server-side defaults
    do not clobber values already resolved upstream.
    """
    if self._sampling_params_provided_keys is None:
        return True
    return name in self._sampling_params_provided_keys

GenerateStreamChoiceBase

Bases: BaseModel

Fields shared by every output_mode of a streaming choice.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class GenerateStreamChoiceBase(BaseModel):
    """Fields shared by every `output_mode` of a streaming choice."""

    index: int
    finish_reason: str | None = None
    token_ids: list[int] | None = None
    routed_experts: str | None = None
    sampling_mask: list[list[int]] | None = None

GenerateTextChoice

Bases: GenerateChoiceBase

Attributes:

  • logprobs (ChatCompletionLogProbs | None) –

    Logprobs with decoded token strings and bytes.

  • text (str) –

    Detokenized output. Excludes a matched stop string unless

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class GenerateTextChoice(GenerateChoiceBase):
    text: str
    """Detokenized output. Excludes a matched stop string unless
    `include_stop_str_in_output` is set, while `token_ids` keeps every
    generated token."""

    logprobs: ChatCompletionLogProbs | None = None
    """Logprobs with decoded token strings and `bytes`."""

logprobs = None class-attribute instance-attribute

Logprobs with decoded token strings and bytes.

text instance-attribute

Detokenized output. Excludes a matched stop string unless include_stop_str_in_output is set, while token_ids keeps every generated token.

GenerateTextStreamChoice

Bases: GenerateStreamChoiceBase

Attributes:

  • text (str) –

    Text delta since the previous chunk of this choice.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class GenerateTextStreamChoice(GenerateStreamChoiceBase):
    text: str
    """Text delta since the previous chunk of this choice."""

    logprobs: ChatCompletionLogProbs | None = None

text instance-attribute

Text delta since the previous chunk of this choice.

GenerateTokensChoice

Bases: GenerateChoiceBase

Attributes:

  • logprobs (ChatCompletionLogProbs | None) –

    Logprobs whose tokens are token_id:N placeholders.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class GenerateTokensChoice(GenerateChoiceBase):
    logprobs: ChatCompletionLogProbs | None = None
    """Logprobs whose tokens are `token_id:N` placeholders."""

logprobs = None class-attribute instance-attribute

Logprobs whose tokens are token_id:N placeholders.

MultiModalFeatures

Bases: BaseModel

Lightweight multimodal metadata produced by the render step.

Carries hashes (for cache lookup / identification) and placeholder positions so the downstream /generate service knows where in the token sequence each multimodal item lives.

Attributes:

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class MultiModalFeatures(BaseModel):
    """Lightweight multimodal metadata produced by the render step.

    Carries hashes (for cache lookup / identification) and placeholder
    positions so the downstream `/generate` service knows *where* in
    the token sequence each multimodal item lives.
    """

    mm_hashes: dict[str, list[str]]
    """Per-modality item hashes, e.g. `{"image": ["abc", "def"]}`."""

    mm_placeholders: dict[str, list[PlaceholderRangeInfo]]
    """Per-modality placeholder ranges in the token sequence."""

    kwargs_data: dict[str, list[str | None]] | None = None
    """Per-modality serialized tensor data.

    Each value is a list parallel to `mm_hashes[modality]`.  A `str`
    entry is a base64-encoded `MultiModalKwargsItem`; `None` means
    the item should be resolved from cache.  The entire field is
    `None` for metadata-only (cache-hit) responses.
    """

    mm_metadata: dict[str, list[str | None]] | None = None
    """Per-modality serialized metadata for disaggregated prefill.

    Each value is a list parallel to `mm_hashes[modality]`. A `str`
    entry is a base64-encoded `MultiModalKwargsItem` containing only
    placeholder-metadata and `keep_on_cpu` fields. `None` means that
    the metadata is unavailable for that item. Prefill can use this
    instead of `kwargs_data` only when `ec_transfer_params` is also
    set, so embeddings arrive through the EC connector rather than from
    `pixel_values`.
    """

    @model_validator(mode="after")
    def _validate_parallel_fields(self) -> "MultiModalFeatures":
        modalities = set(self.mm_hashes)
        if set(self.mm_placeholders) != modalities:
            raise ValueError(
                "mm_hashes and mm_placeholders must use the same modalities"
            )
        if self.kwargs_data is not None and set(self.kwargs_data) != modalities:
            raise ValueError("kwargs_data must use the same modalities as mm_hashes")
        if self.mm_metadata is not None and set(self.mm_metadata) != modalities:
            raise ValueError("mm_metadata must use the same modalities as mm_hashes")

        flattened_ranges: list[tuple[int, int]] = []
        for modality in modalities:
            num_hashes = len(self.mm_hashes[modality])
            num_placeholders = len(self.mm_placeholders[modality])
            if num_hashes != num_placeholders:
                raise ValueError(
                    f"{modality} mm_hashes and mm_placeholders must have "
                    "the same length"
                )
            if (
                self.kwargs_data is not None
                and len(self.kwargs_data[modality]) != num_hashes
            ):
                raise ValueError(
                    f"{modality} kwargs_data and mm_hashes must have the same length"
                )
            if (
                self.mm_metadata is not None
                and len(self.mm_metadata[modality]) != num_hashes
            ):
                raise ValueError(
                    f"{modality} mm_metadata and mm_hashes must have the same length"
                )
            flattened_ranges.extend(
                (placeholder.offset, placeholder.offset + placeholder.length)
                for placeholder in self.mm_placeholders[modality]
            )

        flattened_ranges.sort()
        for (offset, end), (next_offset, _) in zip(
            flattened_ranges, flattened_ranges[1:]
        ):
            if next_offset < end:
                raise ValueError(
                    "mm_placeholders must be globally non-overlapping and sorted"
                )
        return self

kwargs_data = None class-attribute instance-attribute

Per-modality serialized tensor data.

Each value is a list parallel to mm_hashes[modality]. A str entry is a base64-encoded MultiModalKwargsItem; None means the item should be resolved from cache. The entire field is None for metadata-only (cache-hit) responses.

mm_hashes instance-attribute

Per-modality item hashes, e.g. {"image": ["abc", "def"]}.

mm_metadata = None class-attribute instance-attribute

Per-modality serialized metadata for disaggregated prefill.

Each value is a list parallel to mm_hashes[modality]. A str entry is a base64-encoded MultiModalKwargsItem containing only placeholder-metadata and keep_on_cpu fields. None means that the metadata is unavailable for that item. Prefill can use this instead of kwargs_data only when ec_transfer_params is also set, so embeddings arrive through the EC connector rather than from pixel_values.

mm_placeholders instance-attribute

Per-modality placeholder ranges in the token sequence.

PlaceholderRangeInfo

Bases: BaseModel

Serializable placeholder location for a single multi-modal item.

Attributes:

  • length (int) –

    Number of placeholder tokens.

  • offset (int) –

    Start index of the placeholder tokens in the prompt.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
class PlaceholderRangeInfo(BaseModel):
    """Serializable placeholder location for a single multi-modal item."""

    offset: int = Field(ge=0)
    """Start index of the placeholder tokens in the prompt."""

    length: int = Field(gt=0)
    """Number of placeholder tokens."""

length = Field(gt=0) class-attribute instance-attribute

Number of placeholder tokens.

offset = Field(ge=0) class-attribute instance-attribute

Start index of the placeholder tokens in the prompt.

output_mode_or_tokens(value)

Discriminator for GenerateResponse and GenerateStreamResponse.

A missing output_mode means tokens. Existing clients and servers that predate the field never send it.

Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
def output_mode_or_tokens(value: Any) -> Any:
    """Discriminator for `GenerateResponse` and `GenerateStreamResponse`.

    A missing `output_mode` means `tokens`. Existing clients and servers
    that predate the field never send it.
    """
    if isinstance(value, dict):
        return value.get("output_mode", "tokens")
    return getattr(value, "output_mode", "tokens")