vllm.entrypoints.scale_out.token_in_token_out.protocol
¶
Classes:
-
DerenderChatRequest–Request for the /v1/chat/completions/derender endpoint (non-streaming).
-
DerenderChatStreamRequest–One chunk streaming derender request for /v1/chat/completions/derender.
-
DerenderChatStreamResponse–Response for one streaming chat derender chunk.
-
DerenderCompletionRequest–Request for the /v1/completions/derender endpoint (non-streaming).
-
DerenderCompletionStreamRequest–One chunk streaming derender request for /v1/completions/derender.
-
DerenderCompletionStreamResponse–Response for one streaming completions derender chunk.
-
DerenderStreamState–Per sequence state for stateless streaming derender.
-
GenerateChoiceBase–Fields shared by every
output_modeof a non-streaming choice. -
GenerateRequest– -
GenerateStreamChoiceBase–Fields shared by every
output_modeof a streaming choice. -
GenerateTextChoice– -
GenerateTextStreamChoice– -
GenerateTokensChoice– -
MultiModalFeatures–Lightweight multimodal metadata produced by the render step.
-
PlaceholderRangeInfo–Serializable placeholder location for a single multi-modal item.
Functions:
-
output_mode_or_tokens–Discriminator for
GenerateResponseandGenerateStreamResponse.
DerenderChatRequest
¶
Bases: BaseModel
Request for the /v1/chat/completions/derender endpoint (non-streaming).
Wraps a complete GenerateResponse and caller supplied metadata needed to produce a fully formed ChatCompletionResponse without a GPU.
Attributes:
-
chat_request(ChatCompletionRequest | None) –The original (post-adjust_request) ChatCompletionRequest from /render.
-
generate_response(GenerateTokensResponse) –The complete token-in / token-out engine response to derender.
-
model(str | None) –Served model name. Defaults to the server's served model name.
-
prompt_token_ids(list[NonNegativeInt] | None) –Prompt token IDs (
GenerateRequest.token_idsfrom /render). Seeds -
prompt_tokens(int | None) –Prompt token count for usage; defaults to 0 if omitted.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
chat_request = None
class-attribute
instance-attribute
¶
The original (post-adjust_request) ChatCompletionRequest from /render.
Required by the parsing so that tool/reasoning parsers can receive the full request context they expect (request.tools, request.tool_choice, request._grammar_from_parser, etc.).
generate_response
instance-attribute
¶
The complete token-in / token-out engine response to derender.
Only output_mode="tokens" responses are accepted. Other modes are
already detokenized and re-decoding their token_ids would undo the
engine's stop string truncation.
model = None
class-attribute
instance-attribute
¶
Served model name. Defaults to the server's served model name.
prompt_token_ids = None
class-attribute
instance-attribute
¶
Prompt token IDs (GenerateRequest.token_ids from /render). Seeds
detokenization from the prompt tail so the first output token keeps its
leading space on SentencePiece tokenizers. Falls back to
generate_response.prompt_token_ids, then to unseeded decoding.
Only the last few IDs are read, so a suffix of the prompt is enough.
prompt_tokens = None
class-attribute
instance-attribute
¶
Prompt token count for usage; defaults to 0 if omitted.
GenerateResponse carries only output tokens; the caller already has len(GenerateRequest.token_ids) from the render step.
DerenderChatStreamRequest
¶
Bases: BaseModel
One chunk streaming derender request for /v1/chat/completions/derender.
The client sends one request per SSE chunk received from
/inference/v1/generate. Each request carries the generate chunk
plus the stream_state returned by the previous call (None on the
first call). The response contains the derendered chunk and the updated
state to be passed to the next call.
This implements stateless no server side session. All mutable state lives in
the client carried stream_state.
Attributes:
-
chat_request(ChatCompletionRequest | None) –The original (post adjust_request) ChatCompletionRequest from /render.
-
generate_chunk(GenerateTokensStreamResponse) –One
output_mode="tokens"SSE chunk from/inference/v1/generate -
prompt_token_ids(list[NonNegativeInt] | None) –Prompt token IDs. Required by the parser path's
parse_deltato -
prompt_tokens(int | None) –Prompt token count for usage. Forwarded from the render step.
-
stream_state(DerenderStreamState | None) –Client carried detok state from the previous call.
Noneon first.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
chat_request = None
class-attribute
instance-attribute
¶
The original (post adjust_request) ChatCompletionRequest from /render.
generate_chunk
instance-attribute
¶
One output_mode="tokens" SSE chunk from /inference/v1/generate
(stream=True).
prompt_token_ids = None
class-attribute
instance-attribute
¶
Prompt token IDs. Required by the parser path's parse_delta to
settle its initial reasoning state (e.g. chat templates that pre-open
<think>). prompt_tokens is a usage count and cannot serve this
purpose. Sourced from GenerateRequest.token_ids at the render step.
Rejected with a 400 (by ServingDerender) when a tool or reasoning
parser is configured and this is omitted. Without it, parse_delta
cannot tell whether the prompt left reasoning open and would silently
misclassify reasoning content as plain content. On all paths it also
seeds detokenization on the first chunk, falling back to
generate_chunk.prompt_token_ids when omitted (see
DerenderChatRequest.prompt_token_ids).
With a parser configured, send the full list. Parsers look back to the last reasoning marker which can be anywhere in the prompt, so a suffix can flip the initial reasoning state. Without a parser, a suffix is enough.
prompt_tokens = None
class-attribute
instance-attribute
¶
Prompt token count for usage. Forwarded from the render step.
stream_state = None
class-attribute
instance-attribute
¶
Client carried detok state from the previous call. None on first.
DerenderChatStreamResponse
¶
Bases: BaseModel
Response for one streaming chat derender chunk.
Pairs the derendered SSE chunk with the updated client carried state to pass to the next call.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
DerenderCompletionRequest
¶
Bases: BaseModel
Request for the /v1/completions/derender endpoint (non-streaming).
Parallel to DerenderChatRequest but handles the multi-prompt completions case: one GenerateResponse per prompt, mirroring the list[GenerateRequest] returned by /v1/completions/render.
Attributes:
-
completion_request(CompletionRequest | None) –The original (post-adjust_request) CompletionRequest from /render.
-
generate_responses(list[GenerateTokensResponse]) –One
output_mode="tokens"response per prompt, parallel to the -
model(str | None) –Served model name. Defaults to the server's served model name.
-
prompt_token_ids(list[list[NonNegativeInt] | None] | None) –One prompt token ID list per response, used to seed detokenization.
-
prompt_tokens(list[int] | None) –One prompt token count per response; each defaults to 0 if omitted.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
completion_request = None
class-attribute
instance-attribute
¶
The original (post-adjust_request) CompletionRequest from /render.
Mirrors chat_request on DerenderChatRequest. Required by the parsing so parsers receive the full request context.
generate_responses
instance-attribute
¶
One output_mode="tokens" response per prompt, parallel to the
list[GenerateRequest] returned by /v1/completions/render.
model = None
class-attribute
instance-attribute
¶
Served model name. Defaults to the server's served model name.
prompt_token_ids = None
class-attribute
instance-attribute
¶
One prompt token ID list per response, used to seed detokenization.
See DerenderChatRequest.prompt_token_ids.
If provided, len(prompt_token_ids) must equal len(generate_responses).
A None entry falls back to generate_responses[i].prompt_token_ids.
prompt_tokens = None
class-attribute
instance-attribute
¶
One prompt token count per response; each defaults to 0 if omitted.
If provided, len(prompt_tokens) must equal len(generate_responses).
DerenderCompletionStreamRequest
¶
Bases: BaseModel
One chunk streaming derender request for /v1/completions/derender.
Parallel to DerenderChatStreamRequest for the completions endpoint.
Each call processes one SSE chunk (one output sequence's delta) and
returns the derendered chunk plus updated state.
Attributes:
-
completion_request(CompletionRequest | None) –The original (post adjust_request) CompletionRequest from /render.
-
generate_chunk(GenerateTokensStreamResponse) –One
output_mode="tokens"SSE chunk from/inference/v1/generate. -
prompt_token_ids(list[NonNegativeInt] | None) –Prompt token IDs, used on the first chunk to seed detokenization.
-
prompt_tokens(int | None) –Prompt token count for usage.
-
stream_state(DerenderStreamState | None) –Client-carried detok state.
Noneon the first call.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
completion_request = None
class-attribute
instance-attribute
¶
The original (post adjust_request) CompletionRequest from /render.
generate_chunk
instance-attribute
¶
One output_mode="tokens" SSE chunk from /inference/v1/generate.
prompt_token_ids = None
class-attribute
instance-attribute
¶
Prompt token IDs, used on the first chunk to seed detokenization.
Falls back to generate_chunk.prompt_token_ids. See
DerenderChatRequest.prompt_token_ids.
prompt_tokens = None
class-attribute
instance-attribute
¶
Prompt token count for usage.
stream_state = None
class-attribute
instance-attribute
¶
Client-carried detok state. None on the first call.
DerenderCompletionStreamResponse
¶
Bases: BaseModel
Response for one streaming completions derender chunk.
Parallel to DerenderChatStreamResponse for the completions endpoint.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
DerenderStreamState
¶
Bases: BaseModel
Per sequence state for stateless streaming derender.
The client carries this between successive per chunk HTTP calls to the streaming derender endpoint. All fields are plain JSON serializable data. No opaque tokenizer or parser internals are stored here.
Two separate sets of fields support two different streaming modes:
- For plain streaming (no parser configured including the completions path):
prev_tokens,prefix_offsetandread_offsetmaintain a bounded incremental decoding window. This requires O(window) transport and O(delta) computation per chunk. - For parser enabled chat streaming:
output_token_ids,output_chunk_lens,tools_streamedandlast_tool_call_idsare used to replayparse_delta()from scratch on every chunk because parser state cannot be serialized. This incurs O(n) transport per chunk (O(n²) per generation) and O(n²)parse_delta()calls per generation. Since many parsers re-scan the entire accumulated text on each invocation,parse_delta()itself is O(n) yielding a true worst case compute cost of O(n³) per generation. No caching is performed. Work is bounded bymax_model_len. SeeOnlineDerenderer._derender_chat_stream_parsed.
The detokenization strategy carries the incremental decode offsets
directly rather than re-sending the whole token history each chunk.
detokenize_incrementally only ever reads the trailing token window
prev_tokens[prefix_offset:], so we carry just that tail plus the two
offsets. Each chunk resumes exactly where the last one stopped, including
any partially processed multi-byte character (tracked by read_offset),
then trims and rebases the window so it never grows with generation length.
Performance:
- Compute per chunk is O(delta). One detokenize_incrementally call per
new token, independent of how many tokens preceded it.
- Transport per chunk is O(window). The carried tail is bounded by the
incremental detokenization offset, so cumulative bytes over the wire are
O(n) rather than the O(n^2) a full history round trip would incur.
Attributes:
-
last_tool_call_ids(list[str]) –Stable tool call IDs, assigned once when each call first appears.
-
logprob_context_token_ids(list[int]) –Trailing sampled token IDs carried across chunks so byte-fallback
-
logprob_text_offset(int) –Cumulative emitted text length. Seeds
text_offsetfor completion -
output_chunk_lens(list[Annotated[int, Field(gt=0)]]) –Token count of each chunk in
output_token_ids. Parser path only. -
output_token_ids(list[int]) –All output tokens seen so far. Parser path only.
-
prefix_offset(int) –Prefix offset into
prev_tokensfor incremental detokenization. -
prev_tokens(list[str]) –Trailing decode window. Token strings from
prefix_offsetonward. -
read_offset(int) –Read offset into
prev_tokensfor incremental detokenization. -
role_sent(bool) –True once the initial
role: "assistant"delta has been emitted. -
tools_streamed(bool) –True once a tool call delta has been emitted. Parser path only.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 | |
last_tool_call_ids = Field(default_factory=list)
class-attribute
instance-attribute
¶
Stable tool call IDs, assigned once when each call first appears.
Indexed by tool call index. Parser path only. Prevents ID regeneration when replay reprocesses a tool call that already has a pinned ID.
logprob_context_token_ids = Field(default_factory=list, max_length=4)
class-attribute
instance-attribute
¶
Trailing sampled token IDs carried across chunks so byte-fallback
(U+FFFD) correction during logprob placeholder resolution has context
at chunk boundaries. Bounded to the 4-token window that
_correct_decoded_token reads.
logprob_text_offset = Field(default=0, ge=0)
class-attribute
instance-attribute
¶
Cumulative emitted text length. Seeds text_offset for completion
streaming logprobs so offsets stay absolute across chunks, mirroring
initial_text_offset in the generate streaming path.
output_chunk_lens = Field(default_factory=list)
class-attribute
instance-attribute
¶
Token count of each chunk in output_token_ids. Parser path only.
Replay uses these to reproduce the original parse_delta call
boundaries. Must sum to len(output_token_ids).
output_token_ids = Field(default_factory=list)
class-attribute
instance-attribute
¶
All output tokens seen so far. Parser path only.
Replay buffer: each chunk rebuilds a fresh parser and replays every
token in here through parse_delta (discarding the result) before
processing the current chunk's tokens since parser internal state
cannot be serialized into this stateless model. Unavoidably O(n)
bounded by max_model_len (enforced server side, not by a field
validator here since the bound is model dependent).
prefix_offset = Field(default=0, ge=0)
class-attribute
instance-attribute
¶
Prefix offset into prev_tokens for incremental detokenization.
prev_tokens = Field(default_factory=list)
class-attribute
instance-attribute
¶
Trailing decode window. Token strings from prefix_offset onward.
Bounded, trimmed and rebased each chunk to the tail
detokenize_incrementally still reads, so it does not grow with the
number of chunks.
read_offset = Field(default=0, ge=0)
class-attribute
instance-attribute
¶
Read offset into prev_tokens for incremental detokenization.
role_sent = False
class-attribute
instance-attribute
¶
True once the initial role: "assistant" delta has been emitted.
Prevents re-emitting the role on subsequent chunks even when the detok window is transiently empty (e.g. usage only final chunk).
tools_streamed = False
class-attribute
instance-attribute
¶
True once a tool call delta has been emitted. Parser path only.
Drives the finish_reason -> "tool_calls" rewrite on the final
chunk mirroring the generate streaming path.
GenerateChoiceBase
¶
Bases: BaseModel
Fields shared by every output_mode of a non-streaming choice.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
GenerateRequest
¶
Bases: BaseModel
Methods:
-
is_sampling_param_provided–Whether the caller explicitly set
sampling_params.<name>.
Attributes:
-
content_parts(list[dict[str, Any]] | None) –Raw multimodal input; server resolves media. Mutually exclusive
-
features(MultiModalFeatures | None) –Multimodal hashes and placeholder positions (populated for MM inputs).
-
sampling_params(SamplingParams) –The sampling parameters for the model.
-
token_ids(list[int]) –The token ids to generate text from.
-
token_offsets(list[tuple[int, int]] | None) –Char-level (start, end) offsets per token, relative to the
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 | |
content_parts = None
class-attribute
instance-attribute
¶
Raw multimodal input; server resolves media. Mutually exclusive
with features.
features = None
class-attribute
instance-attribute
¶
Multimodal hashes and placeholder positions (populated for MM inputs).
sampling_params
instance-attribute
¶
The sampling parameters for the model.
token_ids = Field(min_length=1)
class-attribute
instance-attribute
¶
The token ids to generate text from.
token_offsets = None
class-attribute
instance-attribute
¶
Char-level (start, end) offsets per token, relative to the
tokenized source string. Present only when the request set
return_token_offsets=True and the renderer was able to compute
them (Fast tokenizer, text input, no multimodal data). List length
equals token_ids length when present. None otherwise.
is_sampling_param_provided(name)
¶
Whether the caller explicitly set sampling_params.<name>.
For requests parsed from a JSON body, this reflects the raw input
dict. For requests constructed with a pre-built SamplingParams
instance, all fields are considered provided so server-side defaults
do not clobber values already resolved upstream.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
GenerateStreamChoiceBase
¶
Bases: BaseModel
Fields shared by every output_mode of a streaming choice.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
GenerateTextChoice
¶
Bases: GenerateChoiceBase
Attributes:
-
logprobs(ChatCompletionLogProbs | None) –Logprobs with decoded token strings and
bytes. -
text(str) –Detokenized output. Excludes a matched stop string unless
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
GenerateTextStreamChoice
¶
Bases: GenerateStreamChoiceBase
Attributes:
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
text
instance-attribute
¶
Text delta since the previous chunk of this choice.
GenerateTokensChoice
¶
Bases: GenerateChoiceBase
Attributes:
-
logprobs(ChatCompletionLogProbs | None) –Logprobs whose tokens are
token_id:Nplaceholders.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
logprobs = None
class-attribute
instance-attribute
¶
Logprobs whose tokens are token_id:N placeholders.
MultiModalFeatures
¶
Bases: BaseModel
Lightweight multimodal metadata produced by the render step.
Carries hashes (for cache lookup / identification) and placeholder
positions so the downstream /generate service knows where in
the token sequence each multimodal item lives.
Attributes:
-
kwargs_data(dict[str, list[str | None]] | None) –Per-modality serialized tensor data.
-
mm_hashes(dict[str, list[str]]) –Per-modality item hashes, e.g.
{"image": ["abc", "def"]}. -
mm_metadata(dict[str, list[str | None]] | None) –Per-modality serialized metadata for disaggregated prefill.
-
mm_placeholders(dict[str, list[PlaceholderRangeInfo]]) –Per-modality placeholder ranges in the token sequence.
Source code in vllm/entrypoints/scale_out/token_in_token_out/protocol.py
kwargs_data = None
class-attribute
instance-attribute
¶
Per-modality serialized tensor data.
Each value is a list parallel to mm_hashes[modality]. A str
entry is a base64-encoded MultiModalKwargsItem; None means
the item should be resolved from cache. The entire field is
None for metadata-only (cache-hit) responses.
mm_hashes
instance-attribute
¶
Per-modality item hashes, e.g. {"image": ["abc", "def"]}.
mm_metadata = None
class-attribute
instance-attribute
¶
Per-modality serialized metadata for disaggregated prefill.
Each value is a list parallel to mm_hashes[modality]. A str
entry is a base64-encoded MultiModalKwargsItem containing only
placeholder-metadata and keep_on_cpu fields. None means that
the metadata is unavailable for that item. Prefill can use this
instead of kwargs_data only when ec_transfer_params is also
set, so embeddings arrive through the EC connector rather than from
pixel_values.
mm_placeholders
instance-attribute
¶
Per-modality placeholder ranges in the token sequence.
PlaceholderRangeInfo
¶
output_mode_or_tokens(value)
¶
Discriminator for GenerateResponse and GenerateStreamResponse.
A missing output_mode means tokens. Existing clients and servers
that predate the field never send it.