Derenderer APIs¶
The derenderer API is the post processing counterpart to the Renderer APIs. Where /render turns a request into token ID (preprocessing), /derender turns generated token IDs back into a fully formed OpenAI compatible response (detokenization, reasoning parsing, tool call parsing), all without a GPU.
This closes the loop for a token-in / token-out engine in disaggregated serving:
- GPU less post processing: Detokenization, reasoning parsing, and tool call parsing run on the same GPU less frontend that hosts
/render - Parser parity: The derenderer reuses vLLM's tool and reasoning parsers, so a disaggregated deployment produces the same
content/reasoning/tool_callssplit as a standardvllm serveserver - Streaming and one-shot parsing: Non-streaming calls send a complete
GenerateResponsewith all token IDs present and get one-shot parsing. Both endpoints also acceptstream: true, taking oneGenerateStreamResponsedelta plus a client carriedstream_stateand returning{chunk, stream_state}. The chat endpoint's streaming path supports reasoning and tool call parsing while emitting the samereasoning/content/tool_callsdeltas the generate streaming path would for the same token IDs (see Streaming)
Both endpoints are hosted by the GPU less rendering server started with vllm launch render, alongside the /render
endpoints.
Pipeline¶
render generate derender
request ───────────────▶ token_ids ─────────▶ token_ids ──────────▶ response
(chat / (GPU less) (token-in / (GPU less) (OpenAI
completion) │ token-out engine) ▲ compatible)
└─────────────── request + prompt_tokens ──┘
The derender step needs more than the engine's token_ids. It also consumes the original chat_request/completion_request and prompt_tokens carried over from the render step (see Request format) so the tool and reasoning parsers have the context they need. The chat streaming path additionally needs prompt_token_ids (see Streaming cost) when a parser is configured.
All derender endpoints also accept an optional prompt_token_ids (the token_ids of the GenerateRequest returned by /render). It's used to seed detokenization from the end of the prompt, the same way the engine does. Without it, the first output token is decoded as if it started a new sequence. On SentencePiece tokenizers (Llama-2, Phi-3) that drops its leading space compared to /v1/chat/completions and /v1/completions. If it's omitted, derender falls back to the generate response's prompt_token_ids (set when the GenerateRequest has return_token_ids), then to decoding without prompt context.
Seeding only reads the last few IDs (7 today), so a suffix of the prompt is enough and a long prompt doesn't need to travel back to derender. The exception is streaming chat with a tool or reasoning parser configured. There the parser also reads prompt_token_ids to decide whether the prompt left reasoning open and it can look back to any point in the prompt. In this case, send the full list (see Streaming cost).
API Reference¶
- Chat Completions Derender API (
/v1/chat/completions/derender)- Post process a single
GenerateResponseinto aChatCompletionResponse - With
stream: true, post process oneGenerateStreamResponsechunk into aChatCompletionStreamResponsechunk
- Post process a single
- Completions Derender API (
/v1/completions/derender)- Post process a list of
GenerateResponseobjects (one per prompt) into aCompletionResponse - With
stream: true, post process oneGenerateStreamResponsechunk into aCompletionStreamResponsechunk
- Post process a list of
Request format¶
Each request wraps the engine's GenerateResponse(s) together with the caller metadata needed to reconstruct the final response without a GPU.
Only output_mode: "tokens" responses are accepted which is the default. A response from a request with output_mode: "text" gets a 400 because its text is already detokenized with stop strings removed. Responses without output_mode, from older servers, are treated as tokens.
/v1/chat/completions/derender:
Code
stream: Literal[False] = False
model: str | None = None
"""Served model name. Defaults to the server's served model name."""
generate_response: GenerateTokensResponse
"""The complete token-in / token-out engine response to derender.
Only `output_mode="tokens"` responses are accepted. Other modes are
already detokenized and re-decoding their `token_ids` would undo the
engine's stop string truncation.
"""
prompt_tokens: int | None = None
"""Prompt token count for usage; defaults to 0 if omitted.
GenerateResponse carries only output tokens; the caller already has
len(GenerateRequest.token_ids) from the render step.
"""
prompt_token_ids: list[NonNegativeInt] | None = None
"""Prompt token IDs (`GenerateRequest.token_ids` from /render). Seeds
detokenization from the prompt tail so the first output token keeps its
leading space on SentencePiece tokenizers. Falls back to
`generate_response.prompt_token_ids`, then to unseeded decoding.
Only the last few IDs are read, so a suffix of the prompt is enough.
"""
chat_request: ChatCompletionRequest | None = None
"""The original (post-adjust_request) ChatCompletionRequest from /render.
Required by the parsing so that tool/reasoning parsers can receive the full
request context they expect (request.tools, request.tool_choice,
request._grammar_from_parser, etc.).
"""
/v1/completions/derender:
Code
stream: Literal[False] = False
model: str | None = None
"""Served model name. Defaults to the server's served model name."""
generate_responses: list[GenerateTokensResponse]
"""One `output_mode="tokens"` response per prompt, parallel to the
list[GenerateRequest] returned by /v1/completions/render."""
prompt_tokens: list[int] | None = None
"""One prompt token count per response; each defaults to 0 if omitted.
If provided, len(prompt_tokens) must equal len(generate_responses).
"""
prompt_token_ids: list[list[NonNegativeInt] | None] | None = None
"""One prompt token ID list per response, used to seed detokenization.
See `DerenderChatRequest.prompt_token_ids`.
If provided, len(prompt_token_ids) must equal len(generate_responses).
A `None` entry falls back to `generate_responses[i].prompt_token_ids`.
"""
completion_request: CompletionRequest | None = None
"""The original (post-adjust_request) CompletionRequest from /render.
Mirrors chat_request on DerenderChatRequest. Required by the parsing
so parsers receive the full request context.
"""
Streaming requests set stream: true and carry one generate chunk (generate_chunk) plus stream_state instead of a complete response.
/v1/chat/completions/derender with stream: true:
Code
stream: Literal[True]
model: str | None = None
generate_chunk: GenerateTokensStreamResponse
"""One `output_mode="tokens"` SSE chunk from `/inference/v1/generate`
(`stream=True`)."""
stream_state: DerenderStreamState | None = None
"""Client carried detok state from the previous call. `None` on first."""
prompt_tokens: int | None = None
"""Prompt token count for usage. Forwarded from the render step."""
prompt_token_ids: list[NonNegativeInt] | None = None
"""Prompt token IDs. Required by the parser path's `parse_delta` to
settle its initial reasoning state (e.g. chat templates that pre-open
`<think>`). `prompt_tokens` is a usage count and cannot serve this
purpose. Sourced from `GenerateRequest.token_ids` at the render step.
Rejected with a 400 (by `ServingDerender`) when a tool or reasoning
parser is configured and this is omitted. Without it, `parse_delta`
cannot tell whether the prompt left reasoning open and would silently
misclassify reasoning content as plain content. On all paths it also
seeds detokenization on the first chunk, falling back to
`generate_chunk.prompt_token_ids` when omitted (see
`DerenderChatRequest.prompt_token_ids`).
With a parser configured, send the full list. Parsers look back to the
last reasoning marker which can be anywhere in the prompt, so a suffix
can flip the initial reasoning state. Without a parser, a suffix is
enough.
"""
chat_request: ChatCompletionRequest | None = None
"""The original (post adjust_request) ChatCompletionRequest from /render."""
/v1/completions/derender with stream: true:
Code
stream: Literal[True]
model: str | None = None
generate_chunk: GenerateTokensStreamResponse
"""One `output_mode="tokens"` SSE chunk from `/inference/v1/generate`."""
stream_state: DerenderStreamState | None = None
"""Client-carried detok state. `None` on the first call."""
prompt_tokens: int | None = None
"""Prompt token count for usage."""
prompt_token_ids: list[NonNegativeInt] | None = None
"""Prompt token IDs, used on the first chunk to seed detokenization.
Falls back to `generate_chunk.prompt_token_ids`. See
`DerenderChatRequest.prompt_token_ids`.
"""
completion_request: CompletionRequest | None = None
"""The original (post adjust_request) CompletionRequest from /render."""
Both return {"chunk": ..., "stream_state": ...}. chunk is a ChatCompletionStreamResponse or CompletionStreamResponse and stream_state goes with the next call.
Oversized payloads are rejected with a 400 before any tokenizer.decode() or parser runs.
Parser configuration¶
The derenderer builds its tool and reasoning parsers from its own server flags plus the chat_request sent with each call. It can't see how the request was rendered or served anywhere else, so for its output to match a vllm serve server:
- Start the render server with the same
--tool-call-parser,--reasoning-parser,--enable-auto-tool-choice,--chat-templateand--default-chat-template-kwargsyou'd givevllm servefor this model. If/renderand/derenderrun on different servers, give both the same values. - Send the full
chat_requestthat went to/render, not justmessagesandtools. Fields likechat_template_kwargs,reasoning_effort,tool_choiceandinclude_reasoningchange how the output is parsed.
A mismatch doesn't fail. The parser just splits reasoning, content and tool_calls differently from what vllm serve would return.
/v1/completions/derender only detokenizes. It never runs tool or reasoning parsers, the same as /v1/completions on vllm serve. It only reads skip_special_tokens from completion_request.
Streaming¶
Streaming derender mirrors /inference/v1/generate. Set stream: true in the body on the same path and the endpoint takes one generate stream chunk instead of a complete response. It answers with JSON, not SSE. Each call turns one generate chunk into one derendered chunk and the client forwards that to its own caller.
The server keeps no state between calls. Everything the next call needs is in stream_state which the client carries:
- Echo
stream_stateback. Leave it out (or sendnull) on the first call, then send thestream_statefrom each response with the next request, unchanged. State that doesn't add up (e.g.output_chunk_lensthat don't sum to the number of output tokens) is rejected with a400. - One choice per chunk.
/inference/v1/generatesends one choice per SSE event and each derender call accepts at most one. More than one is rejected with a400. n > 1is N streams. Generate interleaves chunks for different choices and tags each with itsindex. Keep onestream_stateper index and send each chunk with the state for its index. Each index gets its ownroledelta on its first chunk.- Send every chunk through, including the last ones. The finish reason usually arrives with the final tokens. A chunk with a
finish_reasonand no tokens is accepted too and flushes any buffered tool call arguments. Withstream_options: {"include_usage": true}, generate ends with a usage only chunk (choices: []). Send it through derender like any other chunk to get the usage chunk for the response.usage.prompt_tokensis the request'sprompt_tokensif set, otherwise the generate chunk's. - Send the same context on every call.
chat_requestandprompt_token_idsaren't kept between calls, so they go with every chunk including the usage chunk. Both are required when a tool or reasoning parser is configured.prompt_token_idsis thetoken_idsof theGenerateRequestreturned by/render. - Send
prompt_token_idson the first chunk, even without a parser. It seeds detokenization so the first token's leading space matches the coupled endpoint (see Request format). Later chunks already carry that context instream_state. - Don't forward
[DONE]. It marks the end of the generate stream and isn't a chunk.
Streaming derender resolves logprobs on the plain detokenization path, including token_id:N placeholders, same as the non-streaming endpoints (see Streaming state and logprobs). The parser path doesn't resolve them yet and drops them.
Streaming cost¶
Streaming derender threads a client carried stream_state across per-chunk calls instead of keeping session state on the server.
Plain detokenization (no parser configured) carries only a small, bounded incremental decode window in stream_state independent of generation length.
When a tool or reasoning parser is configured, parser internal state (buffered markup, reasoning/tool phase) can't be serialized. stream_state therefore instead carries the full output_token_ids seen so far plus output_chunk_lens, the token count of each chunk they arrived in. Each chunk rebuilds a fresh parser and replays that history through parse_delta, one call per original chunk, before processing the new tokens for real. For the parser path only, this means:
- Transport:
output_token_idsandoutput_chunk_lensround-trip in full in both directions on every call.output_chunk_lenshas one entry per chunk, which is one per token without speculative decoding. This means O(n) bytes per chunk, O(n²) bytes over a full generation. Bounded bymax_model_len.prompt_token_idsis sent in full on every call too and it isn't trimmed asoutput_token_idsgrows. This means that for most of a stream it dominates the per chunk payload. A 100k token prompt with 1k tokens of output meansprompt_token_idsis ~99% of the request body on every chunk. - Compute: replay is O(n)
parse_deltacalls per chunk (O(n²) per generation).parse_deltaitself is O(n) for parsers that re-scan accumulated text (e.g. Hermes tool-call JSON, DeepSeek-R1 reasoning). The per-generation cost is O(n³) character work, not O(n²). This is a deliberately minimal first implementation with no caching layer. - The parser path also requires
prompt_token_idssoparse_deltacan settle whether the prompt left reasoning open or not. Since parser state can't be carried across calls, it re-scans the full prompt once per chunk. Send the full list and not a suffix because the last reasoning marker can be anywhere in the prompt. - Replay runs off the event loop on the renderer's executor (
renderer_num_workers, default1). Size it for the expected number of concurrent parser configured streams.
output_token_ids and prompt_token_ids are both bounded by max_model_len but callers streaming long reasoning traces through a parser configured model should expect materially more state transport and CPU cost than the plain detokenization path.
Streaming state and logprobs¶
Streaming derender is stateless on the server side: all mutable state lives in the client-carried stream_state (DerenderStreamState), passed back on every per-chunk call. Besides the bounded incremental detokenization window (prev_tokens, prefix_offset, read_offset), role_sent, and the parser path's replay fields (see Streaming cost), the state carries two fields for logprob handling:
logprob_context_token_ids: the trailing sampled token IDs (at most 4) from previous chunks, used to seed byte-fallback (U+FFFD) correction so multi-byte characters whose tokens split across chunk boundaries still resolve to real stringslogprob_text_offset: the cumulative emitted text length, sotext_offsetin completion streaming logprobs stays absolute across chunks instead of restarting at 0
When a streamed GenerateTokensStreamChoice carries logprobs, the token_id:N placeholders are resolved per chunk and the resolved logprobs are attached to the corresponding streamed choice. For chat it is as ChatCompletionLogProbs and for completions it is converted to the flat CompletionLogProbs lists. Chunks without logprobs produce choices with logprobs: null.
Example¶
The example below drives the full render → generate → derender round trip for a chat request against a GPU less render server (/render, /derender) and a token-in / token-out engine (/inference/v1/generate).
Launch the two servers first:
vllm launch render meta-llama/Llama-3.2-1B-Instruct --port 8100
vllm serve meta-llama/Llama-3.2-1B-Instruct --tokens-only --port 8200
import httpx
MODEL = "meta-llama/Llama-3.2-1B-Instruct"
RENDER = "http://localhost:8100" # vllm launch render ...
ENGINE = "http://localhost:8200" # token-in / token-out engine
chat_request = {
"model": MODEL,
"messages": [{"role": "user", "content": "What is 2+2?"}],
"max_tokens": 32,
}
with httpx.Client(timeout=60.0) as client:
# 1. Render: request -> token IDs (GPU less)
generate_request = client.post(
f"{RENDER}/v1/chat/completions/render", json=chat_request
).json()
prompt_token_ids = generate_request["token_ids"]
# 2. Generate: token IDs -> token IDs (token-in / token-out engine)
generate_response = client.post(
f"{ENGINE}/inference/v1/generate", json=generate_request
).json()
# 3. Derender: token IDs -> ChatCompletionResponse (GPU less)
response = client.post(
f"{RENDER}/v1/chat/completions/derender",
json={
"model": MODEL,
"generate_response": generate_response,
"prompt_tokens": len(prompt_token_ids),
"prompt_token_ids": prompt_token_ids,
"chat_request": chat_request,
},
).json()
print(response["choices"][0]["message"]["content"])
Passing chat_request lets the derenderer run the configured tool and reasoning parsers. This means response["choices"][0]["message"] carries the same content / reasoning / tool_calls split a vllm serve server would produce. chat_request can only be omitted for a model with no tool or reasoning parser configured. A parser configured model rejects a missing chat_request with a 400 rather than silently falling back to plain detokenization.
Streaming example¶
The same round trip, streamed. /render keeps stream and stream_options on the GenerateRequest it returns, so the generate call streams too. Each generate chunk goes through derender with the stream_state from the previous call (see Streaming).
import json
import httpx
MODEL = "meta-llama/Llama-3.2-1B-Instruct"
RENDER = "http://localhost:8100" # vllm launch render ...
ENGINE = "http://localhost:8200" # token-in / token-out engine
chat_request = {
"model": MODEL,
"messages": [{"role": "user", "content": "What is 2+2?"}],
"max_tokens": 32,
"stream": True,
"stream_options": {"include_usage": True},
}
with httpx.Client(timeout=60.0) as client:
# 1. Render: request -> token IDs (GPU less)
generate_request = client.post(
f"{RENDER}/v1/chat/completions/render", json=chat_request
).json()
prompt_token_ids = generate_request["token_ids"]
# 2. Generate: stream token IDs (token-in / token-out engine)
stream_state = None # one per choice index when n > 1
with client.stream(
"POST", f"{ENGINE}/inference/v1/generate", json=generate_request
) as generate_stream:
for line in generate_stream.iter_lines():
if not line.startswith("data: ") or line == "data: [DONE]":
continue
# 3. Derender: one generate chunk -> one chat completion chunk
derendered = client.post(
f"{RENDER}/v1/chat/completions/derender",
json={
"stream": True,
"model": MODEL,
"generate_chunk": json.loads(line[len("data: ") :]),
"stream_state": stream_state,
"prompt_tokens": len(prompt_token_ids),
"prompt_token_ids": prompt_token_ids,
"chat_request": chat_request,
},
).json()
stream_state = derendered["stream_state"]
chunk = derendered["chunk"]
for choice in chunk["choices"]:
print(choice["delta"].get("content") or "", end="", flush=True)
if chunk.get("usage"):
print(f"\n{chunk['usage']}")
With a tool or reasoning parser configured, delta also carries reasoning and tool_calls, the same deltas /v1/chat/completions streams for those tokens.