vllm.renderers.inkling
¶
Native Inkling chat renderer for the Python frontend.
Mirrors the Rust frontend's native Inkling renderer
(rust/src/chat/src/renderer/inkling/mod.rs): chat messages are rendered
directly to token ids — Inkling has no Jinja chat template and no faithful
text form.
The encoding logic lives in inkling_encoding.py behind a narrow
"OpenAI messages + tools -> token ids" call; see the swap-point comment
in :meth:InklingRenderer._render for adopting a standalone Inkling
input-processing library (mistral-common style) later.
_HfBackedTmlTokenizer
¶
Adapts an HF tokenizer to the encoding core's tokenizer protocol.
Special-token ids are resolved from the tokenizer vocab at
construction (never hardcoded), trying each known spelling — the Inkling
HF vocab exposes some semantic slots as <|unused_NNNNNN|> tokens.