Qwen3 parser for tool calls and reasoning.
Qwen3 XML tool call format::
<tool_call>
<function=func_name>
<parameter=key>value</parameter>
</function>
</tool_call>
The argument body consists of <parameter=NAME>VALUE</parameter> tags.
The _qwen3_arg_converter parses these into a JSON object.
Classes:
-
Qwen3Parser
–
Qwen3 parser: <think>/</think> reasoning +
Qwen3Parser
Bases: ParserEngine
Qwen3 parser: <think>/</think> reasoning +
<tool_call> XML tool calls in a single engine.
<tool_call> as implicit reasoning end (a grammar transition, so it
also feeds is_reasoning_end and the structured-output gate)
Subclasses that share the grammar but differ only in the four wrapper
token strings (reasoning + tool-call) override the class attributes
below; everything else is inherited unchanged.
Source code in vllm/parser/qwen3.py
| class Qwen3Parser(ParserEngine):
"""Qwen3 parser: ``<think>``/``</think>`` reasoning +
``<tool_call>`` XML tool calls in a single engine.
- ``<tool_call>`` as implicit reasoning end (a grammar transition, so it
also feeds ``is_reasoning_end`` and the structured-output gate)
Subclasses that share the grammar but differ only in the four wrapper
token strings (reasoning + tool-call) override the class attributes
below; everything else is inherited unchanged.
"""
CONFIG_NAME = "qwen3"
THINK_START = THINK_START
THINK_END = THINK_END
TOOL_START = TOOL_CALL_START
TOOL_END = TOOL_CALL_END
TURN_BOUNDARIES: frozenset[str] = CHATML_TURN_BOUNDARIES
def __init__(
self,
tokenizer: TokenizerLike,
tools: list[Tool] | None = None,
**kwargs,
) -> None:
chat_kwargs = kwargs.get("chat_template_kwargs", {}) or {}
self.thinking_enabled = chat_kwargs.get("enable_thinking", True)
kwargs.setdefault(
"parser_engine_config",
qwen3_config(
thinking=self.thinking_enabled,
name=self.CONFIG_NAME,
think_start=self.THINK_START,
think_end=self.THINK_END,
tool_start=self.TOOL_START,
tool_end=self.TOOL_END,
turn_boundary_tokens=self.TURN_BOUNDARIES,
),
)
super().__init__(
tokenizer,
tools,
**kwargs,
)
def extract_reasoning(
self,
model_output: str,
request: ChatCompletionRequest | ResponsesRequest,
) -> tuple[str | None, str | None]:
if not self.thinking_enabled:
return None, model_output
return super().extract_reasoning(model_output, request)
|
_trim_wrapping_newlines(value)
Strip one leading and one trailing newline (the Qwen3 template markup).
Source code in vllm/parser/qwen3.py
| def _trim_wrapping_newlines(value: str) -> str:
"""Strip one leading and one trailing newline (the Qwen3 template markup)."""
if value.startswith("\n"):
value = value[1:]
if value.endswith("\n"):
value = value[:-1]
return value
|