Skip to content

vllm.config.watermarking

Classes:

WatermarkConfig

Configuration for text watermark generation.

Attributes:

  • algorithm (WatermarkingAlgorithm) –

    Algorithm used to watermark generated text.

  • allow_target_only_watermarking (bool) –

    Allow speculative decoding without watermarking draft tokens.

  • alpha (float) –

    Probability of selecting key B for dual-key watermarking.

  • context_width (int) –

    Number of prior tokens used by the watermark PRF.

  • deduplicate_contexts (WatermarkContextScope) –

    Which history is searched for a repeated context before a token is

  • deduplicate_contexts_max_history (int | None) –

    Number of most recent history positions searched (default 8192), or

  • key (int) –

    Secret key used to watermark generated text.

  • prf (WatermarkPRFName) –

    Pseudorandom function used by the watermarking algorithm.

Source code in vllm/config/watermarking.py
@config
class WatermarkConfig:
    """Configuration for text watermark generation."""

    key: int = Field(ge=0, repr=False, exclude=True)
    """Secret key used to watermark generated text."""
    algorithm: WatermarkingAlgorithm = "gumbel"
    """Algorithm used to watermark generated text."""
    alpha: float = Field(default=0.1, ge=0, le=1)
    """Probability of selecting key B for dual-key watermarking."""
    context_width: int = Field(default=4, ge=1)
    """Number of prior tokens used by the watermark PRF."""
    deduplicate_contexts: WatermarkContextScope = "single_turn"
    """Which history is searched for a repeated context before a token is
    sampled; a repeated context is sampled without the watermark. `none`
    disables the search. `single_turn` (default) searches this request's
    generated tokens. `all` also searches the prompt and samples the first
    `context_width` generated tokens without watermarking."""
    deduplicate_contexts_max_history: int | None = Field(default=8192, ge=1)
    """Number of most recent history positions searched (default 8192), or
    `None` for the whole scope. Each position is compared over the
    `context_width` tokens before it. Ignored when `deduplicate_contexts` is
    `none`."""
    prf: WatermarkPRFName = "philox"
    """Pseudorandom function used by the watermarking algorithm."""
    allow_target_only_watermarking: bool = False
    """Allow speculative decoding without watermarking draft tokens."""

    @property
    def supports_speculative_decoding(self) -> bool:
        return self.algorithm == "dual_key_gumbel"

    @model_validator(mode="after")
    def validate_watermark_settings(self) -> Self:
        if self.key > 2**64 - 1:
            raise ValueError("philox keys must fit in 64 bits")
        history_is_too_short = (
            self.deduplicate_contexts_max_history is not None
            and self.deduplicate_contexts_max_history < _MIN_RECOMMENDED_DEDUP_HISTORY
        )
        if self.algorithm in ("gumbel", "dual_key_gumbel") and (
            self.deduplicate_contexts == "none" or history_is_too_short
        ):
            logger.warning_once(
                "Gumbel-max watermarking with context deduplication "
                "disabled or limited to fewer than "
                f"{_MIN_RECOMMENDED_DEDUP_HISTORY} positions may increase the "
                "frequency of degenerate generations, including repetition loops. "
                "Use deduplicate_contexts='single_turn' or 'all' with "
                "deduplicate_contexts_max_history at least "
                f"{_MIN_RECOMMENDED_DEDUP_HISTORY} or null to mitigate this.",
                scope="global",
            )
        return self

algorithm = 'gumbel' class-attribute instance-attribute

Algorithm used to watermark generated text.

allow_target_only_watermarking = False class-attribute instance-attribute

Allow speculative decoding without watermarking draft tokens.

alpha = Field(default=0.1, ge=0, le=1) class-attribute instance-attribute

Probability of selecting key B for dual-key watermarking.

context_width = Field(default=4, ge=1) class-attribute instance-attribute

Number of prior tokens used by the watermark PRF.

deduplicate_contexts = 'single_turn' class-attribute instance-attribute

Which history is searched for a repeated context before a token is sampled; a repeated context is sampled without the watermark. none disables the search. single_turn (default) searches this request's generated tokens. all also searches the prompt and samples the first context_width generated tokens without watermarking.

deduplicate_contexts_max_history = Field(default=8192, ge=1) class-attribute instance-attribute

Number of most recent history positions searched (default 8192), or None for the whole scope. Each position is compared over the context_width tokens before it. Ignored when deduplicate_contexts is none.

key = Field(ge=0, repr=False, exclude=True) class-attribute instance-attribute

Secret key used to watermark generated text.

prf = 'philox' class-attribute instance-attribute

Pseudorandom function used by the watermarking algorithm.