vllm.config.multimodal
¶
Classes:
-
AudioDummyOptions–Options for generating dummy audio data during profiling.
-
BaseDummyOptions–Base options for generating dummy data during profiling.
-
ImageDummyOptions–Options for generating dummy image data during profiling.
-
MultiModalConfig–Controls the behavior of multimodal models.
-
MultiModalDummyOptions–Dummy data options for each modality.
-
VideoDummyOptions–Options for generating dummy video data during profiling.
Attributes:
-
MMProcessorDevice(TypeAlias) –"auto","cpu", or the platform's own accelerator name
MMProcessorDevice = str
module-attribute
¶
"auto", "cpu", or the platform's own accelerator name
(current_platform.device_type, e.g. "cuda" on CUDA and ROCm,
"xpu" on XPU). Validated against that set by the CLI.
AudioDummyOptions
¶
Bases: BaseDummyOptions
Options for generating dummy audio data during profiling.
Source code in vllm/config/multimodal.py
BaseDummyOptions
¶
ImageDummyOptions
¶
Bases: BaseDummyOptions
Options for generating dummy image data during profiling.
Source code in vllm/config/multimodal.py
MultiModalConfig
¶
Controls the behavior of multimodal models.
Methods:
-
compute_hash–WARNING: Whenever a new field is added to this config,
-
fold_mm_processor_device–Fold the
mm_processor_deviceconvenience flag into the kwargs. -
get_limit_per_prompt–Get the maximum number of input items allowed per prompt
-
get_mm_processor_device_type–The torch device type
mm_processor_kwargs["device"]names. -
get_video_pruning_spec–Return
(method, rate)when video pruning is enabled, else None. -
merge_mm_processor_kwargs–Get the keyword arguments to pass to the multi-modal processor
-
use_gpu_video_backend–Return whether the configured video loader or codec uses the GPU.
-
validate_mm_processor_device–Check
mm_processor_kwargs["device"]for this deployment.
Attributes:
-
allow_missing_mm_embeddings(bool) –Whether a pre-computed-embedding input may omit the
*_embedstensor. -
enable_mm_embeds(bool) –If
True, enables passing multimodal embeddings: -
interleave_mm_strings(bool) –Enable fully interleaved support for multimodal prompts, while using
-
language_model_only(bool) –If True, disables all multimodal inputs by setting all modality limits to 0.
-
limit_per_prompt(MultiModalDummyOptions) –The maximum number of input items and options allowed per
-
media_io_kwargs(dict[str, dict[str, Any]]) –Additional args passed to process media inputs, keyed by modalities.
-
mm_device_do_normalize(bool | None) –Move the do_normalize computation in the mm preprocessing to before the ViT,
-
mm_encoder_attn_backend(AttentionBackendEnum | None) –Optional override for the multi-modal encoder attention backend when
-
mm_encoder_attn_dtype(Literal['fp8'] | None) –Optional dtype override for ViT encoder attention. Set to
"fp8"to -
mm_encoder_fp8_scale_path(str | None) –Path to a JSON file containing per-layer FP8 Q/K/V scales for ViT
-
mm_encoder_fp8_scale_save_margin(float) –Safety margin multiplied onto scales when auto-saving. A value > 1
-
mm_encoder_fp8_scale_save_path(str | None) –When set with dynamic FP8 scaling (
mm_encoder_attn_dtype="fp8" -
mm_encoder_only(bool) –When enabled, skips the language component of the model.
-
mm_encoder_tp_mode(MMEncoderTPMode) –Indicates how to optimize multi-modal encoder inference using tensor
-
mm_hasher_algorithm(MMHasherAlgorithm) –Hash algorithm to use for multi-modal input caching. Use
"sha256"or -
mm_ipc_gpu_memory_gb(float) –Amount of GPU memory (in GiB) sequestered on the engine's device for
-
mm_processor_cache_gb(float) –The size (in GiB) of the multi-modal processor cache, which is used to
-
mm_processor_cache_type(MMCacheType) –Type of cache to use for the multi-modal preprocessor/mapper. If
shm, -
mm_processor_kwargs(dict[str, object] | None) –Arguments to be forwarded to the model's processor for multi-modal data,
-
mm_shm_cache_max_object_size_mb(int) –Size limit (in MiB) for each object stored in the multi-modal processor
-
mm_tensor_ipc(MMTensorIPC) –IPC (inter-process communication) method for multimodal tensors.
-
skip_mm_profiling(bool) –When enabled, skips multimodal memory profiling and only profiles with
-
video_pruning_method(VideoPruningMethod) –Video token pruning algorithm applied when
video_pruning_rate> 0: -
video_pruning_rate(float | None) –Fraction of video tokens to prune from each video. Value sits in range
Source code in vllm/config/multimodal.py
300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 | |
allow_missing_mm_embeddings = False
class-attribute
instance-attribute
¶
Whether a pre-computed-embedding input may omit the *_embeds tensor.
In an encode/prefill/decode (EPD) deployment the encoder instance publishes embeddings through the EC connector. An EC consumer loads those embeddings from the connector, while a KV consumer receives the resulting prompt KV cache. Their requests only need the grid/size metadata that sizes the placeholder range.
Derived, not user-settable: VllmConfig.__post_init__ sets this to True
on EC and KV consumers. Everywhere else it stays False so that a request
which forgets its embeddings still fails fast in the frontend, with a clear
error, rather than deep inside the model.
enable_mm_embeds = False
class-attribute
instance-attribute
¶
If True, enables passing multimodal embeddings:
for LLM class, this refers to tensor inputs under multi_modal_data;
for the OpenAI-compatible server, this refers to chat messages with content
"type": "*_embeds".
When enabled with --limit-mm-per-prompt set to 0 for a modality,
precomputed embeddings skip count validation for that modality,
saving memory by not loading encoder modules while still enabling
embeddings as an input. Limits greater than 0 still apply to embeddings.
WARNING: The vLLM engine may crash if incorrect shape of embeddings is passed. Only enable this flag for trusted users!
interleave_mm_strings = False
class-attribute
instance-attribute
¶
Enable fully interleaved support for multimodal prompts, while using --chat-template-content-format=string.
language_model_only = False
class-attribute
instance-attribute
¶
If True, disables all multimodal inputs by setting all modality limits to 0.
Equivalent to setting --limit-mm-per-prompt to 0 for every modality.
limit_per_prompt = Field(default_factory=MultiModalDummyOptions)
class-attribute
instance-attribute
¶
The maximum number of input items and options allowed per prompt for each modality.
Defaults to 999 for each modality.
Legacy format (count only):
Configurable format (with options): {"video": {"count": 1, "num_frames": 32, "width": 512, "height": 512}, "image": {"count": 5, "width": 512, "height": 512}}
Mixed format (combining both): {"image": 16, "video": {"count": 1, "num_frames": 32, "width": 512, "height": 512}}
media_io_kwargs = Field(default_factory=dict)
class-attribute
instance-attribute
¶
Additional args passed to process media inputs, keyed by modalities.
For example, to set num_frames for video, set
--media-io-kwargs '{"video": {"num_frames": 40} }'
mm_device_do_normalize = True
class-attribute
instance-attribute
¶
Move the do_normalize computation in the mm preprocessing to before the ViT, and let the device do it, so that CPU computation can be saved.
mm_encoder_attn_backend = None
class-attribute
instance-attribute
¶
Optional override for the multi-modal encoder attention backend when
using vision transformers. Accepts any value from
vllm.v1.attention.backends.registry.AttentionBackendEnum (e.g. FLASH_ATTN).
mm_encoder_attn_dtype = None
class-attribute
instance-attribute
¶
Optional dtype override for ViT encoder attention. Set to "fp8" to
enable FP8 quantization via the FlashInfer cuDNN backend. When set to
"fp8" without a scale file, dynamic scaling is used automatically.
See docs/features/quantization/fp8_vit_attn.md for details.
mm_encoder_fp8_scale_path = None
class-attribute
instance-attribute
¶
Path to a JSON file containing per-layer FP8 Q/K/V scales for ViT
encoder attention. When provided (with mm_encoder_attn_dtype="fp8"),
static scaling is used. When omitted, dynamic scaling is used.
mm_encoder_fp8_scale_save_margin = Field(default=1.5, gt=0.0)
class-attribute
instance-attribute
¶
Safety margin multiplied onto scales when auto-saving. A value > 1 leaves headroom so that inputs with larger activations than the calibration set do not overflow FP8 range. Default 1.5.
mm_encoder_fp8_scale_save_path = None
class-attribute
instance-attribute
¶
When set with dynamic FP8 scaling (mm_encoder_attn_dtype="fp8"
and no mm_encoder_fp8_scale_path), saves the calibrated scales to
this file after the amax history buffer is full. The saved file can
then be used as mm_encoder_fp8_scale_path in subsequent runs.
mm_encoder_only = False
class-attribute
instance-attribute
¶
When enabled, skips the language component of the model.
This is usually only valid in disaggregated Encoder process.
mm_encoder_tp_mode = 'weights'
class-attribute
instance-attribute
¶
Indicates how to optimize multi-modal encoder inference using tensor parallelism (TP).
"weights": Within the same vLLM engine, split the weights of each layer across TP ranks. (default TP behavior)"data": Within the same vLLM engine, split the batched input data across TP ranks to process the data in parallel, while hosting the full weights on each TP rank. This batch-level DP is not to be confused with API request-level DP (which is controlled by--data-parallel-size). This is only supported on a per-model basis and falls back to"weights"if the encoder does not support DP.
mm_hasher_algorithm = 'blake3'
class-attribute
instance-attribute
¶
Hash algorithm to use for multi-modal input caching. Use "sha256" or
"sha512" for FIPS-compliant deployments.
mm_ipc_gpu_memory_gb = Field(default=0, ge=0)
class-attribute
instance-attribute
¶
Amount of GPU memory (in GiB) sequestered on the engine's device for GPU-side multimodal work in the API-server (frontend) process, such as hardware video decoding.
This budget is carved out of the engine's KV-cache memory so the headroom physically exists, and frontend GPU decode paths acquire from a blocking byte-counting semaphore of this size before allocating on the device.
Set to 0 (default) to disable frontend GPU multimodal memory gating.
mm_processor_cache_gb = Field(default=4, ge=0)
class-attribute
instance-attribute
¶
The size (in GiB) of the multi-modal processor cache, which is used to avoid re-processing past multi-modal inputs.
This cache is duplicated for each API process and engine core process,
resulting in a total memory usage of
mm_processor_cache_gb * (api_server_count + data_parallel_size).
A single processed item larger than this budget is served uncached (with a warning) instead of failing. Raise this value to cache such items.
Set to 0 to disable this cache completely (not recommended).
mm_processor_cache_type = 'lru'
class-attribute
instance-attribute
¶
Type of cache to use for the multi-modal preprocessor/mapper. If shm,
use shared memory FIFO cache. If lru, use mirrored LRU cache.
mm_processor_kwargs = None
class-attribute
instance-attribute
¶
Arguments to be forwarded to the model's processor for multi-modal data,
e.g., image processor. Overrides for the multi-modal processor obtained
from transformers.AutoProcessor.from_pretrained.
The available overrides depend on the model that is being run.
For example, for Phi-3-Vision:
{"num_crops": 4}.
mm_shm_cache_max_object_size_mb = Field(default=128, ge=0)
class-attribute
instance-attribute
¶
Size limit (in MiB) for each object stored in the multi-modal processor
shared memory cache. Only effective when mm_processor_cache_type is
"shm".
mm_tensor_ipc = 'direct_rpc'
class-attribute
instance-attribute
¶
IPC (inter-process communication) method for multimodal tensors. - "direct_rpc": Use msgspec serialization via RPC - "torch_shm": Use torch.multiprocessing shared memory for zero-copy IPC Defaults to "direct_rpc".
skip_mm_profiling = False
class-attribute
instance-attribute
¶
When enabled, skips multimodal memory profiling and only profiles with language backbone model during engine initialization.
This reduces engine startup time but shifts the responsibility to users for estimating the peak memory usage of the activation of multimodal encoder and embedding cache.
video_pruning_method = 'evs'
class-attribute
instance-attribute
¶
Video token pruning algorithm applied when video_pruning_rate > 0:
- "evs": Efficient Video Sampling.
- "vidcom2": Video Compression Commander.
video_pruning_rate = Field(default=None, ge=0.0, lt=1.0)
class-attribute
instance-attribute
¶
Fraction of video tokens to prune from each video. Value sits in range
[0;1); pruning is enabled when it is greater than 0. The pruning algorithm
is selected by video_pruning_method.
compute_hash()
¶
WARNING: Whenever a new field is added to this config, ensure that it is included in the factors list if it affects the computation graph.
Provide a hash that uniquely identifies all the configs that affect the structure of the computation graph from input ids/embeddings to the final hidden states, excluding anything before input ids/embeddings and after the final hidden states.
Source code in vllm/config/multimodal.py
fold_mm_processor_device(mm_processor_kwargs, mm_processor_device)
staticmethod
¶
Fold the mm_processor_device convenience flag into the kwargs.
The flag keeps no state of its own: mm_processor_kwargs["device"] is
the only representation of where the processor runs, so an explicit
device there always wins and "auto" stays unresolved for
VllmConfig, which is where the EC role needed to resolve it lives.
Parameters:
-
(mm_processor_kwargs¶dict[str, Any] | None) –The kwargs as given, or None.
-
(mm_processor_device¶MMProcessorDevice | None) –The flag's value, or None when unset.
Returns:
-
dict[str, Any] | None–The kwargs to build the config with, unchanged unless the flag adds
-
dict[str, Any] | None–a
device.
Source code in vllm/config/multimodal.py
get_limit_per_prompt(modality)
¶
Get the maximum number of input items allowed per prompt for the given modality (backward compatible).
Source code in vllm/config/multimodal.py
get_mm_processor_device_type()
¶
The torch device type mm_processor_kwargs["device"] names.
mm_processor_kwargs is untyped, so device may be any form torch
accepts -- "cuda", "cuda:1", torch.device(...), or a bare index.
Normalising through torch rather than parsing the string keeps the
non-string forms from slipping past a caller's comparison.
Returns:
-
str | None–The device type, or None when no device is requested.
Raises:
-
ValueError–If
deviceis not somethingtorch.deviceaccepts.validate_mm_processor_deviceis what surfaces this during startup, so the value is only parsed once.
Source code in vllm/config/multimodal.py
get_video_pruning_spec()
¶
Return (method, rate) when video pruning is enabled, else None.
rate is the fraction of video tokens to prune.
Source code in vllm/config/multimodal.py
merge_mm_processor_kwargs(inference_kwargs)
¶
Get the keyword arguments to pass to the multi-modal processor according to the extra arguments passed during inference.
Nested mappings are merged recursively, with inference-time values taking precedence over configured values.
Source code in vllm/config/multimodal.py
use_gpu_video_backend()
¶
Return whether the configured video loader or codec uses the GPU.
Source code in vllm/config/multimodal.py
validate_mm_processor_device(ec_config)
¶
Check mm_processor_kwargs["device"] for this deployment.
The only place the requested device is validated, so it runs even on a CPU-only platform: the value is parsed before any early return.
Parameters:
-
(ec_config¶ECTransferConfig | None) –The deployment's EC config, or None when it is not an encode/prefill/decode deployment. Passed in because it is not reachable from here, and because a field assigned after construction would not re-trigger this config's validators.
Raises:
-
ValueError–If the requested device is not a torch device, or if it is the accelerator on an instance that also runs the language model.
Source code in vllm/config/multimodal.py
MultiModalDummyOptions
¶
Bases: dict[str, BaseDummyOptions]
Dummy data options for each modality.
Lookups of the modalities predefined by vLLM return their own options
class, while any other modality returns
BaseDummyOptions.
Source code in vllm/config/multimodal.py
VideoDummyOptions
¶
Bases: BaseDummyOptions
Options for generating dummy video data during profiling.
Source code in vllm/config/multimodal.py
_merge_mm_processor_kwargs(config_kwargs, inference_kwargs)
¶
Merge configured and inference-time multi-modal processor kwargs.
The merge is performed in two parts:
- The shared flat kwargs are merged recursively, with inference-time values taking precedence.
- Each existing HuggingFace processor kwarg scope is merged separately. For keys present in either the configured or inference-time scope, precedence is: configured flat < configured scoped < inference flat < inference scoped.
For example:
configured = {
"size": {"shortest_edge": 256},
"videos_kwargs": {
"size": {"longest_edge": 1024},
"num_frames": 16,
},
}
inference = {
"size": {"shortest_edge": 512},
"videos_kwargs": {"fps": 4},
}
produces:
{
"size": {"shortest_edge": 512},
"videos_kwargs": {
"size": {
"shortest_edge": 512,
"longest_edge": 1024,
},
"num_frames": 16,
"fps": 4,
},
}
Scoped values do not modify the shared flat kwargs or sibling scopes. Missing scopes, None-valued scopes, and empty scope mappings are treated as absent and do not create a scope.
Source code in vllm/config/multimodal.py
97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 | |
_merge_mm_processor_scope(config_flat, config_scoped, inference_flat, inference_scoped)
¶
Merge kwargs for one existing HuggingFace processor kwarg scope.
Flat kwargs contribute only to keys present in either the configured or inference-time scope. For those keys, precedence is: configured flat < configured scoped < inference flat < inference scoped.
Nested mappings are merged recursively; otherwise the higher-priority value replaces the lower-priority value.
Source code in vllm/config/multimodal.py
_recursively_merge_mm_processor_kwargs(defaults, overrides)
¶
Merge two mappings recursively, with overrides taking precedence.
Nested values are merged only when both values are mappings; otherwise the override replaces the existing value.