vllm.config.lora
¶
Classes:
-
LoRAConfig–Configuration for LoRA.
LoRAConfig
¶
Configuration for LoRA.
Methods:
-
compute_hash–WARNING: Whenever a new field is added to this config,
Attributes:
-
default_mm_loras(dict[str, str] | None) –Dictionary mapping specific modalities to LoRA model paths; this field
-
enable_mixed_moe_lora_format(bool) –If True, force the engine to use the universal 2D MoE LoRA wrapper
-
enable_moe_shared_loras(bool) –If True, load MoE expert adapters in the "shared-outer" layout, where the
-
enable_tower_connector_lora(bool) –If
True, LoRA support for the tower (vision encoder) and connector -
fully_sharded_loras(bool) –By default, only half of the LoRA computation is sharded with tensor
-
lora_dtype(dtype | LoRADType) –Data type for LoRA. If auto, will default to base model dtype.
-
max_cpu_loras(int | None) –Maximum number of LoRAs to store in CPU memory. Must be >= than
-
max_lora_cls_labels(int | None) –Maximum output size for LoRA classification heads. Defaults to the
-
max_lora_rank(MaxLoRARanks) –Max LoRA rank.
-
max_loras(int) –Max number of LoRAs in a single batch.
-
specialize_active_lora(bool) –Whether to construct lora kernel grid by the number of active LoRA adapters.
-
target_modules(list[str] | None) –Restrict LoRA to specific module suffixes (e.g., ["o_proj", "qkv_proj"]).
Source code in vllm/config/lora.py
31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 | |
default_mm_loras = None
class-attribute
instance-attribute
¶
Dictionary mapping specific modalities to LoRA model paths; this field is only applicable to multimodal models and should be leveraged when a model always expects a LoRA to be active when a given modality is present. Note that currently, if a request provides multiple additional modalities, each of which have their own LoRA, we do NOT apply default_mm_loras because we currently only support one lora adapter per prompt. When run in offline mode, the lora IDs for n modalities will be automatically assigned to 1-n with the names of the modalities in alphabetic order.
enable_mixed_moe_lora_format = False
class-attribute
instance-attribute
¶
If True, force the engine to use the universal 2D MoE LoRA wrapper
(FusedMoEWithLoRA) regardless of the model's is_3d_moe_weight flag, so
that 2D-format and 3D-format MoE LoRA adapters can be served in the same
deployment. Only meaningful for MoE models; ignored otherwise. Default False
keeps the existing model-driven behavior.
enable_moe_shared_loras = False
class-attribute
instance-attribute
¶
If True, load MoE expert adapters in the "shared-outer" layout, where the
gate/up (w1/w3) lora_A and the down (w2) lora_B are shared across all
experts (stored once with expert-dim 1) instead of per-expert. The shared
factors are broadcast to the expert count at kernel time. Only meaningful for
MoE models whose adapters use this layout; ignored otherwise.
enable_tower_connector_lora = False
class-attribute
instance-attribute
¶
If True, LoRA support for the tower (vision encoder) and connector
of multimodal models will be enabled. This is an experimental feature and
currently only supports some MM models such as the Qwen VL series. The default
is False.
fully_sharded_loras = False
class-attribute
instance-attribute
¶
By default, only half of the LoRA computation is sharded with tensor parallelism. Enabling this will use the fully sharded layers. At high sequence length, max rank or tensor parallel size, this is likely faster.
lora_dtype = 'auto'
class-attribute
instance-attribute
¶
Data type for LoRA. If auto, will default to base model dtype.
max_cpu_loras = None
class-attribute
instance-attribute
¶
Maximum number of LoRAs to store in CPU memory. Must be >= than
max_loras.
max_lora_cls_labels = Field(default=None, ge=1)
class-attribute
instance-attribute
¶
Maximum output size for LoRA classification heads. Defaults to the base classification head size.
max_lora_rank = 16
class-attribute
instance-attribute
¶
Max LoRA rank.
max_loras = Field(default=1, ge=1)
class-attribute
instance-attribute
¶
Max number of LoRAs in a single batch.
specialize_active_lora = False
class-attribute
instance-attribute
¶
Whether to construct lora kernel grid by the number of active LoRA adapters. When set to True, separate cuda graphs will be captured for different counts of active LoRAs (powers of 2 up to max_loras), which can improve performance for variable LoRA usage patterns at the cost of increased startup time and memory usage. Only takes effect when cudagraph_specialize_lora is True.
target_modules = None
class-attribute
instance-attribute
¶
Restrict LoRA to specific module suffixes (e.g., ["o_proj", "qkv_proj"]). If None, all supported LoRA modules are used. This allows deployment-time control over which modules have LoRA applied, useful for performance tuning.
compute_hash()
¶
WARNING: Whenever a new field is added to this config, ensure that it is included in the factors list if it affects the computation graph.
Provide a hash that uniquely identifies all the configs that affect the structure of the computation graph from input ids/embeddings to the final hidden states, excluding anything before input ids/embeddings and after the final hidden states.