vllm.lora.request
¶
Classes:
-
LoRARequest–Request for a LoRA adapter.
LoRARequest
¶
Bases: Struct
Request for a LoRA adapter.
lora_int_id must be globally unique for a given adapter. This is currently not enforced in vLLM.
If True, forces reloading the adapter even if one
with the same lora_int_id already exists in the cache. This replaces the existing adapter in-place. If False (default), only loads if the adapter is not already loaded.
Methods:
-
__eq__–Overrides the equality method to compare LoRARequest
-
__hash__–Overrides the hash method to hash LoRARequest instances
Attributes:
-
is_3d_lora_weight(bool) –Whether this adapter's MoE weights are stored in the 3D fused
Source code in vllm/lora/request.py
is_3d_lora_weight = False
class-attribute
instance-attribute
¶
Whether this adapter's MoE weights are stored in the 3D fused
gate_up_proj / down_proj layout (one fused tensor per layer) or the
2D per-expert split layout (separate gate_proj / up_proj / down_proj
tensors per expert). Only consulted when the engine is started with
enable_mixed_moe_lora_format=True; otherwise it is ignored and the
on-disk format is inferred from the base model.
__eq__(value)
¶
Overrides the equality method to compare LoRARequest instances based on lora_name. This allows for identification and comparison lora adapter across engines.
Source code in vllm/lora/request.py
__hash__()
¶
Overrides the hash method to hash LoRARequest instances based on lora_name. This ensures that LoRARequest instances can be used in hash-based collections such as sets and dictionaries, identified by their names across engines.