vllm.model_executor.model_loader.base_loader
¶
Classes:
-
BaseModelLoader–Base class for model loaders.
Functions:
-
log_model_inspection–Log model structure if VLLM_LOG_MODEL_INSPECTION=1.
-
log_online_quantization–Log the online-quantized layer count and types, when applicable.
BaseModelLoader
¶
Bases: ABC
Base class for model loaders.
Methods:
-
create_model–Create a model with the given configurations.
-
download_model–Download a model so that it can be immediately loaded.
-
get_external_weight_memory–Get weights memory from external process;
-
load_model–Load a model with the given configurations.
-
load_weights–Load weights into a model. This standalone API allows
Source code in vllm/model_executor/model_loader/base_loader.py
create_model(vllm_config, model_config, prefix='')
¶
Create a model with the given configurations.
Source code in vllm/model_executor/model_loader/base_loader.py
download_model(model_config)
abstractmethod
¶
get_external_weight_memory(vllm_config)
¶
Get weights memory from external process; 0 when the weights are not external.
load_model(vllm_config, model_config, prefix='')
¶
Load a model with the given configurations.
Source code in vllm/model_executor/model_loader/base_loader.py
load_weights(model, model_config)
abstractmethod
¶
Load weights into a model. This standalone API allows inplace weights loading for an already-initialized model
Source code in vllm/model_executor/model_loader/base_loader.py
log_model_inspection(model)
¶
Log model structure if VLLM_LOG_MODEL_INSPECTION=1.
Source code in vllm/model_executor/model_loader/base_loader.py
log_online_quantization(vllm_config)
¶
Log the online-quantized layer count and types, when applicable.