vllm.utils.import_utils
¶
Contains helpers related to importing modules.
This is similar in concept to the importlib module.
Classes:
-
LazyLoader–LazyLoadermodule borrowed from [Tensorflow] -
PlaceholderModule–A placeholder object to use when a module does not exist.
Functions:
-
check_moonep_system_support–Raise if the current device cannot run MoonEP.
-
check_torchcodec_available–Whether the optional
torchcodecpackage is available. -
get_triton_kernels_version–The triton_kernels MoE-API generation ("3.5.1"/"3.6"/"3.8"), or None.
-
has_aiter–Whether the optional
aiterpackage is available. -
has_arctic_inference–Whether the optional
arctic_inferencepackage is available. -
has_cutedsl–Whether the optional
cutelasspackage is available. -
has_deep_ep–Whether the optional
deep_eppackage is available. -
has_deep_ep_v2–Whether deep_ep with ElasticBuffer (v2 API) is available.
-
has_deep_gemm–Whether the optional
deep_gemmpackage is available. -
has_fbgemm_gpu–Whether the optional
fbgemm_gpupackage is available. -
has_helion–Whether the optional
helionpackage is available. -
has_humming–Whether the optional
hummingpackage is available. -
has_moonep–Whether the optional
mooneppackage is available. -
has_mori–Whether the optional
moripackage is available. -
has_nixl_ep–Whether the optional
nixl_eppackage is available. -
has_quark–Whether the optional
quarkpackage is available. -
has_tilelang–Whether the optional
tilelangpackage is available. -
has_triton_kernels–Whether the optional
triton_kernelspackage is available. -
import_from_path–Import a Python file according to its file path.
-
import_plugin–Import a user-defined plugin.
-
import_pynvml–Historical comments:
-
import_triton_kernels–For convenience, prioritize triton_kernels that is available in
-
is_numba_available–Whether the optional
numbapackage is available. -
resolve_obj_by_qualname–Resolve an object by its fully-qualified class name.
LazyLoader
¶
Bases: ModuleType
LazyLoader module borrowed from [Tensorflow]
(https://github.com/tensorflow/tensorflow/blob/main/tensorflow/python/util/lazy_loader.py)
with an addition of "module caching".
Lazily import a module, mainly to avoid pulling in large dependencies.
Modules such as xgrammar might do additional side effects, so we
only want to use this when it is needed, delaying all eager effects.
Source code in vllm/utils/import_utils.py
PlaceholderModule
¶
Bases: _PlaceholderBase
A placeholder object to use when a module does not exist.
This enables more informative errors when trying to access attributes of a module that does not exist.
Source code in vllm/utils/import_utils.py
_PlaceholderBase
¶
Disallows downstream usage of placeholder modules.
We need to explicitly override each dunder method because
__getattr__
is not called when they are accessed.
Methods:
-
__getattr__–The main class should implement this to throw an error
Source code in vllm/utils/import_utils.py
143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 | |
__getattr__(key)
¶
The main class should implement this to throw an error for attribute accesses representing downstream usage.
_get_runtime_nccl_version()
¶
Get the runtime NCCL version by loading the actual library.
Returns the raw version int (e.g. 23004 for 2.30.4), or None on failure. torch.cuda.nccl.version() is a compile-time constant from the PyTorch wheel and does not reflect a separately installed NCCL.
Source code in vllm/utils/import_utils.py
_has_module(module_name)
cached
¶
Return True if module_name can be imported in the current environment.
Uses importlib.util.find_spec as a fast pre-check, then performs a
trial import to verify that native dependencies (shared libraries, etc.)
are also satisfied. Any failure during the trial import is treated as the
module being unavailable. The result is cached so that subsequent queries
for the same module incur no additional overhead.
Source code in vllm/utils/import_utils.py
_has_module_spec(module_name)
cached
¶
Return True if module_name is installed, without importing it.
Unlike _has_module, this only
resolves the import spec. It therefore does not pay the import cost of
heavyweight modules, at the price of not verifying that native
dependencies (shared libraries, etc.) are satisfied. The result is cached.
Source code in vllm/utils/import_utils.py
check_moonep_system_support()
¶
Raise if the current device cannot run MoonEP.
MoonEP's symmetric-memory buffers require CUDA VMM plus NVSwitch multicast
(SHARP) on every EP rank; NVLink-only topologies without NVSwitch (e.g.
4x H100 NV6) fail deep inside moonep.Buffer otherwise.
Source code in vllm/utils/import_utils.py
check_torchcodec_available()
¶
Whether the optional torchcodec package is available.
Source code in vllm/utils/import_utils.py
get_triton_kernels_version()
cached
¶
The triton_kernels MoE-API generation ("3.5.1"/"3.6"/"3.8"), or None.
Inferred by capability since the package exposes no usable version: 3.8
replaced matmul_ogs with matmul, and 3.5.1 predates SparseMatrix.
Source code in vllm/utils/import_utils.py
has_aiter()
¶
has_arctic_inference()
¶
has_cutedsl()
¶
has_deep_ep()
¶
has_deep_ep_v2()
¶
Whether deep_ep with ElasticBuffer (v2 API) is available.
Requires both the ElasticBuffer class in the deep_ep module and NCCL >= 2.30.4 (GIN backend), checked against the runtime library.
Source code in vllm/utils/import_utils.py
has_deep_gemm()
¶
Whether the optional deep_gemm package is available.
Prefers an externally installed deep_gemm package (so users can
override with a newer version), then falls back to the vendored copy
bundled in the vLLM wheel.
Source code in vllm/utils/import_utils.py
has_fbgemm_gpu()
¶
has_helion()
¶
Whether the optional helion package is available.
Helion is a Python-embedded DSL for writing ML kernels. See: https://github.com/pytorch/helion
Usage
if has_helion(): import helion import helion.language as hl # use helion...
Source code in vllm/utils/import_utils.py
has_humming()
¶
has_moonep()
¶
has_mori()
¶
has_nixl_ep()
¶
has_quark()
¶
has_tilelang()
cached
¶
Whether the optional tilelang package is available.
Only the import spec is checked: importing tilelang is expensive, so
callers must import it lazily at their point of use rather than relying
on this function to have imported it already.
Source code in vllm/utils/import_utils.py
has_triton_kernels()
¶
Whether the optional triton_kernels package is available.
Source code in vllm/utils/import_utils.py
import_from_path(module_name, file_path)
¶
Import a Python file according to its file path.
Based on the official recipe: https://docs.python.org/3/library/importlib.html#importing-a-source-file-directly
Source code in vllm/utils/import_utils.py
import_plugin(plugin_path)
¶
Import a user-defined plugin.
Plugin can be either: * a module in site-packages * a Python file specified by its path
Source code in vllm/utils/import_utils.py
import_pynvml()
¶
Historical comments:
libnvml.so is the library behind nvidia-smi, and
pynvml is a Python wrapper around it. We use it to get GPU
status without initializing CUDA context in the current process.
Historically, there are two packages that provide pynvml:
- nvidia-ml-py (https://pypi.org/project/nvidia-ml-py/): The official
wrapper. It is a dependency of vLLM, and is installed when users
install vLLM. It provides a Python module named pynvml.
- pynvml (https://pypi.org/project/pynvml/): An unofficial wrapper.
Prior to version 12.0, it also provides a Python module pynvml,
and therefore conflicts with the official one. What's worse,
the module is a Python package, and has higher priority than
the official one which is a standalone Python file.
This causes errors when both of them are installed.
Starting from version 12.0, it migrates to a new module
named pynvml_utils to avoid the conflict.
It is so confusing that many packages in the community use the
unofficial one by mistake, and we have to handle this case.
For example, nvcr.io/nvidia/pytorch:24.12-py3 uses the unofficial
one, and it will cause errors, see the issue
https://github.com/vllm-project/vllm/issues/12847 for example.
After all the troubles, we decide to copy the official pynvml
module to our codebase, and use it directly.
Source code in vllm/utils/import_utils.py
import_triton_kernels()
cached
¶
For convenience, prioritize triton_kernels that is available in
site-packages. Use vllm.third_party.triton_kernels as a fall-back.
Source code in vllm/utils/import_utils.py
is_numba_available()
¶
resolve_obj_by_qualname(qualname)
¶
Resolve an object by its fully-qualified class name.