vllm.benchmarks.startup
¶
Benchmark the cold and warm startup time of vLLM models.
This script measures total startup time (including model loading, compilation, and cache operations) for both cold and warm scenarios: - Cold startup: Fresh start with no caches (temporary cache directories) - Warm startup: Using cached compilation and model info
Classes:
-
MetricDesc–Descriptor for a metric to collect from each iteration.
-
MetricStats–Aggregated statistics for a single benchmark metric.
Functions:
-
cold_startup–Context manager to measure cold startup time:
-
run_startup_in_subprocess–Run LLM startup in a subprocess and return timing metrics via a queue.
MetricDesc
¶
Bases: NamedTuple
Descriptor for a metric to collect from each iteration.
Source code in vllm/benchmarks/startup.py
MetricStats
¶
Bases: NamedTuple
Aggregated statistics for a single benchmark metric.
Source code in vllm/benchmarks/startup.py
cold_startup()
¶
Context manager to measure cold startup time: 1. Uses a temporary directory for vLLM cache to avoid any pollution between cold startup iterations. 2. Uses inductor's fresh_cache to clear torch.compile caches.
Source code in vllm/benchmarks/startup.py
run_startup_in_subprocess(engine_args, result_queue)
¶
Run LLM startup in a subprocess and return timing metrics via a queue. This ensures complete isolation between iterations.