Production Metrics¶
vLLM exposes a number of metrics that can be used to monitor the health of the
system. These metrics are exposed via the /metrics endpoint on the vLLM
OpenAI compatible API server.
You can start the server using Python, or using Docker:
Then query the endpoint to get the latest metrics from the server:
Output
$ curl http://0.0.0.0:8000/metrics
# HELP vllm:iteration_tokens_total Histogram of number of tokens per engine_step.
# TYPE vllm:iteration_tokens_total histogram
vllm:iteration_tokens_total_sum{model_name="unsloth/Llama-3.2-1B-Instruct"} 0.0
vllm:iteration_tokens_total_bucket{le="1.0",model_name="unsloth/Llama-3.2-1B-Instruct"} 3.0
vllm:iteration_tokens_total_bucket{le="8.0",model_name="unsloth/Llama-3.2-1B-Instruct"} 3.0
vllm:iteration_tokens_total_bucket{le="16.0",model_name="unsloth/Llama-3.2-1B-Instruct"} 3.0
vllm:iteration_tokens_total_bucket{le="32.0",model_name="unsloth/Llama-3.2-1B-Instruct"} 3.0
vllm:iteration_tokens_total_bucket{le="64.0",model_name="unsloth/Llama-3.2-1B-Instruct"} 3.0
vllm:iteration_tokens_total_bucket{le="128.0",model_name="unsloth/Llama-3.2-1B-Instruct"} 3.0
vllm:iteration_tokens_total_bucket{le="256.0",model_name="unsloth/Llama-3.2-1B-Instruct"} 3.0
vllm:iteration_tokens_total_bucket{le="512.0",model_name="unsloth/Llama-3.2-1B-Instruct"} 3.0
...
The following metrics are exposed:
General Metrics¶
| Metric Name | Type | Description |
|---|---|---|
vllm:corrupted_requests |
Counter | Corrupted requests, in terms of total number of requests with NaNs in logits. |
vllm:external_prefix_cache_hits |
Counter | External prefix cache hits from KV connector cross-instance cache sharing, in terms of number of cached tokens. |
vllm:external_prefix_cache_queries |
Counter | External prefix cache queries from KV connector cross-instance cache sharing, in terms of number of queried tokens. |
vllm:generation_tokens |
Counter | Number of generation tokens processed. |
vllm:mm_cache_hits |
Counter | Multi-modal cache hits, in terms of number of cached items. |
vllm:mm_cache_queries |
Counter | Multi-modal cache queries, in terms of number of queried items. |
vllm:num_preemptions |
Counter | Cumulative number of preemption from the engine. |
vllm:prefix_cache_hits |
Counter | Prefix cache hits, in terms of number of cached tokens. |
vllm:prefix_cache_queries |
Counter | Prefix cache queries, in terms of number of queried tokens. |
vllm:prompt_tokens |
Counter | Number of prefill tokens processed. |
vllm:prompt_tokens_by_source |
Counter | Number of prompt tokens by source. |
vllm:prompt_tokens_cached |
Counter | Number of cached prompt tokens (local + external). |
vllm:request_success |
Counter | Count of successfully processed requests. |
vllm:engine_sleep_state |
Gauge | Engine sleep state; awake = 0 means engine is sleeping; awake = 1 means engine is awake; weights_offloaded = 1 means sleep level 1; discard_all = 1 means sleep level 2. |
vllm:kv_cache_usage_perc |
Gauge | KV-cache usage. 1 means 100 percent usage. |
vllm:lora_requests_info |
Gauge | Running stats on lora requests. |
vllm:num_requests_kv_fetch_by_stage |
Gauge | Number of waiting requests by async KV load stage. Stage labels: 'waiting_to_start' = needs an async KV load that has not started; 'in_progress' = load started and not yet reported finished; 'completed_waiting' = load finished, request not running yet. |
vllm:num_requests_running |
Gauge | Number of requests in model execution batches. |
vllm:num_requests_waiting |
Gauge | Number of requests waiting to be processed. |
vllm:num_requests_waiting_by_reason |
Gauge | Number of waiting requests by reason. Reason labels: 'capacity' = waiting for scheduling capacity; 'deferred' = deferred by transient constraints (LoRA budget, KV transfer, blocked status). Sum of all reasons equals vllm:num_requests_waiting. |
vllm:e2e_request_latency_seconds |
Histogram | Histogram of e2e request latency in seconds. |
vllm:inter_token_latency_seconds |
Histogram | Histogram of inter-token latency in seconds. |
vllm:iteration_tokens_total |
Histogram | Histogram of number of tokens per engine_step. |
vllm:kv_block_idle_before_evict_seconds |
Histogram | Histogram of idle time before KV cache block eviction. Sampled metrics (controlled by --kv-cache-metrics-sample). |
vllm:kv_block_lifetime_seconds |
Histogram | Histogram of KV cache block lifetime from allocation to eviction. Sampled metrics (controlled by --kv-cache-metrics-sample). |
vllm:kv_block_reuse_gap_seconds |
Histogram | Histogram of time gaps between consecutive KV cache block accesses. Only the most recent accesses are recorded (ring buffer). Sampled metrics (controlled by --kv-cache-metrics-sample). |
vllm:request_decode_time_seconds |
Histogram | Histogram of time spent in DECODE phase for request. |
vllm:request_generation_tokens |
Histogram | Number of generation tokens processed. |
vllm:request_inference_time_seconds |
Histogram | Histogram of time spent in RUNNING phase for request. |
vllm:request_max_num_generation_tokens |
Histogram | Histogram of maximum number of requested generation tokens. |
vllm:request_num_preemptions |
Histogram | Histogram of the number of times a request was preempted. |
vllm:request_params_max_tokens |
Histogram | Histogram of the max_tokens request parameter. |
vllm:request_params_n |
Histogram | Histogram of the n request parameter. |
vllm:request_prefill_kv_computed_tokens |
Histogram | Histogram of new KV tokens computed during prefill (excluding cached tokens). |
vllm:request_prefill_time_seconds |
Histogram | Histogram of time spent in PREFILL phase for request. |
vllm:request_prompt_tokens |
Histogram | Number of prefill tokens processed. |
vllm:request_queue_time_seconds |
Histogram | Histogram of time spent in WAITING phase for request. |
vllm:request_time_per_output_token_seconds |
Histogram | Histogram of time_per_output_token_seconds per request. |
vllm:time_to_first_token_seconds |
Histogram | Histogram of time to first token in seconds. |
Speculative Decoding Metrics¶
| Metric Name | Type | Description |
|---|---|---|
vllm:spec_decode_num_accepted_tokens_per_pos |
Counter | Accepted tokens per draft position. |
NIXL KV Connector Metrics¶
| Metric Name | Type | Description |
|---|---|---|
vllm:nixl_num_failed_notifications |
Counter | Number of failed NIXL KV Cache notifications. Retained for compatibility; these failures are also included in vllm:nixl_num_failed_transfers. |
vllm:nixl_num_failed_transfers |
Counter | Number of failed NIXL KV Cache transfers, including handshake and notification failures. NOTE: KV expiry is tracked separately in vllm:nixl_num_kv_expired_reqs. |
vllm:nixl_num_kv_expired_reqs |
Counter | Number of requests that had their KV expire. NOTE: This metric is tracked on the P instance. |
vllm:nixl_num_notifications_after_expiry |
Counter | Number of completion notifications for requests that were no longer tracked, usually because their KV lease expired. The KV blocks may have been reused before the transfer finished. Counted per notification (one per remote rank), not per request. |
vllm:nixl_bytes_transferred |
Histogram | Histogram of bytes transferred per NIXL KV Cache transfers. |
vllm:nixl_num_descriptors |
Histogram | Histogram of number of descriptors per NIXL KV Cache transfers. |
vllm:nixl_post_time_seconds |
Histogram | Histogram of transfer post time for NIXL KV Cache transfers. |
vllm:nixl_xfer_time_seconds |
Histogram | Histogram of transfer duration for NIXL KV Cache transfers. |
Simple CPU Offload Connector Metrics¶
These metrics are exposed when the SimpleCPUOffloadConnector KV connector
is configured (e.g. --kv-transfer-config='{"kv_connector":
"SimpleCPUOffloadConnector", "kv_role": "kv_both", "kv_connector_extra_config":
{"kv_offload_backend": "disk", "disk_path": "/mnt/nvme/kv"}}'). They are
updated once per engine step.
Caveats to keep in mind when interpreting them:
- A "completed" store means the write syscalls returned; it is not fsync-durable, and the disk backend's file is process-lifetime scratch, unlinked at startup and shutdown.
save_outcomes_totalclassifies eager-mode boundary hand-off stores; lazy-mode stores are not classified.used_blockscounts blocks pinned by in-flight transfers or cache hits; warm cached blocks that are evictable are not counted. Use thecapacity_blockslabel ofsimple_kv_offload_infoas the denominator.- Counters and gauges are quantized to engine steps. They are reported by the scheduler process and reflect engine-wide logical block counts, not values pooled from individual tensor-parallel workers.
| Metric Name | Type | Description |
|---|---|---|
vllm:simple_kv_offload_load_blocks_total |
Counter | KV blocks restored to GPU from the offload pool, counted when a load finishes. This is prefill work avoided through offload cache hits. |
vllm:simple_kv_offload_save_outcomes_total |
Counter | Store admission decisions in SimpleCPUOffloadConnector, by outcome. Outcomes classify eager-mode boundary hand-off stores; lazy-mode stores are not classified. |
vllm:simple_kv_offload_info |
Gauge | SimpleCPUOffloadConnector deployment facts. Value is always 1. backend is cpu or disk; page_cache and lazy_offload are true/false; capacity_blocks is the offload pool size. |
vllm:simple_kv_offload_pending_store_blocks |
Gauge | Offload blocks in pending stores: queued for dispatch, issued to workers, or awaiting release after a cache reset. A persistently growing value indicates a stuck transfer. |
vllm:simple_kv_offload_used_blocks |
Gauge | Offload-pool blocks currently pinned by in-flight transfers or cache hits (capacity minus free). Evictable cached blocks are not counted; capacity_blocks is on vllm:simple_kv_offload_info. |
Model Flops Utilization (MFU) Performance Metrics¶
These metrics are available via --enable-mfu-metrics:
| Metric Name | Type | Description |
|---|---|---|
vllm:estimated_flops_per_gpu_total |
Counter | Estimated number of floating point operations per GPU (for Model Flops Utilization calculations). |
vllm:estimated_read_bytes_per_gpu_total |
Counter | Estimated number of bytes read from memory per GPU (for Model Flops Utilization calculations). |
vllm:estimated_write_bytes_per_gpu_total |
Counter | Estimated number of bytes written to memory per GPU (for Model Flops Utilization calculations). |
Custom Histogram Buckets¶
The core engine histograms ship with default bucket boundaries tuned for
typical serving workloads. The --custom-histogram-buckets option replaces
the boundaries of one or more bucket families — exactly the histograms
listed in the table below — with your own list; histograms owned by other
subsystems (for example, the NIXL connector metrics) are not affected. Use it,
for example, to track sub-300ms latency SLOs with the request-phase
histograms, whose smallest default boundary is 0.3s:
vllm serve Qwen/Qwen3-0.6B \
--custom-histogram-buckets '{"request_latency": [0.01, 0.05, 0.1, 0.25, 0.5, 1.0, 5.0, 30.0]}'
Each family key overrides a group of related histograms:
| Family key | Histograms |
|---|---|
request_latency |
vllm:e2e_request_latency_seconds, vllm:request_queue_time_seconds, vllm:request_inference_time_seconds, vllm:request_prefill_time_seconds, vllm:request_decode_time_seconds |
time_to_first_token |
vllm:time_to_first_token_seconds |
inter_token_latency |
vllm:inter_token_latency_seconds, vllm:request_time_per_output_token_seconds |
iteration_tokens |
vllm:iteration_tokens_total |
request_params_n |
vllm:request_params_n |
request_num_preemptions |
vllm:request_num_preemptions |
request_tokens |
vllm:request_prompt_tokens, vllm:request_generation_tokens, vllm:request_max_num_generation_tokens, vllm:request_params_max_tokens, vllm:request_prefill_kv_computed_tokens |
kv_cache_residency |
vllm:kv_block_lifetime_seconds, vllm:kv_block_idle_before_evict_seconds, vllm:kv_block_reuse_gap_seconds |
Bucket values must be positive, finite, and strictly increasing; unknown
family keys are rejected at startup. Families you do not list keep their
default boundaries. The request_tokens defaults normally scale with
--max-model-len; an override replaces that computed list. The
kv_cache_residency family only takes effect when --kv-cache-metrics is
enabled.
Bucket cardinality
Every bucket boundary creates one extra time series per metric and per label combination (model and engine index, multiplied under data-parallel deployments). Long bucket lists inflate Prometheus storage, scrape sizes, and query costs. Keep custom lists short, and only override the families you actively monitor.
Deprecation Policy¶
Note: when metrics are deprecated in version X.Y, they are hidden in version X.Y+1
but can be re-enabled using the --show-hidden-metrics-for-version=X.Y escape hatch,
and are then removed in version X.Y+2.