vllm.config.observability
¶
Classes:
-
ObservabilityConfig–Configuration for observability - metrics and tracing.
ObservabilityConfig
¶
Configuration for observability - metrics and tracing.
Methods:
-
compute_hash–WARNING: Whenever a new field is added to this config,
Attributes:
-
collect_detailed_traces(list[DetailedTraceModules] | None) –It makes sense to set this only if
--otlp-traces-endpointis set. If -
collect_model_execute_time(bool) –Whether to collect model execute time for the request.
-
collect_model_forward_time(bool) –Whether to collect model forward time for the request.
-
cudagraph_metrics(bool) –Enable CUDA graph metrics (number of padded/unpadded tokens, runtime cudagraph
-
custom_histogram_buckets(dict[str, list[float]] | None) –Custom Prometheus histogram bucket boundaries, as a JSON mapping from
-
enable_layerwise_nvtx_tracing(bool) –Enable layerwise NVTX tracing. This traces the execution of each layer or
-
enable_logging_iteration_details(bool) –Enable detailed logging of iteration details.
-
enable_mfu_metrics(bool) –Enable Model FLOPs Utilization (MFU) metrics.
-
enable_mm_processor_stats(bool) –Enable collection of timing statistics for multimodal processor operations.
-
jit_monitor_mode(Literal['warn', 'error']) –How to handle post-warmup JIT compilation events.
-
jit_monitor_verbose(bool) –Log every monitored JIT compile with runtime details. This can emit many
-
kv_cache_metrics(bool) –Enable KV cache residency metrics (lifetime, idle time, reuse gaps).
-
kv_cache_metrics_sample(float) –Sampling rate for KV cache metrics (0.0, 1.0]. Default 0.01 = 1% of blocks.
-
otlp_traces_endpoint(str | None) –Target URL to which OpenTelemetry traces will be sent.
-
per_request_spec_decode_metrics(Literal['none', 'summary', 'detailed']) –Include per-request speculative-decoding acceptance metrics in the
-
show_hidden_metrics(bool) –Check if the hidden metrics should be shown.
-
show_hidden_metrics_for_version(str | None) –Enable deprecated Prometheus metrics that have been hidden since the
Source code in vllm/config/observability.py
19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 | |
collect_detailed_traces = None
class-attribute
instance-attribute
¶
It makes sense to set this only if --otlp-traces-endpoint is set. If
set, it will collect detailed traces for the specified modules. This
involves use of possibly costly and or blocking operations and hence might
have a performance impact.
Note that collecting detailed timing information for each request can be expensive.
collect_model_execute_time
cached
property
¶
Whether to collect model execute time for the request.
collect_model_forward_time
cached
property
¶
Whether to collect model forward time for the request.
cudagraph_metrics = False
class-attribute
instance-attribute
¶
Enable CUDA graph metrics (number of padded/unpadded tokens, runtime cudagraph dispatch modes, and their observed frequencies at every logging interval).
custom_histogram_buckets = None
class-attribute
instance-attribute
¶
Custom Prometheus histogram bucket boundaries, as a JSON mapping from
bucket-family key to a list of strictly increasing, positive, finite
upper bounds. When a family key is present, its list replaces the default
buckets for every histogram in that family; families not listed keep
their defaults. Known families: request_latency, time_to_first_token,
inter_token_latency, iteration_tokens, request_params_n,
request_num_preemptions, request_tokens, kv_cache_residency. Example:
--custom-histogram-buckets '{"request_latency": [0.01, 0.05, 0.1, 0.5]}'.
Note that every extra bucket adds one time series per metric and label
set.
enable_layerwise_nvtx_tracing = False
class-attribute
instance-attribute
¶
Enable layerwise NVTX tracing. This traces the execution of each layer or module in the model and attach information such as input/output shapes to nvtx range markers. Noted that this doesn't work with CUDA graphs enabled.
enable_logging_iteration_details = False
class-attribute
instance-attribute
¶
Enable detailed logging of iteration details. If set, vllm EngineCore will log iteration details This includes number of context/generation requests and tokens and the elapsed cpu time for the iteration.
enable_mfu_metrics = False
class-attribute
instance-attribute
¶
Enable Model FLOPs Utilization (MFU) metrics.
enable_mm_processor_stats = False
class-attribute
instance-attribute
¶
Enable collection of timing statistics for multimodal processor operations. This is for internal use only (e.g., benchmarks) and is not exposed as a CLI argument.
jit_monitor_mode = 'warn'
class-attribute
instance-attribute
¶
How to handle post-warmup JIT compilation events.
jit_monitor_verbose = False
class-attribute
instance-attribute
¶
Log every monitored JIT compile with runtime details. This can emit many logs and add overhead, so it is intended for debugging.
kv_cache_metrics = False
class-attribute
instance-attribute
¶
Enable KV cache residency metrics (lifetime, idle time, reuse gaps). Uses sampling to minimize overhead. Requires log stats to be enabled (i.e., --disable-log-stats not set).
kv_cache_metrics_sample = Field(default=0.01, gt=0, le=1)
class-attribute
instance-attribute
¶
Sampling rate for KV cache metrics (0.0, 1.0]. Default 0.01 = 1% of blocks.
otlp_traces_endpoint = None
class-attribute
instance-attribute
¶
Target URL to which OpenTelemetry traces will be sent.
per_request_spec_decode_metrics = 'none'
class-attribute
instance-attribute
¶
Include per-request speculative-decoding acceptance metrics in the
response under metrics.speculative_decoding. none disables; summary adds mean
acceptance length, draft acceptance rate, and the step-by-draft-length
histogram; detailed additionally records the ordered per-step
accepted/proposed arrays (one entry per verify step). Only reported for
single-sequence requests (n == 1), mirroring the timing metrics. No effect
unless speculative decoding is enabled. Independent of --disable-log-stats.
This is the per-request response-body counterpart of the aggregate
vllm:spec_decode_* Prometheus metrics. The response field is experimental
and its shape may change in a future release.
show_hidden_metrics
cached
property
¶
Check if the hidden metrics should be shown.
show_hidden_metrics_for_version = None
class-attribute
instance-attribute
¶
Enable deprecated Prometheus metrics that have been hidden since the
specified version. For example, if a previously deprecated metric has been
hidden since the v0.7.0 release, you use
--show-hidden-metrics-for-version=0.7 as a temporary escape hatch while
you migrate to new metrics. The metric is likely to be removed completely
in an upcoming release.
_reject_bool_histogram_bounds(value)
classmethod
¶
Reject booleans before pydantic silently coerces them to floats.
Source code in vllm/config/observability.py
compute_hash()
¶
WARNING: Whenever a new field is added to this config, ensure that it is included in the factors list if it affects the computation graph.
Provide a hash that uniquely identifies all the configs that affect the structure of the computation graph from input ids/embeddings to the final hidden states, excluding anything before input ids/embeddings and after the final hidden states.