vllm.multimodal.cache
¶
Modules:
-
base– -
factories– -
lru–Implementation of Key-Replicated Cache (see docs/configuration/optimization.md).
-
shm–Implementation of Shared Memory Cache (see docs/configuration/optimization.md).
Classes:
-
BaseMultiModalProcessorCache–The required interface for caches on P0.
-
BaseMultiModalReceiverCache–The required interface for caches on P1.
-
LruKeyReplicatedReceiverCache–The cache which is used on P1 when LRU caching is enabled.
-
LruKeyReplicatedSenderCache–The cache which is used on P0 when LRU caching is enabled.
-
MultiModalCacheMissError–Raised by the P1 receiver cache when items are requested with no data and
-
MultiModalProcessorOnlyCache–The cache which is used on P0 when IPC caching is disabled.
-
ShmObjectStoreReceiverCache–The cache which is used on P1 Worker Process when SHM caching is enabled.
-
ShmObjectStoreSenderCache–The cache which is used on P0 when SHM caching is enabled.
Functions:
-
engine_receiver_cache_from_config–Return a
BaseMultiModalReceiverCachefor the engine process. -
processor_cache_from_config–Return a
BaseMultiModalProcessorCache, if enabled. -
processor_only_cache_from_config–Return a
MultiModalProcessorOnlyCache, if enabled. -
worker_receiver_cache_from_config–Return a
BaseMultiModalReceiverCachefor the worker process.
BaseMultiModalProcessorCache
¶
Bases: BaseMultiModalCache[MultiModalProcessorCacheInItem, MultiModalProcessorCacheOutItem]
The required interface for caches on P0.
Methods:
-
close–Close the underlying cache, if needed.
-
invalidate–Drop
mm_hashfrom this P0 shadow cache to recover from P0/P1 drift. -
is_cached–Check whether a sequence of multi-modal items are
-
is_cached_item–Check whether a multi-modal item is
-
make_stats–Get (and reset) the multi-modal cache stats.
-
release_sender_touches–Release the items that were touched but not updated afterwards.
-
touch_sender_cache_item–Update the cache eviction order for a multi-modal item.
-
validate_input_item–Validate externally supplied cache metadata before engine handoff.
Source code in vllm/multimodal/cache/base.py
312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 | |
close()
¶
invalidate(mm_hash)
¶
Drop mm_hash from this P0 shadow cache to recover from P0/P1 drift.
No-op by default; shadow caches that can drift from P1 override this.
is_cached(mm_hashes)
¶
Check whether a sequence of multi-modal items are in the underlying cache.
This DOES NOT update the cache eviction order.
Parameters:
Returns:
Source code in vllm/multimodal/cache/base.py
is_cached_item(mm_hash)
abstractmethod
¶
Check whether a multi-modal item is in the underlying cache.
This DOES NOT update the cache eviction order.
Parameters:
Returns:
-
bool–Trueif the item is cached, otherwiseFalse.
Source code in vllm/multimodal/cache/base.py
make_stats(*, delta=False)
abstractmethod
¶
Get (and reset) the multi-modal cache stats.
Returns:
-
CacheInfo–The current multi-modal caching stats.
release_sender_touches()
¶
touch_sender_cache_item(mm_hash)
abstractmethod
¶
Update the cache eviction order for a multi-modal item.
This is used to touch the item in the cache without changing its value.
Parameters:
Source code in vllm/multimodal/cache/base.py
validate_input_item(mm_item, mm_hash)
¶
Validate externally supplied cache metadata before engine handoff.
BaseMultiModalReceiverCache
¶
Bases: BaseMultiModalCache[MultiModalKwargsItem | None, MultiModalKwargsItem]
The required interface for caches on P1.
Methods:
-
get_and_update_features–Update multimodal features with cached encoder outputs.
-
touch_receiver_cache_item–Update the cache eviction order for a multi-modal item.
Source code in vllm/multimodal/cache/base.py
get_and_update_features(mm_features)
¶
Update multimodal features with cached encoder outputs. Touch all identifier at first before update to avoid item in updated list evict during update.
Uses mm_hash for cache key to share across LoRAs (falls back to identifier for backward compatibility).
Source code in vllm/multimodal/cache/base.py
touch_receiver_cache_item(mm_hash, mm_item=None)
abstractmethod
¶
Update the cache eviction order for a multi-modal item.
This is used to touch the item in the cache without changing its value.
Parameters:
-
(mm_hash¶str) –The hash of the multi-modal item.
-
(mm_item¶MultiModalKwargsItem | None, default:None) –The multi-modal item itself. This is optional and may not be needed by some cache implementations.
Source code in vllm/multimodal/cache/base.py
LruKeyReplicatedReceiverCache
¶
Bases: BaseMultiModalReceiverCache
The cache which is used on P1 when LRU caching is enabled.
How to update each item:
- If the caller sent tensor data, store it (replacing any cached item under the same key) and return that data. P0 can miss after independent LRU eviction and resend a different item for the same identity.
- If the caller sent no data and the item is cached, return the cached item.
- If the caller sent no data and the item is not cached, raise
MultiModalCacheMissError.
Source code in vllm/multimodal/cache/lru.py
LruKeyReplicatedSenderCache
¶
Bases: BaseMultiModalProcessorCache
The cache which is used on P0 when LRU caching is enabled.
How to update each item:
-
If the item is already in the cache, clear the input to avoid unnecessary IPC.
-
If the item is not in the cache, store the metadata of that item so that the eviction policy remains the same as the cache on P1, and return the input. By only storing the metadata, we avoid keeping the data itself in memory inside P0.
Source code in vllm/multimodal/cache/lru.py
MultiModalCacheMissError
¶
Bases: RuntimeError
Raised by the P1 receiver cache when items are requested with no data and are not cached.
P0 (frontend) keeps a metadata-only shadow of P1 (engine) and sends
data=None on a shadow hit. The two caches are updated in different orders
across processes, so they can drift -- leaving P0 referencing items P1 has
evicted. Raising (instead of asserting) lets the engine return a retryable
response and have P0 drop the stale entries
(BaseMultiModalProcessorCache.invalidate) so the client resends the data.
Carries every drifted mm_hash in the request so P0 can drop them all in one
pass -- one retry then recovers the whole request, not one item per retry.
Source code in vllm/multimodal/cache/base.py
MultiModalProcessorOnlyCache
¶
Bases: BaseMultiModalProcessorCache
The cache which is used on P0 when IPC caching is disabled.
How to update each item:
- If the item is in the cache, replace the input with the cached item.
- If the item is not in the cache, store that item (which includes tensor data and metadata) into the cache, and return the input.
Source code in vllm/multimodal/cache/base.py
ShmObjectStoreReceiverCache
¶
Bases: BaseMultiModalReceiverCache
The cache which is used on P1 Worker Process when SHM caching is enabled.
How to update each item:
- If the item has an address, replace the input with the cached item.
- If not, return the input.
Methods:
-
touch_receiver_cache_item–Validate the item's handle in shared memory cache.
Source code in vllm/multimodal/cache/shm.py
248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 | |
touch_receiver_cache_item(mm_hash, mm_item=None)
¶
Validate the item's handle in shared memory cache.
Source code in vllm/multimodal/cache/shm.py
ShmObjectStoreSenderCache
¶
Bases: BaseMultiModalProcessorCache
The cache which is used on P0 when SHM caching is enabled.
How to update each item:
-
If the item is already in the cache, clear the input to avoid unnecessary IPC.
-
If the item is not in the cache, store the data in shared memory.
Methods:
-
remove_dangling_items–Remove items that are no longer in the shared memory cache.
-
touch_sender_cache_item–Touch the item in shared memory cache to prevent eviction.
Source code in vllm/multimodal/cache/shm.py
64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 | |
engine_receiver_cache_from_config(vllm_config)
¶
Return a BaseMultiModalReceiverCache for the engine process.
Source code in vllm/multimodal/cache/factories.py
processor_cache_from_config(vllm_config)
¶
Return a BaseMultiModalProcessorCache, if enabled.
Source code in vllm/multimodal/cache/factories.py
processor_only_cache_from_config(vllm_config)
¶
Return a MultiModalProcessorOnlyCache, if enabled.
Source code in vllm/multimodal/cache/factories.py
worker_receiver_cache_from_config(vllm_config, shared_worker_lock)
¶
Return a BaseMultiModalReceiverCache for the worker process.