vllm.v1.kv_offload.cpu.policies.base
¶
Classes:
-
CachePolicy–Encapsulates both chunk organization (data structures) and replacement
-
ChunkStatus–Offloading status for a single chunk of KV data.
Functions:
-
order_request_keys–Return one head-to-tail order across all KV cache groups.
CachePolicy
¶
Bases: ABC
Encapsulates both chunk organization (data structures) and replacement decisions (which chunk to evict). LRU and ARC differ in both dimensions — ARC's ghost lists and target_t1_size live at the intersection of storage and eviction, so they cannot be separated cleanly.
Methods:
-
clear–Remove ALL chunks regardless of ref_cnt.
-
evict–Evict exactly n chunks, skipping any in protected.
-
get–Find chunk in data structures. Returns None if not present.
-
insert–Add a newly allocated chunk. For ARC: also removes from ghost lists.
-
mark_evictable–Called when a chunk's ref_cnt transitions to 0.
-
mark_non_evictable–Called when a chunk's ref_cnt transitions from 0.
-
on_request_finished–Apply one request-scoped cache access in prefix order.
-
on_store_miss–Observe store misses before their cache entries are inserted.
-
remove–Remove a chunk (used to clean up after a failed store).
-
touch–Mark chunks as recently used.
Source code in vllm/v1/kv_offload/cpu/policies/base.py
48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 | |
clear()
abstractmethod
¶
Remove ALL chunks regardless of ref_cnt.
Ghost lists and adaptive state are also reset.
evict(n, protected)
abstractmethod
¶
Evict exactly n chunks, skipping any in protected.
Returns a list of (key, chunk) for the evicted chunks, or None if n evictions cannot be satisfied. The operation is atomic: if None is returned, no state changes are made.
For ARC: ghost list cleanup (trimming to cache_capacity) is performed at the end of a successful eviction.
Source code in vllm/v1/kv_offload/cpu/policies/base.py
get(key)
abstractmethod
¶
insert(key, chunk)
abstractmethod
¶
mark_evictable(key)
¶
mark_non_evictable(key)
¶
on_request_finished(key_groups, insertion_only_keys, reused_keys, req_context)
¶
Apply one request-scoped cache access in prefix order.
key_groups contains keys observed for each KV cache group in
head-to-tail order. insertion_only_keys and reused_keys
distinguish chunks only created by this request from ready chunks it
actually reused, which matters for policies such as ARC where reuse
changes frequency but insertion and pending observations do not.
The default forwards one touch per group, preserving compatibility for experimental out-of-tree policies while moving those touches to request finalization. Policies that distinguish insertion from reuse can override this hook and inspect the access classifications.
Source code in vllm/v1/kv_offload/cpu/policies/base.py
on_store_miss(keys, req_context)
¶
Observe store misses before their cache entries are inserted.
The default delegates to touch for compatibility with external
policies. Policies may override this hook when a miss has distinct
semantics, such as ARC adapting to a ghost-list hit.
Source code in vllm/v1/kv_offload/cpu/policies/base.py
remove(key)
abstractmethod
¶
ChunkStatus
¶
Bases: Structure
Offloading status for a single chunk of KV data. Holds the following information:
ref_cnt - the current number of transfers using this chunk as a source. A value of -1 indicates the chunk is not yet ready to be read. chunk_id - index of the physical CPU buffer slot.
Attributes:
Source code in vllm/v1/kv_offload/cpu/policies/base.py
is_ready
property
¶
Returns whether the chunk is ready to be read.
order_request_keys(key_groups, req_context)
¶
Return one head-to-tail order across all KV cache groups.