vllm.distributed.ec_transfer.ec_connector.cpu.scheduler.embedding_cache
¶
EmbeddingCache — named-entry block cache with FIFO eviction.
Manages a fixed pool of block IDs keyed by content identity (mm_hash). Entries transition through: not-ready → ready (evictable) → pinned. Eviction targets ready + unpinned entries in FIFO order.
All public methods are thread-safe.
Classes:
-
CacheEntry–A single cache entry. Read
.block_idsand.readyfreely; -
EmbeddingCache–Fixed-size block cache with FIFO eviction.
CacheEntry
¶
A single cache entry. Read .block_ids and .ready freely;
mutations only through EmbeddingCache methods.
Source code in vllm/distributed/ec_transfer/ec_connector/cpu/scheduler/embedding_cache.py
EmbeddingCache
¶
Fixed-size block cache with FIFO eviction.
Entries are keyed by content identity (e.g. mm_hash). Blocks are allocated from a free-set; when space is needed, ready + unpinned entries are evicted oldest-first.
The caller is responsible for deciding when to call mark_ready
(e.g. after enough engine steps have elapsed for the worker to have
completed the write).
Methods:
-
alloc–Allocate n_blocks for key, evicting as needed.
-
discard–Remove a not-ready in-flight entry, returning its blocks to the pool.
-
get–Return the entry for key, or None if not present.
-
has_held_entries–True if any entry is not-ready or pinned (an in-flight transfer).
-
mark_ready–Mark an entry as ready (data is CPU-visible).
-
pin–Pin an entry (prevent eviction).
-
pin_if_ready–Atomically pin key if present and ready; return its entry.
-
unpin–Unpin an entry. Asserts currently pinned.
Source code in vllm/distributed/ec_transfer/ec_connector/cpu/scheduler/embedding_cache.py
41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 | |
_evict_until(n_blocks)
¶
Evict ready+unpinned entries FIFO until enough space. Lock held.
Source code in vllm/distributed/ec_transfer/ec_connector/cpu/scheduler/embedding_cache.py
alloc(key, n_blocks)
¶
Allocate n_blocks for key, evicting as needed.
The entry starts not-ready. Returns None if there is not enough space even after evicting all evictable entries.
Source code in vllm/distributed/ec_transfer/ec_connector/cpu/scheduler/embedding_cache.py
discard(key)
¶
Remove a not-ready in-flight entry, returning its blocks to the pool.
Used when an in-flight fill (e.g. a NIXL READ) fails before the entry is marked ready. Asserts the entry is present and not ready — ready or pinned entries are reclaimed through eviction/unpin, not discard.
Source code in vllm/distributed/ec_transfer/ec_connector/cpu/scheduler/embedding_cache.py
get(key)
¶
Return the entry for key, or None if not present.
has_held_entries()
¶
True if any entry is not-ready or pinned (an in-flight transfer).
Held entries are exactly those absent from the eviction free list, so they represent saves awaiting mark_ready or loads awaiting unpin.
Source code in vllm/distributed/ec_transfer/ec_connector/cpu/scheduler/embedding_cache.py
mark_ready(key)
¶
Mark an entry as ready (data is CPU-visible).
Source code in vllm/distributed/ec_transfer/ec_connector/cpu/scheduler/embedding_cache.py
pin(key)
¶
Pin an entry (prevent eviction).
Source code in vllm/distributed/ec_transfer/ec_connector/cpu/scheduler/embedding_cache.py
pin_if_ready(key)
¶
Atomically pin key if present and ready; return its entry.
Returns None when the key is absent. An entry returned with ready
False was not pinned: its save is still in flight, which the producer
reports differently from a miss because such a read can be retried.
Source code in vllm/distributed/ec_transfer/ec_connector/cpu/scheduler/embedding_cache.py
unpin(key)
¶
Unpin an entry. Asserts currently pinned.