Skip to content

vllm.v1.kv_offload.cpu.policies.base

Classes:

  • CachePolicy –

    Encapsulates both chunk organization (data structures) and replacement

  • ChunkStatus –

    Offloading status for a single chunk of KV data.

Functions:

CachePolicy

Bases: ABC

Encapsulates both chunk organization (data structures) and replacement decisions (which chunk to evict). LRU and ARC differ in both dimensions — ARC's ghost lists and target_t1_size live at the intersection of storage and eviction, so they cannot be separated cleanly.

Methods:

  • clear –

    Remove ALL chunks regardless of ref_cnt.

  • evict –

    Evict exactly n chunks, skipping any in protected.

  • get –

    Find chunk in data structures. Returns None if not present.

  • insert –

    Add a newly allocated chunk. For ARC: also removes from ghost lists.

  • mark_evictable –

    Called when a chunk's ref_cnt transitions to 0.

  • mark_non_evictable –

    Called when a chunk's ref_cnt transitions from 0.

  • on_request_finished –

    Apply one request-scoped cache access in prefix order.

  • on_store_miss –

    Observe store misses before their cache entries are inserted.

  • remove –

    Remove a chunk (used to clean up after a failed store).

  • touch –

    Mark chunks as recently used.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
class CachePolicy(ABC):
    """Encapsulates both chunk organization (data structures) and replacement
    decisions (which chunk to evict). LRU and ARC differ in both dimensions —
    ARC's ghost lists and target_t1_size live at the intersection of storage
    and eviction, so they cannot be separated cleanly.
    """

    def __init__(self, cache_capacity: int) -> None:
        self.cache_capacity = cache_capacity

    @abstractmethod
    def get(self, key: OffloadKey) -> ChunkStatus | None:
        """Find chunk in data structures. Returns None if not present."""

    @abstractmethod
    def insert(self, key: OffloadKey, chunk: ChunkStatus) -> None:
        """Add a newly allocated chunk. For ARC: also removes from ghost lists."""

    @abstractmethod
    def remove(self, key: OffloadKey) -> None:
        """Remove a chunk (used to clean up after a failed store)."""

    @abstractmethod
    def touch(self, keys: Iterable[OffloadKey], req_context: ReqContext) -> None:
        """Mark chunks as recently used.

        Args:
            keys: Chunks to mark as recently used.
            req_context: Per-request context for the request touching these chunks.

        """

    def on_store_miss(
        self, keys: Iterable[OffloadKey], req_context: ReqContext
    ) -> None:
        """Observe store misses before their cache entries are inserted.

        The default delegates to ``touch`` for compatibility with external
        policies. Policies may override this hook when a miss has distinct
        semantics, such as ARC adapting to a ghost-list hit.
        """
        self.touch(keys, req_context)

    def on_request_finished(
        self,
        key_groups: Sequence[Sequence[OffloadKey]],
        insertion_only_keys: set[OffloadKey],
        reused_keys: set[OffloadKey],
        req_context: ReqContext,
    ) -> None:
        """Apply one request-scoped cache access in prefix order.

        ``key_groups`` contains keys observed for each KV cache group in
        head-to-tail order. ``insertion_only_keys`` and ``reused_keys``
        distinguish chunks only created by this request from ready chunks it
        actually reused, which matters for policies such as ARC where reuse
        changes frequency but insertion and pending observations do not.

        The default forwards one touch per group, preserving compatibility
        for experimental out-of-tree policies while moving those touches to
        request finalization. Policies that distinguish insertion from reuse
        can override this hook and inspect the access classifications.
        """
        del insertion_only_keys, reused_keys
        for keys in key_groups:
            self.touch(keys, req_context)

    @abstractmethod
    def evict(
        self, n: int, protected: set[OffloadKey]
    ) -> list[tuple[OffloadKey, ChunkStatus]] | None:
        """Evict exactly n chunks, skipping any in protected.

        Returns a list of (key, chunk) for the evicted chunks,
        or None if n evictions cannot be satisfied. The operation is atomic:
        if None is returned, no state changes are made.

        For ARC: ghost list cleanup (trimming to cache_capacity) is performed
        at the end of a successful eviction.
        """

    @abstractmethod
    def clear(self) -> None:
        """Remove ALL chunks regardless of ref_cnt.

        Ghost lists and adaptive state are also reset.
        """

    def mark_evictable(self, key: OffloadKey) -> None:
        """Called when a chunk's ref_cnt transitions to 0."""
        return

    def mark_non_evictable(self, key: OffloadKey) -> None:
        """Called when a chunk's ref_cnt transitions from 0."""
        return

clear() abstractmethod

Remove ALL chunks regardless of ref_cnt.

Ghost lists and adaptive state are also reset.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
@abstractmethod
def clear(self) -> None:
    """Remove ALL chunks regardless of ref_cnt.

    Ghost lists and adaptive state are also reset.
    """

evict(n, protected) abstractmethod

Evict exactly n chunks, skipping any in protected.

Returns a list of (key, chunk) for the evicted chunks, or None if n evictions cannot be satisfied. The operation is atomic: if None is returned, no state changes are made.

For ARC: ghost list cleanup (trimming to cache_capacity) is performed at the end of a successful eviction.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
@abstractmethod
def evict(
    self, n: int, protected: set[OffloadKey]
) -> list[tuple[OffloadKey, ChunkStatus]] | None:
    """Evict exactly n chunks, skipping any in protected.

    Returns a list of (key, chunk) for the evicted chunks,
    or None if n evictions cannot be satisfied. The operation is atomic:
    if None is returned, no state changes are made.

    For ARC: ghost list cleanup (trimming to cache_capacity) is performed
    at the end of a successful eviction.
    """

get(key) abstractmethod

Find chunk in data structures. Returns None if not present.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
@abstractmethod
def get(self, key: OffloadKey) -> ChunkStatus | None:
    """Find chunk in data structures. Returns None if not present."""

insert(key, chunk) abstractmethod

Add a newly allocated chunk. For ARC: also removes from ghost lists.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
@abstractmethod
def insert(self, key: OffloadKey, chunk: ChunkStatus) -> None:
    """Add a newly allocated chunk. For ARC: also removes from ghost lists."""

mark_evictable(key)

Called when a chunk's ref_cnt transitions to 0.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
def mark_evictable(self, key: OffloadKey) -> None:
    """Called when a chunk's ref_cnt transitions to 0."""
    return

mark_non_evictable(key)

Called when a chunk's ref_cnt transitions from 0.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
def mark_non_evictable(self, key: OffloadKey) -> None:
    """Called when a chunk's ref_cnt transitions from 0."""
    return

on_request_finished(key_groups, insertion_only_keys, reused_keys, req_context)

Apply one request-scoped cache access in prefix order.

key_groups contains keys observed for each KV cache group in head-to-tail order. insertion_only_keys and reused_keys distinguish chunks only created by this request from ready chunks it actually reused, which matters for policies such as ARC where reuse changes frequency but insertion and pending observations do not.

The default forwards one touch per group, preserving compatibility for experimental out-of-tree policies while moving those touches to request finalization. Policies that distinguish insertion from reuse can override this hook and inspect the access classifications.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
def on_request_finished(
    self,
    key_groups: Sequence[Sequence[OffloadKey]],
    insertion_only_keys: set[OffloadKey],
    reused_keys: set[OffloadKey],
    req_context: ReqContext,
) -> None:
    """Apply one request-scoped cache access in prefix order.

    ``key_groups`` contains keys observed for each KV cache group in
    head-to-tail order. ``insertion_only_keys`` and ``reused_keys``
    distinguish chunks only created by this request from ready chunks it
    actually reused, which matters for policies such as ARC where reuse
    changes frequency but insertion and pending observations do not.

    The default forwards one touch per group, preserving compatibility
    for experimental out-of-tree policies while moving those touches to
    request finalization. Policies that distinguish insertion from reuse
    can override this hook and inspect the access classifications.
    """
    del insertion_only_keys, reused_keys
    for keys in key_groups:
        self.touch(keys, req_context)

on_store_miss(keys, req_context)

Observe store misses before their cache entries are inserted.

The default delegates to touch for compatibility with external policies. Policies may override this hook when a miss has distinct semantics, such as ARC adapting to a ghost-list hit.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
def on_store_miss(
    self, keys: Iterable[OffloadKey], req_context: ReqContext
) -> None:
    """Observe store misses before their cache entries are inserted.

    The default delegates to ``touch`` for compatibility with external
    policies. Policies may override this hook when a miss has distinct
    semantics, such as ARC adapting to a ghost-list hit.
    """
    self.touch(keys, req_context)

remove(key) abstractmethod

Remove a chunk (used to clean up after a failed store).

Source code in vllm/v1/kv_offload/cpu/policies/base.py
@abstractmethod
def remove(self, key: OffloadKey) -> None:
    """Remove a chunk (used to clean up after a failed store)."""

touch(keys, req_context) abstractmethod

Mark chunks as recently used.

Parameters:

  • keys

    (Iterable[OffloadKey]) –

    Chunks to mark as recently used.

  • req_context

    (ReqContext) –

    Per-request context for the request touching these chunks.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
@abstractmethod
def touch(self, keys: Iterable[OffloadKey], req_context: ReqContext) -> None:
    """Mark chunks as recently used.

    Args:
        keys: Chunks to mark as recently used.
        req_context: Per-request context for the request touching these chunks.

    """

ChunkStatus

Bases: Structure

Offloading status for a single chunk of KV data. Holds the following information:

ref_cnt - the current number of transfers using this chunk as a source. A value of -1 indicates the chunk is not yet ready to be read. chunk_id - index of the physical CPU buffer slot.

Attributes:

  • is_ready (bool) –

    Returns whether the chunk is ready to be read.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
class ChunkStatus(ctypes.Structure):
    """Offloading status for a single chunk of KV data.
    Holds the following information:

    ref_cnt - the current number of transfers using this chunk as a source.
        A value of -1 indicates the chunk is not yet ready to be read.
    chunk_id - index of the physical CPU buffer slot.
    """

    _fields_ = [("ref_cnt", ctypes.c_int32), ("chunk_id", ctypes.c_int64)]

    def __init__(self, chunk_id: int):
        super().__init__()
        # initialize chunk as "not ready" (ref_cnt = -1)
        self.ref_cnt = -1
        self.chunk_id = chunk_id

    @property
    def is_ready(self) -> bool:
        """Returns whether the chunk is ready to be read."""
        return self.ref_cnt >= 0

is_ready property

Returns whether the chunk is ready to be read.

order_request_keys(key_groups, req_context)

Return one head-to-tail order across all KV cache groups.

Source code in vllm/v1/kv_offload/cpu/policies/base.py
def order_request_keys(
    key_groups: Sequence[Sequence[OffloadKey]], req_context: ReqContext
) -> list[OffloadKey]:
    """Return one head-to-tail order across all KV cache groups."""
    keys = [key for group in key_groups for key in group]
    positions = {
        key: position
        for key in keys
        if (position := req_context.get_offload_key_position(key)) is not None
    }
    if len(positions) == len(keys):
        keys.sort(key=positions.__getitem__)
    return keys