vllm.config.engram
¶
Classes:
-
EngramConfig–Configuration for Engram embedding storage and sharding.
Functions:
-
model_has_engram_layers–Whether the model carries n-gram embedding layers.
EngramConfig
¶
Configuration for Engram embedding storage and sharding.
Methods:
-
compute_hash–Hash settings that affect embedding execution and graph structure.
-
get_parallel_size–Derive the embedding group size from the parallel configuration.
-
resolve_dp_shared_memory–Share host tables by default wherever the configuration permits.
-
verify_model_config–Reject Engram configuration for models without n-gram embeddings.
-
verify_parallel_config–Reject unsupported embedding parallel topologies.
Attributes:
-
cpu_offload(bool) –Store embedding weights in pinned CPU memory for UVA lookup.
-
dp_shared_memory(bool | None) –Share CPU-offloaded embedding weights between co-located
-
embedding_across_dp(bool) –Shard embeddings across TP and all DP ranks when enabled.
-
use_thp(bool) –Back private CPU-offloaded tables with transparent huge pages (best
Source code in vllm/config/engram.py
cpu_offload = True
class-attribute
instance-attribute
¶
Store embedding weights in pinned CPU memory for UVA lookup.
dp_shared_memory = None
class-attribute
instance-attribute
¶
Share CPU-offloaded embedding weights between co-located DP replicas. Each node stores one copy of every TP shard, reducing host memory without per-step Engram DP collectives. Requires sufficient /dev/shm capacity and a shared IPC namespace. Defaults to enabled whenever the other settings allow it, falling back to per-replica tables when DP replicas are not co-located on one node or /dev/shm cannot hold them.
embedding_across_dp = False
class-attribute
instance-attribute
¶
Shard embeddings across TP and all DP ranks when enabled. Otherwise, each DP rank has a separate TP-sharded embedding replica.
use_thp = False
class-attribute
instance-attribute
¶
Back private CPU-offloaded tables with transparent huge pages (best effort, falls back to ordinary pinned pages). Prefaulting the tables at startup takes longer. Requires cpu_offload without dp_shared_memory.
compute_hash()
¶
get_parallel_size(parallel_config)
¶
Derive the embedding group size from the parallel configuration.
Source code in vllm/config/engram.py
resolve_dp_shared_memory(parallel_config)
¶
Share host tables by default wherever the configuration permits.
Source code in vllm/config/engram.py
verify_model_config(model_config)
¶
Reject Engram configuration for models without n-gram embeddings.
Source code in vllm/config/engram.py
verify_parallel_config(parallel_config)
¶
Reject unsupported embedding parallel topologies.
Source code in vllm/config/engram.py
model_has_engram_layers(model_config)
¶
Whether the model carries n-gram embedding layers.