Skip to content

MooncakeConnector Usage Guide

About Mooncake

Mooncake aims to enhance the inference efficiency of large language models (LLMs), especially in slow object storage environments, by constructing a multi-level caching pool on high-speed interconnected DRAM/SSD resources. Compared to traditional caching systems, Mooncake utilizes (GPUDirect) RDMA technology to transfer data directly in a zero-copy manner, while maximizing the use of multi-NIC resources on a single machine.

For more details about Mooncake, please refer to Mooncake project and Mooncake documents.

Prerequisites

Installation

Install mooncake through pip: uv pip install mooncake-transfer-engine-cuda13.

vLLM defaults to CUDA 13. On a CUDA 12 environment install mooncake-transfer-engine instead — the two are the same release built against different CUDA majors, and the wrong one fails to import with libcudart.so.<major>: cannot open shared object file.

Refer to Mooncake official repository for more installation instructions

Usage

Prefiller Node (192.168.0.2)

vllm serve Qwen/Qwen2.5-7B-Instruct --port 8010 --kv-transfer-config '{"kv_connector":"MooncakeConnector","kv_role":"kv_producer"}'

Decoder Node (192.168.0.3)

vllm serve Qwen/Qwen2.5-7B-Instruct --port 8020 --kv-transfer-config '{"kv_connector":"MooncakeConnector","kv_role":"kv_consumer"}'

Proxy

python examples/disaggregated/mooncake_connector/mooncake_connector_proxy.py --prefill http://192.168.0.2:8010 --decode http://192.168.0.3:8020

Now you can send requests to the proxy server through port 8000.

For multimodal Encode-Prefill-Decode (encoder embeddings rather than KV cache), see ECMooncakeConnector Usage Guide.

Environment Variables

  • VLLM_MOONCAKE_BOOTSTRAP_PORT: Port for Mooncake bootstrap server

    • Default: 8998
    • Required only for prefiller instances
    • For headless instances, must be the same as the master instance
    • Each instance needs a unique port on its host; using the same port number across different hosts is fine
  • VLLM_MOONCAKE_CONNECTOR_TIMEOUT: Per-request timeout (in seconds) for MooncakeConnector HTTP calls to the bootstrap server, including prefiller worker registration at startup. (Optional)

    • Default: 30.0
    • Global rank 0 hosts the bootstrap server in the same process that mounts the Mooncake transfer engine's host segment, so a large MOONCAKE_GLOBAL_SEGMENT_SIZE can delay responses for several seconds during startup
    • Raise this if bootstrap server timeout warnings appear on hosts with very large host-memory segments
    • Bootstrap registration uses up to 3 attempts internally, with exponential backoff capped at 10 seconds; connection errors and timeouts are both retried
    • Once the attempts are exhausted the caller raises instead of blocking, so failures surface rather than hanging
  • WITH_NVIDIA_PEERMEM: Selects how mooncake registers GPU memory for RDMA. Read by mooncake, not vLLM.

    • Default: 1, which uses ibv_reg_mr() and requires the nvidia-peermem kernel module to be loaded
    • Set to 0 to use the DMA-BUF path, which does not need that module. Required on hosts where nvidia-peermem is not loaded, such as GB200
    • With the container image, pass it at run time: docker run -e WITH_NVIDIA_PEERMEM=0 ...
    • Symptom when left unset on such a host: Failed to register memory <addr>: Bad address [14] from rdma_context.cpp, and KV transfers fail
  • VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT: Timeout (in seconds) for automatically releasing the prefiller’s KV cache for a particular request. (Optional)

    • Default: 480
    • If a request is aborted and the decoder has not yet notified the prefiller, the prefill instance will release its KV-cache blocks after this timeout to avoid holding them indefinitely.

KV Transfer Config

KV Role Options

  • kv_producer: For prefiller instances that generate KV caches
  • kv_consumer: For decoder instances that consume KV caches from prefiller
  • kv_both: Enables symmetric functionality where the connector can act as both producer and consumer. This provides flexibility for experimental setups and scenarios where the role distinction is not predetermined.

kv_connector_extra_config

  • num_workers: Size of thread pool for one prefiller worker to transfer KV caches by mooncake. (default 10)
  • mooncake_protocol: Mooncake connector protocol. (default "rdma")
  • device_name: Comma-separated whitelist of RDMA devices (e.g. "mlx5_0,mlx5_1") to restrict topology discovery to. Empty discovers every device. Useful on hosts exposing a mix of InfiniBand and RoCE ports, where both peers must settle on the same link layer.
  • bootstrap_server_address: Optional "host:port" of an already running Mooncake bootstrap server. All prefiller ranks register with this endpoint instead of launching an embedded server. The caller owns the server lifetime and must configure the proxy to use the same endpoint. The example proxy assumes the bootstrap server is on the prefiller API host; a different host requires a proxy that supplies that address in remote_bootstrap_addr. Omit this option to use the embedded server and VLLM_MOONCAKE_BOOTSTRAP_PORT.

Example Scripts/Code

Refer to these example scripts in the vLLM repository: