vllm.v1.worker.gpu.eplb_utils
¶
Functions:
-
preserve_serving_state–Keep the elastic EP warmup out of the request pool and the KV cache.
-
step_eplb_after–Step EPLB after a model runner method completes successfully.
preserve_serving_state(model_runner)
¶
Keep the elastic EP warmup out of the request pool and the KV cache.
Source code in vllm/v1/worker/gpu/eplb_utils.py
step_eplb_after(*, is_dummy=False)
¶
Step EPLB after a model runner method completes successfully.