vllm.utils.gc_utils
¶
Classes:
-
GCDebugConfig–Config for GC Debugger.
-
GCDebugger–Debugger for GC which logs helpful information for GC understanding.
Functions:
-
freeze_gc_for_cudagraph_capture–Freeze and disable gc for the duration of bulk CUDA graph capture.
-
freeze_gc_heap–Freeze all objects tracked by the garbage collector. It should be invoked
-
maybe_attach_gc_debug_callback–Attached a callback for GC debug when VLLM_GC_DEBUG is enabled.
GCDebugConfig
¶
Config for GC Debugger. - 0: disable GC debugger - 1: enable GC debugger with gc.collect elapsed times - '{"top_objects":5}': enable GC debugger with top 5 collected objects
Source code in vllm/utils/gc_utils.py
GCDebugger
¶
Debugger for GC which logs helpful information for GC understanding. To enable, you should call maybe_attach_gc_debug_callback in the process.
Methods:
-
handle–Handles a GC event (e.g. GC start or GC finish).
Source code in vllm/utils/gc_utils.py
handle(phase, info)
¶
Handles a GC event (e.g. GC start or GC finish).
Source code in vllm/utils/gc_utils.py
_compute_detailed_type(o)
¶
Detailed object type.
TODO(Jialin): Further enhance the detailed type with element types for easier debugging. We tried but occasionally it would run into signals which kills the engine.
Source code in vllm/utils/gc_utils.py
_compute_top_gc_collected_objects(objects, top)
¶
Group collected objects by types.
Source code in vllm/utils/gc_utils.py
freeze_gc_for_cudagraph_capture()
¶
Freeze and disable gc for the duration of bulk CUDA graph capture.
A gc cycle during stream capture can invalidate the captured graph, e.g. a finalized Triton kernel unloads its module. Opt out with VLLM_ENABLE_CUDAGRAPH_GC=1.
Source code in vllm/utils/gc_utils.py
freeze_gc_heap()
¶
Freeze all objects tracked by the garbage collector. It should be invoked after server init / warmup, to reduce GC overhead from static objects during serving time.
Source code in vllm/utils/gc_utils.py
maybe_attach_gc_debug_callback()
¶
Attached a callback for GC debug when VLLM_GC_DEBUG is enabled.