Skip to content

vllm.renderers.embed_utils

Attributes:

safe_load_prompt_embeds_async = make_async(safe_load_prompt_embeds) module-attribute

Async variant of safe_load_prompt_embeds that defers the decode to a thread-pool executor, so the asyncio event loop is not blocked by the base64 decode + torch.load work.

_truncated_reason(exc)

exc's message, capped, saying how much was left out when it is capped.

Without the count a capped reason reads like the whole one, so a caller debugging a large payload cannot tell that torch said more than this.

Source code in vllm/renderers/embed_utils.py
def _truncated_reason(exc: Exception) -> str:
    """`exc`'s message, capped, saying how much was left out when it is capped.

    Without the count a capped reason reads like the whole one, so a caller
    debugging a large payload cannot tell that torch said more than this.
    """
    reason: Final = str(exc).strip() or type(exc).__name__
    omitted: Final = len(reason) - _MAX_EMBED_ERROR_REASON_CHARS
    if omitted <= 0:
        return reason
    return (
        f"{reason[:_MAX_EMBED_ERROR_REASON_CHARS]}"
        f"... ({omitted} more characters truncated)"
    )