vllm.v1.pool.flash_maxsim
¶
Fused Triton kernels for late-interaction (MaxSim) scoring.
Modules:
-
flash_maxsim_rerank–Flash-MaxSim rerank: one query vs variable-length docs in a packed tensor.
Functions:
-
flash_maxsim_rerank_direct–TRUE zero-copy: score query against docs scattered in a batch tensor.
flash_maxsim_rerank_direct(Q, batch_tensor, doc_offsets, doc_lengths, max_seqlen_d)
¶
TRUE zero-copy: score query against docs scattered in a batch tensor.
The kernel reads doc embeddings directly from batch_tensor at the positions specified by doc_offsets. No torch.stack, no torch.cat, no copy of any kind. The batch tensor is the model's output.
Memory for doc scoring: 0 bytes additional.
Parameters:
-
(Q¶Tensor) –[Lq, d] — single query embedding (from cache)
-
(batch_tensor¶Tensor) –[total_tokens, d] — the model's projected output tensor. Contains ALL requests' tokens (queries + docs + others).
-
(doc_offsets¶Tensor) –[B] int32 — start token index of each doc in batch_tensor
-
(doc_lengths¶Tensor) –[B] int32 — number of tokens per doc
-
(max_seqlen_d¶int) –int — max(doc_lengths)
Returns:
-
scores(Tensor) –[B] float32 — one MaxSim score per document