vllm.v1.spec_decode.utils
¶
Functions:
-
copy_and_expand_dflash_inputs_kernel–Fused kernel for DFlash first-pass input setup.
-
copy_and_expand_eagle_inputs_kernel–Copy and expand inputs from the target model to the drafting buffers for Eagle
-
eagle_prepare_inputs_padded_kernel–Fused kernel for Eagle prepare_input_padded. This kernel computes the
-
eagle_prepare_next_token_padded_kernel–Fused kernel for Eagle prepare_next_token_ids_padded. This kernel computes the
-
eagle_step_slot_mapping_metadata_kernel–Fused kernel for EAGLE autoregressive step: updates positions, slot mapping,
-
eagle_step_update_slot_mapping_and_metadata–Fused update of slot mapping and metadata for one EAGLE autoregressive step.
-
extend_all_queries_by_N–Creates a new CommonAttentionMetadata with all query lengths increased by N.
-
next_power_of_2–Return the smallest power of 2 >= n.
-
unconditional_to_conditional_rates–Convert per-position unconditional rates to per-position conditional
-
update_num_computed_tokens_for_batch_change–Correct num_computed_tokens for async spec decode drift.
copy_and_expand_dflash_inputs_kernel(next_token_ids_ptr, target_positions_ptr, out_input_ids_ptr, out_context_positions_ptr, out_query_positions_ptr, out_context_slot_mapping_ptr, out_query_slot_mapping_ptr, out_token_indices_ptr, block_table_ptr, block_table_stride, query_start_loc_ptr, num_rejected_tokens_ptr, parallel_drafting_token_id, block_size, num_query_per_req, num_speculative_tokens, total_input_tokens, BLOCK_SIZE, HAS_NUM_REJECTED=False)
¶
Fused kernel for DFlash first-pass input setup.
Per request, this kernel: 1. Copies context positions from target_positions to out_context_positions. 2. Computes query positions (last_target_pos + 1 + offset) and writes them to out_query_positions. 3. Writes input_ids for query tokens: [next_token, mask, mask, ...]. 4. Computes slot_mapping for context and query positions into separate buffers via block_table lookup. 5. Writes token_indices_to_sample for the mask (speculative) tokens.
Source code in vllm/v1/spec_decode/utils.py
639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 | |
copy_and_expand_eagle_inputs_kernel(target_token_ids_ptr, target_positions_ptr, next_token_ids_ptr, out_input_ids_ptr, out_positions_ptr, out_is_rejected_token_mask_ptr, out_is_masked_token_mask_ptr, out_new_token_indices_ptr, out_hidden_state_mapping_ptr, query_start_loc_ptr, query_end_loc_ptr, padding_token_id, parallel_drafting_token_id, total_input_tokens, num_padding_slots_per_request, shift_input_ids, BLOCK_SIZE_TOKENS)
¶
Copy and expand inputs from the target model to the drafting buffers for Eagle speculative decoding. This kernel handles padding slots and parallel drafting tokens, if enabled.
Source code in vllm/v1/spec_decode/utils.py
428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 | |
eagle_prepare_inputs_padded_kernel(cu_num_draft_tokens_ptr, valid_sampled_tokens_count_ptr, query_start_loc_gpu_ptr, token_indices_to_sample_ptr, num_rejected_tokens_gpu_ptr, num_reqs)
¶
Fused kernel for Eagle prepare_input_padded. This kernel computes the token index to sample for each request, taking into account the number of draft tokens and the number of valid sampled tokens (which is one more than the number of accepted tokens).
Source code in vllm/v1/spec_decode/utils.py
eagle_prepare_next_token_padded_kernel(sampled_token_ids_ptr, discard_request_mask_ptr, backup_next_token_ids_ptr, next_token_ids_ptr, valid_sampled_tokens_count_ptr, vocab_size, num_sampled_tokens_per_req, num_reqs, stride_sampled_token_ids, BLOCK_SIZE_TOKENS)
¶
Fused kernel for Eagle prepare_next_token_ids_padded. This kernel computes the number of valid (1 + accepted) tokens for each request, and the corresponding "next" token id to sample from during speculative decoding. This is the "last accepted token" from the sampled tokens, or the backup token if no tokens were accepted or if the request is marked as discarded.
Source code in vllm/v1/spec_decode/utils.py
eagle_step_slot_mapping_metadata_kernel(positions_ptr, block_table_ptr, block_table_stride, seq_lens_ptr, out_clamped_positions_ptr, out_slot_mapping_ptr, block_size, max_model_len, n_blocks_per_req, PAD_ID, batch_size)
¶
Fused kernel for EAGLE autoregressive step: updates positions, slot mapping, and sequence lengths in a single kernel to reduce launch overhead.
Launched with input_batch_size threads. Threads with req_idx >= batch_size are cudagraph padding slots and only write PADDING_SLOT_ID.
Each real thread handles one request in the batch. Computes: - new_position = position + 1, clamped if exceeds max_model_len - slot_mapping from block table lookup - seq_lens += 1, or 1 if position exceeds max
Source code in vllm/v1/spec_decode/utils.py
eagle_step_update_slot_mapping_and_metadata(positions_1d, block_table_tensor, seq_lens, block_size, max_model_len, out_clamped_positions, out_slot_mapping, input_batch_size=None)
¶
Fused update of slot mapping and metadata for one EAGLE autoregressive step. Updates seq_lens in place. Writes to out_clamped_positions and out_slot_mapping.
When input_batch_size > batch_size, threads beyond batch_size write PADDING_SLOT_ID to out_slot_mapping for cudagraph padding.
Parameters:
-
(positions_1d¶Tensor) –[batch_size] current positions (use positions[0] for M-RoPE)
-
(block_table_tensor¶Tensor) –[batch_size, n_blocks_per_req]
-
(seq_lens¶Tensor) –[batch_size] updated in place
-
(block_size¶int) –KV cache block size
-
(max_model_len¶int) –max model length for clamping
-
(out_clamped_positions¶Tensor) –[batch_size] output buffer for clamped positions
-
(out_slot_mapping¶Tensor) –[input_batch_size] output buffer for slot mapping
-
(input_batch_size¶int | None, default:None) –total batch size including cudagraph padding; defaults to batch_size (no padding)
Source code in vllm/v1/spec_decode/utils.py
extend_all_queries_by_N(common_attn_metadata, N, arange, new_slot_mapping)
¶
Creates a new CommonAttentionMetadata with all query lengths increased by N. Also all seq lens are increased by N. This is useful e.g. in speculative decoding with parallel drafting, where we extend each sequence by N tokens and predict all tokens in one pass. The slot mapping is computed externally, as it requires more information.
Source code in vllm/v1/spec_decode/utils.py
next_power_of_2(n)
¶
Return the smallest power of 2 >= n.
unconditional_to_conditional_rates(rates)
¶
Convert per-position unconditional rates to per-position conditional rates for the early-terminating rejection loop (c_i = p_i / p_{i-1}).
Source code in vllm/v1/spec_decode/utils.py
update_num_computed_tokens_for_batch_change(num_computed_tokens, num_accepted_tokens, prev_positions, valid_sampled_token_count, prev_num_draft_tokens, cpu_num_computed_tokens)
¶
Correct num_computed_tokens for async spec decode drift.
Requests that had drafts: corrected = prev_gpu + valid_count. New requests or non-draft (e.g. prefills): use CPU value directly.