vllm.v1.attention.backend
¶
Classes:
-
AttentionBackend–Abstract class for attention backends.
-
AttentionCGSupport–Constants for the cudagraph support of the attention backend
-
AttentionImpl–Standard attention implementation with forward method.
-
AttentionImplBase–Base class for attention implementations.
-
AttentionMetadataBuilder– -
AttentionType–Attention type.
-
CommonAttentionMetadata–Per-batch attention metadata, shared across layers and backends.
-
MLAAttentionImpl–MLA attention implementation with forward_mqa and forward_mha methods.
Functions:
-
max_decode_query_len–Widest request a spec-as-decode builder treats as a decode.
-
subclass_attention_backend–Return a new subclass where
get_builder_clsreturnsbuilder_cls.
AttentionBackend
¶
Bases: ABC
Abstract class for attention backends.
Methods:
-
customize_spec–Adjust the layer's KV cache spec for this backend's kernels. Used when the
-
supported_kv_cache_layouts–Layouts this backend's kernels can consume, most preferred first, or
-
supports_attn_type–Check if backend supports a given attention type.
-
supports_device_cpu_query_lens_mismatch–Whether this backend can run a batch whose device query_start_loc disagrees
-
supports_non_causal–Check if backend supports non-causal (bidirectional) attention
Source code in vllm/v1/attention/backend.py
61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 | |
customize_spec(spec)
classmethod
¶
Adjust the layer's KV cache spec for this backend's kernels. Used when the kernels want KV packed in a specific way.
NOTE: temporary compatibility API. Today the Attention layer builds the spec from the model config and the backend only gets to adjust it post-hoc; the end state is for the backend to build and return the spec directly, at which point this hook goes away.
(see: https://github.com/vllm-project/vllm/issues/42449)
Source code in vllm/v1/attention/backend.py
supported_kv_cache_layouts()
classmethod
¶
Layouts this backend's kernels can consume, most preferred first, or None when the kernels consume any layout and express no preference.
Source code in vllm/v1/attention/backend.py
supports_attn_type(attn_type)
classmethod
¶
Check if backend supports a given attention type.
By default, only supports decoder attention. Backends should override this to support other attention types.
Source code in vllm/v1/attention/backend.py
supports_device_cpu_query_lens_mismatch()
classmethod
¶
Whether this backend can run a batch whose device query_start_loc disagrees with the CPU one; backends that plan off the CPU query lengths must opt out.
Currently only verification requests are affected: adaptive verification trims their drafts on device. On the CPU the draft budget is evenly distributed across requests, so the total draft budget, the decode/prefill split point and the CPU prefill query lengths all stay correct.
SSM backends opt out: their recurrent-state planning is built from the CPU per-request boundaries, which the trimmed batch no longer matches.
Source code in vllm/v1/attention/backend.py
supports_non_causal()
classmethod
¶
Check if backend supports non-causal (bidirectional) attention for decoder models.
Unlike ENCODER_ONLY attention type which implies a different execution model, this refers to non-causal attention within the standard paged-KV-cache decoder path.
Source code in vllm/v1/attention/backend.py
AttentionCGSupport
¶
Bases: Enum
Constants for the cudagraph support of the attention backend Here we do not consider the cascade attention, as currently it is never cudagraph supported.
Attributes:
-
ALWAYS–Cudagraph always supported; supports mixed-prefill-decode
-
NEVER–NO cudagraph support
-
UNIFORM_BATCH–Cudagraph supported for batches the only contain query lengths that are
-
UNIFORM_SINGLE_TOKEN_DECODE–Cudagraph supported for batches the only contain query_len==1 decodes
Source code in vllm/v1/attention/backend.py
ALWAYS = 3
class-attribute
instance-attribute
¶
Cudagraph always supported; supports mixed-prefill-decode
NEVER = 0
class-attribute
instance-attribute
¶
NO cudagraph support
UNIFORM_BATCH = 2
class-attribute
instance-attribute
¶
Cudagraph supported for batches the only contain query lengths that are the same, this can be used for spec-decode i.e. "decodes" are 1 + num_speculative_tokens
UNIFORM_SINGLE_TOKEN_DECODE = 1
class-attribute
instance-attribute
¶
Cudagraph supported for batches the only contain query_len==1 decodes
AttentionImpl
¶
Bases: AttentionImplBase[T], Generic[T]
Standard attention implementation with forward method.
Methods:
-
do_qk_norm_mrope_kvcache_update–Apply QK-norm and MRoPE, then write K/V to the cache.
-
do_qk_norm_rope_kvcache_update–If
fused_qk_norm_rope_kvcache_supportedreturns True, this method -
do_rope_and_kv_cache_update–If
fused_rope_kvcache_supportedreturns True, this method will be called -
fused_output_quant_supported–Does this attention implementation support fused output quantization.
-
fused_qk_norm_mrope_kvcache_supported–Whether this implementation supports fused QKNorm+MRoPE+KVCache.
-
fused_qk_norm_rope_kvcache_supported–Does this attention implementation support fused QKNorm+RoPE+KVCache fusion.
-
fused_rope_kvcache_supported–Does this attention implementation support RoPE+KVCache fusion.
Attributes:
-
kv_quant_mode(KVQuantMode) –Return the KV cache quantization mode for this layer.
Source code in vllm/v1/attention/backend.py
910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 | |
kv_quant_mode
property
¶
Return the KV cache quantization mode for this layer.
do_qk_norm_mrope_kvcache_update(layer, qkv, q_out, positions, q_weight, k_weight, rms_norm_eps, cos_sin_cache, is_neox, mrope_section, is_interleaved, rotary_dim, kv_cache, layer_slot_mapping)
¶
Apply QK-norm and MRoPE, then write K/V to the cache.
Source code in vllm/v1/attention/backend.py
do_qk_norm_rope_kvcache_update(layer, qkv, q_out, k_out, positions, q_weight, k_weight, rms_norm_eps, cos_sin_cache, is_neox, kv_cache, layer_slot_mapping)
¶
If fused_qk_norm_rope_kvcache_supported returns True, this method
will be called by the fused custom op. Applies QK-norm + RoPE and
writes K/V to the KV cache. Results are written to the pre-allocated
q_out and k_out tensors; V is split from QKV at the graph level.
Source code in vllm/v1/attention/backend.py
do_rope_and_kv_cache_update(layer, query, key, value, positions, cos_sin_cache, is_neox, kv_cache, layer_slot_mapping)
¶
If fused_rope_kvcache_supported returns True, this method will be called
by torch.ops.vllm.fused_rope_and_unified_kv_cache_update
to perform the inplace RoPE and KV cache update.
Source code in vllm/v1/attention/backend.py
fused_output_quant_supported(quant_key)
¶
Does this attention implementation support fused output quantization. This is used by the AttnFusionPass to only fuse output quantization onto implementations that support it.
Parameters:
Returns:
-
bool–is fusion supported for this type of quantization
Source code in vllm/v1/attention/backend.py
fused_qk_norm_mrope_kvcache_supported()
¶
fused_qk_norm_rope_kvcache_supported()
¶
Does this attention implementation support fused QKNorm+RoPE+KVCache fusion. This is used by the QkNormRopeKvCachePattern to only fuse the QKNorm ops with the RoPE ops and the KV cache update for implementations that support it.
Source code in vllm/v1/attention/backend.py
fused_rope_kvcache_supported()
¶
Does this attention implementation support RoPE+KVCache fusion. This is used by the RopeKVCacheFusionPass to only fuse the RoPE ops with the KV cache update for implementations that support it.
Source code in vllm/v1/attention/backend.py
AttentionImplBase
¶
Base class for attention implementations.
Contains common attributes and initialization logic shared by both standard AttentionImpl and MLAAttentionImpl. Does not define a forward method - subclasses define their own forward interfaces.
Methods:
-
prepare_for_batch–Prepare implementation-specific state for the current batch.
Source code in vllm/v1/attention/backend.py
806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 | |
AttentionMetadataBuilder
¶
Methods:
-
build–Central method that builds attention metadata.
-
build_for_cudagraph_capture–Build attention metadata for CUDA graph capture. Uses build by default.
-
build_for_drafting–Build attention metadata for draft model. Uses build by default.
-
get_cudagraph_support–Get the cudagraph support level of this builder class.
-
get_varlen_cudagraph_max_query_len–Get the largest per-request query length L of the variable-length
-
update_block_table–Update the block table for the attention metadata.
-
update_draft_decode_metadata–Update step-dependent draft decode metadata in place.
Source code in vllm/v1/attention/backend.py
599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 | |
build(common_prefix_len, common_attn_metadata, fast_build=False)
abstractmethod
¶
Central method that builds attention metadata. Some builders (MLA) require reorder_batch to be called prior to build.
Parameters:
-
(common_prefix_len¶int) –The length of the common prefix of the batch.
-
(common_attn_metadata¶CommonAttentionMetadata) –The common attention metadata.
-
(fast_build¶bool, default:False) –The meta-data will prioritize speed of building over then speed at execution. Can be used for spec-decode where the result of a build call may only be used for few layers/iters.
Source code in vllm/v1/attention/backend.py
build_for_cudagraph_capture(common_attn_metadata)
¶
Build attention metadata for CUDA graph capture. Uses build by default. Subclasses that override this method should call self.build or super().build_for_cudagraph_capture.
Source code in vllm/v1/attention/backend.py
build_for_drafting(common_attn_metadata, draft_index)
¶
Build attention metadata for draft model. Uses build by default.
Parameters:
-
(common_attn_metadata¶CommonAttentionMetadata) –The common attention metadata.
-
(draft_index¶int) –The index of the current draft operation. When speculating a chain of tokens, this index refers to the draft attempt for the i-th token. For tree-based attention, this index instead refers to the draft attempt for the i-th level in the tree of tokens.
Source code in vllm/v1/attention/backend.py
get_cudagraph_support(vllm_config, kv_cache_spec)
classmethod
¶
Get the cudagraph support level of this builder class.
Source code in vllm/v1/attention/backend.py
get_varlen_cudagraph_max_query_len(vllm_config, kv_cache_spec)
classmethod
¶
Get the largest per-request query length L of the variable-length decode batches a FULL cudagraph of this builder class can replay.
The graph must replay any decode batch in which every real request has between 1 and L query tokens and every padding request has 0. Lengths come from the device query_start_loc; host metadata carries only the token count and an upper bound on the per-request length. Batches with a prefill never replay these graphs. Whether host metadata may understate device query lengths at all is a separate backend question; see supports_device_cpu_query_lens_mismatch().
Returns:
-
int | None–L, or None for builders reporting ALWAYS, which replay any batch,
-
int | None–and for builders that cannot replay variable-length batches.
Source code in vllm/v1/attention/backend.py
update_block_table(metadata, blk_table, slot_mapping)
¶
Update the block table for the attention metadata. Faster when theres multiple kv-cache groups that create virtually the same metadata but just with different block tables.
Only needs to be implemented if supports_update_block_table is True.
Source code in vllm/v1/attention/backend.py
update_draft_decode_metadata(metadata)
¶
Update step-dependent draft decode metadata in place.
The fused draft loop may call this method during full CUDA graph capture. CUDA graph replay does not run this Python method, so implementations must emit capture-safe operations and keep replayed tensor state in persistent storage.
Source code in vllm/v1/attention/backend.py
AttentionType
¶
Attention type.
Use string to be compatible with torch.compile.
Attributes:
-
DECODER–Decoder attention between previous layer Q/K/V.
-
ENCODER–Encoder attention between previous layer Q/K/V for encoder-decoder.
-
ENCODER_DECODER–Attention between dec. Q and enc. K/V for encoder-decoder.
-
ENCODER_ONLY–Encoder attention between previous layer Q/K/V.
Source code in vllm/v1/attention/backend.py
DECODER = 'decoder'
class-attribute
instance-attribute
¶
Decoder attention between previous layer Q/K/V.
ENCODER = 'encoder'
class-attribute
instance-attribute
¶
Encoder attention between previous layer Q/K/V for encoder-decoder.
ENCODER_DECODER = 'encoder_decoder'
class-attribute
instance-attribute
¶
Attention between dec. Q and enc. K/V for encoder-decoder.
ENCODER_ONLY = 'encoder_only'
class-attribute
instance-attribute
¶
Encoder attention between previous layer Q/K/V.
CommonAttentionMetadata
dataclass
¶
Per-batch attention metadata, shared across layers and backends. AttentionMetadataBuilder instances use it to construct per-layer metadata.
For many of the tensors we keep both GPU and CPU versions.
Methods:
-
compute_num_computed_tokens–Compute num_computed_tokens on device (seq_lens - query_lens).
-
naive_query_lens–Naive because it assumes that query ends where the next query starts.
-
token_to_req_indices–Build or reuse the per-token request index mapping.
Attributes:
-
dcp_local_seq_lens(Tensor | None) –Sequence lengths of the local rank in decode context parallelism world
-
dcp_local_seq_lens_cpu_upper_bound(Tensor | None) –(batch_size,) CPU upper bound on dcp_local_seq_lens. Under PCP+DCP it
-
is_prefilling(Tensor | None) –(batch_size,) bool tensor: True if request is still in prefill phase
-
max_query_len(int) –Longest query in batch
-
max_seq_len(int) –Longest context length (may be an upper bound)
-
mm_req_doc_ranges(dict[int, list[tuple[int, int]]] | None) –PrefixLM bidirectional ranges for multimodal tokens. Maps
-
num_actual_tokens(int) –Total number of tokens in batch
-
num_reqs(int) –Number of requests
-
positions(Tensor | None) –(num_actual_tokens,) token positions. Optional; set when the caller
-
query_start_loc_cpu(Tensor) –(batch_size + 1,), the start location of each request in query Tensor
-
replayssm_decode_base_cpu(Tensor | None) –(batch_size,) CPU ring origin for Mamba2 ReplaySSM decode: num_computed
-
req_idx(ndarray | None) –(batch_size,) index of each row's request in the runner's request
-
rswa_prefix_lens(Tensor | None) –(batch_size,) per-request prefix length (prompt/image token count) for
-
seq_lens(Tensor) –(batch_size,), the number of computed tokens for each request
-
seq_lens_cpu_upper_bound(Tensor | None) –(batch_size,) CPU upper bound on seq_lens. Precise for prefill rows
Source code in vllm/v1/attention/backend.py
388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 | |
dcp_local_seq_lens = None
class-attribute
instance-attribute
¶
Sequence lengths of the local rank in decode context parallelism world
dcp_local_seq_lens_cpu_upper_bound = None
class-attribute
instance-attribute
¶
(batch_size,) CPU upper bound on dcp_local_seq_lens. Under PCP+DCP it holds, on every row, the largest shard any DCP rank has of the row's whole request, identical on every PCP rank, so the sparse backends pad their KV gather to it.
is_prefilling = None
class-attribute
instance-attribute
¶
(batch_size,) bool tensor: True if request is still in prefill phase (num_computed_tokens < num_prompt_tokens). Used by some backends to distinguish actual decodes from short extends.
max_query_len
instance-attribute
¶
Longest query in batch
max_seq_len
instance-attribute
¶
Longest context length (may be an upper bound)
mm_req_doc_ranges = None
class-attribute
instance-attribute
¶
PrefixLM bidirectional ranges for multimodal tokens. Maps request index to list of (start, end) token position ranges where bidirectional attention should apply. None for text-only batches or non-PrefixLM models. A request's ranges must not overlap.
num_actual_tokens
instance-attribute
¶
Total number of tokens in batch
num_reqs
instance-attribute
¶
Number of requests
positions = None
class-attribute
instance-attribute
¶
(num_actual_tokens,) token positions. Optional; set when the caller has positions available so that builders can pre-compute position-dependent sparse metadata for DeepSeek V4 C128A layers.
query_start_loc_cpu
instance-attribute
¶
(batch_size + 1,), the start location of each request in query Tensor
replayssm_decode_base_cpu = None
class-attribute
instance-attribute
¶
(batch_size,) CPU ring origin for Mamba2 ReplaySSM decode: num_computed at the current decode run's last full-state write. write_pos counts from here, so a preemption-resumed request re-anchors past the prompt boundary.
req_idx = None
class-attribute
instance-attribute
¶
(batch_size,) index of each row's request in the runner's request table. Rows of one request are adjacent, so equal neighbours are PCP chunks sharing one KV context.
rswa_prefix_lens = None
class-attribute
instance-attribute
¶
(batch_size,) per-request prefix length (prompt/image token count) for
Reference Sliding Window Attention (R-SWA). Tokens with logical index below
this stay globally visible; later (generated) tokens additionally see a
fixed sliding window. None disables R-SWA. The attention backend copies this
into its own persistent buffer and reads rswa_window from model config.
seq_lens
instance-attribute
¶
(batch_size,), the number of computed tokens for each request
seq_lens_cpu_upper_bound = None
class-attribute
instance-attribute
¶
(batch_size,) CPU upper bound on seq_lens. Precise for prefill rows and for all rows outside async spec decode; optimistic for async-spec decode rows (assumes every draft was accepted). Not safe for kernels that need exact per-row context lengths on decode rows.
compute_num_computed_tokens()
¶
Compute num_computed_tokens on device (seq_lens - query_lens).
Source code in vllm/v1/attention/backend.py
naive_query_lens()
¶
Naive because it assumes that query ends where the next query starts.
token_to_req_indices(buffer)
¶
Build or reuse the per-token request index mapping.
Source code in vllm/v1/attention/backend.py
MLAAttentionImpl
¶
Bases: AttentionImplBase[T], Generic[T]
MLA attention implementation with forward_mqa and forward_mha methods.
Methods:
-
forward_mha–MHA-style prefill forward pass.
-
forward_mqa–MQA-style decode forward pass.
-
fused_output_quant_supported–Does this attention implementation support fused output quantization.
Source code in vllm/v1/attention/backend.py
1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 | |
forward_mha(q, kv_c_normed, k_pe, kv_c_and_k_pe_cache, attn_metadata, k_scale, output, output_scale=None)
¶
MHA-style prefill forward pass.
Source code in vllm/v1/attention/backend.py
forward_mqa(q, kv_c_and_k_pe_cache, attn_metadata, layer)
abstractmethod
¶
MQA-style decode forward pass.
Source code in vllm/v1/attention/backend.py
fused_output_quant_supported(quant_key)
¶
Does this attention implementation support fused output quantization. Since MLA quantization is done manually in forward_impl (common code), all MLA backends support it by default.
Source code in vllm/v1/attention/backend.py
max_decode_query_len(vllm_config)
¶
Widest request a spec-as-decode builder treats as a decode.
On model runner V2 this is a verification request, 1 + num_speculative_tokens: no V2 speculator builds a wider query. The reorder threshold and the code that mirrors the builders' decode/prefill split all read this, so they cannot drift apart.
Source code in vllm/v1/attention/backend.py
subclass_attention_backend(name_prefix, attention_backend_cls, builder_cls)
¶
Return a new subclass where get_builder_cls returns builder_cls.