vllm.model_executor.models.transformers.fusers.attention
¶
Attention fuser: the module that dispatches to the attention interface.
Classes:
-
AttentionFuser–A module that dispatches through the Transformers attention interface.
Functions:
-
interface_call–The attention interface call in
forward, if it makes exactly one.
AttentionFuser
dataclass
¶
Bases: BaseFuser
A module that dispatches through the Transformers attention interface.
Methods:
-
layer_index–The layer
modulecomputes attention for, if it declares one. -
scale–The softmax scale
modulepasses to the interface, orNone. -
sinks–The per-head sink tensor
modulepasses to the interface, orNone. -
validate–Whether
modulewill actually dispatch to vLLM.
Attributes:
-
s_aux_expr(expr | None) –Source of the
s_aux=the module hands the interface, if it hands one. -
scale_expr(expr | None) –Source of the
scaling=the module hands the interface, if it hands one. -
source_cls(str) –Class of the HF module that dispatches (for logging).
Source code in vllm/model_executor/models/transformers/fusers/attention.py
81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 | |
s_aux_expr = None
class-attribute
instance-attribute
¶
Source of the s_aux= the module hands the interface, if it hands one.
scale_expr = None
class-attribute
instance-attribute
¶
Source of the scaling= the module hands the interface, if it hands one.
source_cls
instance-attribute
¶
Class of the HF module that dispatches (for logging).
layer_index(module)
¶
The layer module computes attention for, if it declares one.
Source code in vllm/model_executor/models/transformers/fusers/attention.py
scale(module)
¶
The softmax scale module passes to the interface, or None.
Source code in vllm/model_executor/models/transformers/fusers/attention.py
sinks(module)
¶
The per-head sink tensor module passes to the interface, or None.
Source code in vllm/model_executor/models/transformers/fusers/attention.py
validate(module, vllm_config)
¶
Whether module will actually dispatch to vLLM.
Source code in vllm/model_executor/models/transformers/fusers/attention.py
_is_interface_lookup(node)
¶
Whether node reads an entry out of ALL_ATTENTION_FUNCTIONS.
Source code in vllm/model_executor/models/transformers/fusers/attention.py
_resolve(node, module)
¶
The value of node on module, for literals and self.<attr>.
Source code in vllm/model_executor/models/transformers/fusers/attention.py
interface_call(forward)
cached
¶
The attention interface call in forward, if it makes exactly one.