vllm.entrypoints.speech_to_text.transcription.protocol
¶
Classes:
-
TranscriptionDiarizedSegment–A speaker-attributed transcription segment.
-
TranscriptionRequest– -
TranscriptionResponse– -
TranscriptionResponseDiarized–OpenAI-compatible diarized transcription response.
-
TranscriptionResponseVerbose– -
TranscriptionSegment– -
TranscriptionWord–
TranscriptionDiarizedSegment
¶
Bases: OpenAIBaseModel
A speaker-attributed transcription segment.
Source code in vllm/entrypoints/speech_to_text/transcription/protocol.py
TranscriptionRequest
¶
Bases: OpenAIBaseModel
Attributes:
-
file(UploadFile) –The audio file object (not file name) to transcribe, in one of these
-
frequency_penalty(float | None) –The frequency penalty to use for sampling.
-
hotwords(str | None) –hotwords refers to a list of important words or phrases that the model
-
include_stop_str_in_output(bool) –Whether to include the stop strings in output text.
-
language(str | None) –The language of the input audio.
-
length_penalty(float) –Length penalty to be used for beam search.
-
max_completion_tokens(int | None) –The maximum number of tokens to generate.
-
min_p(float | None) –Filters out tokens with a probability lower than
min_p, ensuring a -
model(str | None) –ID of the model to use.
-
n(int) –The number of beams to be used in beam search.
-
presence_penalty(float | None) –The presence penalty to use for sampling.
-
prompt(str) –An optional text to guide the model's style or continue a previous audio
-
repetition_penalty(float | None) –The repetition penalty to use for sampling.
-
response_format(TranscriptionResponseFormat) –The format of the output, in one of these options:
json,text,srt, -
seed(int | None) –The seed to use for sampling.
-
stream(bool | None) –When set, it will enable output to be streamed in a similar fashion
-
temperature(float) –The sampling temperature, between 0 and 1.
-
timestamp_granularities(list[Literal['word', 'segment']]) –The timestamp granularities to populate for this transcription.
-
to_language(str | None) –The language of the output audio we transcribe to.
-
top_k(int | None) –Limits sampling to the
kmost probable tokens at each step. -
top_p(float | None) –Enables nucleus (top-p) sampling, where tokens are selected from the
-
use_beam_search(bool) –Whether or not beam search should be used.
Source code in vllm/entrypoints/speech_to_text/transcription/protocol.py
52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 | |
file
instance-attribute
¶
The audio file object (not file name) to transcribe, in one of these formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, or webm.
frequency_penalty = 0.0
class-attribute
instance-attribute
¶
The frequency penalty to use for sampling.
hotwords = None
class-attribute
instance-attribute
¶
hotwords refers to a list of important words or phrases that the model should pay extra attention to during transcription.
include_stop_str_in_output = False
class-attribute
instance-attribute
¶
Whether to include the stop strings in output text.
language = None
class-attribute
instance-attribute
¶
The language of the input audio.
Supplying the input language in ISO-639-1 format will improve accuracy and latency.
length_penalty = 1.0
class-attribute
instance-attribute
¶
Length penalty to be used for beam search.
max_completion_tokens = None
class-attribute
instance-attribute
¶
The maximum number of tokens to generate.
min_p = None
class-attribute
instance-attribute
¶
Filters out tokens with a probability lower than min_p, ensuring a
minimum likelihood threshold during sampling.
model = None
class-attribute
instance-attribute
¶
ID of the model to use.
n = 1
class-attribute
instance-attribute
¶
The number of beams to be used in beam search.
presence_penalty = 0.0
class-attribute
instance-attribute
¶
The presence penalty to use for sampling.
prompt = Field(default='')
class-attribute
instance-attribute
¶
An optional text to guide the model's style or continue a previous audio segment.
The prompt should match the audio language.
repetition_penalty = None
class-attribute
instance-attribute
¶
The repetition penalty to use for sampling.
response_format = Field(default='json')
class-attribute
instance-attribute
¶
The format of the output, in one of these options: json, text, srt,
verbose_json, or vtt.
seed = Field(None, ge=_LONG_INFO.min, le=_LONG_INFO.max)
class-attribute
instance-attribute
¶
The seed to use for sampling.
stream = False
class-attribute
instance-attribute
¶
When set, it will enable output to be streamed in a similar fashion as the Chat Completion endpoint.
temperature = Field(default=0.0)
class-attribute
instance-attribute
¶
The sampling temperature, between 0 and 1.
Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused / deterministic. If set to 0, the model will use log probability to automatically increase the temperature until certain thresholds are hit.
timestamp_granularities = Field(alias='timestamp_granularities[]', default=[])
class-attribute
instance-attribute
¶
The timestamp granularities to populate for this transcription.
response_format must be set verbose_json to use timestamp granularities.
Either or both of these options are supported: word, or segment. Note:
There is no additional latency for segment timestamps, but generating word
timestamps incurs additional latency.
to_language = None
class-attribute
instance-attribute
¶
The language of the output audio we transcribe to.
Please note that this is not currently used by supported models at this time, but it is a placeholder for future use, matching translation api.
top_k = None
class-attribute
instance-attribute
¶
Limits sampling to the k most probable tokens at each step.
top_p = None
class-attribute
instance-attribute
¶
Enables nucleus (top-p) sampling, where tokens are selected from the
smallest possible set whose cumulative probability exceeds p.
use_beam_search = False
class-attribute
instance-attribute
¶
Whether or not beam search should be used.
TranscriptionResponse
¶
Bases: OpenAIBaseModel
Attributes:
Source code in vllm/entrypoints/speech_to_text/transcription/protocol.py
text
instance-attribute
¶
The transcribed text.
TranscriptionResponseDiarized
¶
Bases: OpenAIBaseModel
OpenAI-compatible diarized transcription response.
Source code in vllm/entrypoints/speech_to_text/transcription/protocol.py
TranscriptionResponseVerbose
¶
Bases: OpenAIBaseModel
Attributes:
-
duration(float) –The duration of the input audio.
-
language(str) –The language of the input audio.
-
segments(list[TranscriptionSegment] | None) –Segments of the transcribed text and their corresponding details.
-
text(str) –The transcribed text.
-
words(list[TranscriptionWord] | None) –Extracted words and their corresponding timestamps.
Source code in vllm/entrypoints/speech_to_text/transcription/protocol.py
duration
instance-attribute
¶
The duration of the input audio.
language
instance-attribute
¶
The language of the input audio.
segments = None
class-attribute
instance-attribute
¶
Segments of the transcribed text and their corresponding details.
text
instance-attribute
¶
The transcribed text.
words = None
class-attribute
instance-attribute
¶
Extracted words and their corresponding timestamps.
TranscriptionSegment
¶
Bases: OpenAIBaseModel
Attributes:
-
avg_logprob(float) –Average logprob of the segment.
-
compression_ratio(float) –Compression ratio of the segment.
-
end(float) –End time of the segment in seconds.
-
id(int) –Unique identifier of the segment.
-
no_speech_prob(float | None) –Probability of no speech in the segment.
-
seek(int) –Seek offset of the segment.
-
start(float) –Start time of the segment in seconds.
-
temperature(float) –Temperature parameter used for generating the segment.
-
text(str) –Text content of the segment.
-
tokens(list[int]) –Array of token IDs for the text content.
Source code in vllm/entrypoints/speech_to_text/transcription/protocol.py
avg_logprob
instance-attribute
¶
Average logprob of the segment.
If the value is lower than -1, consider the logprobs failed.
compression_ratio
instance-attribute
¶
Compression ratio of the segment.
If the value is greater than 2.4, consider the compression failed.
end
instance-attribute
¶
End time of the segment in seconds.
id
instance-attribute
¶
Unique identifier of the segment.
no_speech_prob = None
class-attribute
instance-attribute
¶
Probability of no speech in the segment.
If the value is higher than 1.0 and the avg_logprob is below -1, consider
this segment silent.
seek
instance-attribute
¶
Seek offset of the segment.
start
instance-attribute
¶
Start time of the segment in seconds.
temperature
instance-attribute
¶
Temperature parameter used for generating the segment.
text
instance-attribute
¶
Text content of the segment.
tokens
instance-attribute
¶
Array of token IDs for the text content.