vllm.models.glm5next.common.model
¶
_LOGIT_SCALE = 1.0
module-attribute
¶
Output logit scale. A GLM-5.3-Flash trained value that neither the
checkpoint nor Glm5NextTextConfig carries.
_MHC_POST_MULT_VALUE = 2.0
module-attribute
¶
mHC post-multiplier. A GLM-5.3-Flash trained value that neither the
checkpoint nor Glm5NextTextConfig carries.
_MHC_TAU = 0.05
module-attribute
¶
mHC routing temperature. A GLM-5.3-Flash trained value that neither the
checkpoint nor Glm5NextTextConfig carries.
_VISION_RMS_NORM_EPS = 1e-06
module-attribute
¶
Vision tower RMSNorm epsilon.
GLM-5.3-Flash checkpoints ship vision_config.rms_norm_eps = 1e-5, but the
vision tower was trained with 1e-6. Serving with 1e-5 drifts the RMSNorm and
produces repetitive/degraded image descriptions, so force the trained value
regardless of the checkpoint field.
_dequant_fp8_block(weight_fp8, scale_inv, block_size=128)
¶
Dequantize a block-FP8 (e4m3) weight with per-block scale to BF16.
Unlike scaled_dequantize this tolerates a non-divisible (partial last
block) shape by zero-padding to a multiple of block_size before the
scale broadcast and trimming back afterwards (e.g. kv_a_proj_with_mqa is
576 rows = 4*128 + 64).
Source code in vllm/models/glm5next/common/model.py
_try_load_fp8_attn_proj(name, tensor, buf, params_dict, loaded_params, kv_a_pad_size)
¶
Dequantize FP8 q_a_proj / kv_a_proj_with_mqa / o_proj to BF16 on load.
The FP8 checkpoint stores these as block-FP8 (weight + weight_scale_inv),
but the model holds them in BF16 (fused_qkv_a_proj is always BF16 via
DeepSeekV2FusedQkvAProjLinear; o_proj is excluded by
modules_to_not_convert). When the model target is BF16 (no
weight_scale_inv param) we dequantize; otherwise we return False so the
normal stacked/direct path loads the FP8 tensor as-is.
Source code in vllm/models/glm5next/common/model.py
1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 | |
_validate_supported_config(config)
¶
Reject checkpoints using config options this implementation lacks.
The kpool indexer kernels always keep the incomplete trailing pool, so a checkpoint asking otherwise would be served silently wrong.