dependabot[bot] opened a new pull request, #39371:
URL: https://github.com/apache/beam/pull/39371

   Bumps [vllm](https://github.com/vllm-project/vllm) from 0.10.1.1 to 0.24.0.
   <details>
   <summary>Release notes</summary>
   <p><em>Sourced from <a 
href="https://github.com/vllm-project/vllm/releases";>vllm's 
releases</a>.</em></p>
   <blockquote>
   <h2>v0.24.0</h2>
   <h1>vLLM v0.24.0 Release Notes</h1>
   <h2>Highlights</h2>
   <p>This release features 571 commits from 256 contributors (77 new)!</p>
   <ul>
   <li><strong>MiniMax-M3</strong>: Added support for the new 
<strong>MiniMax-M3</strong> model (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45381";>#45381</a>), 
with a fast follow-on of BF16/FP8 indexer via MSA (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45892";>#45892</a>), 
MXFP4 support (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45896";>#45896</a>), 
FP8 sparse GQA (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45744";>#45744</a>), 
and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45725";>#45725</a>), 
fp8_per_channel for bf16 weights on MI300X (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45854";>#45854</a>), 
FP8 KV-cache fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45720";>#45720</a>), 
and packed-modules mapping (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45794";>#45794</a>). 
A MiniMax-M
 2 perf regression was also fixed (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45935";>#45935</a>).</li>
   <li><strong>DeepSeek-V4 keeps maturing</strong>: Following its debut, 
DeepSeek-V4 received another large optimization pass — a FlashInfer sparse 
index cache (2–4% TTFT) (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45863";>#45863</a>), 
prefill chunk-planning optimization (4% E2E throughput) (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45061";>#45061</a>), 
a cluster-cooperative topK kernel for low-latency (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43008";>#43008</a>), 
contiguous per-block KV allocations (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44577";>#44577</a>), 
TEP=16 for the block-FP8 shared expert (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46001";>#46001</a>), 
and native DSA indexer decode for <code>next_n &gt; 2</code> on SM100 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45322";>#45322</a>). 
It is now enabled on <strong>SM120</strong> alongside GLM-5.1 (<a href="h
 ttps://redirect.github.com/vllm-project/vllm/issues/43477">#43477</a>), with 
XPU (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44144";>#44144</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44517";>#44517</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45240";>#45240</a>) 
and ROCm (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44899";>#44899</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45103";>#45103</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45681";>#45681</a>) 
attention/MoE paths added.</li>
   <li><strong>Model Runner V2 (MRv2) continues to expand</strong>: MRv2 now 
<strong>supports quantized models by default</strong> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44446";>#44446</a>), 
enables <strong>GraniteMoE by default</strong> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45461";>#45461</a>), 
and gained migration of Qwen + DeepSeek-V2 MoE models (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42667";>#42667</a>), 
DFlash speculative decoding (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44586";>#44586</a>), 
and more accurate FP32 Gumbel sampling (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45996";>#45996</a>).</li>
   <li><strong>Streaming Parser Engine</strong>: A new streaming parser engine 
unifies tool-call/reasoning parsing across models, with parsers for Qwen3 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45413";>#45413</a>), 
MiniMax-M2 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45701";>#45701</a>), 
GLM-4.7/5.1/5.2 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45915";>#45915</a>), 
and Nemotron V3 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45755";>#45755</a>).</li>
   <li><strong>Diffusion LLMs</strong>: Added <strong>DiffusionGemma</strong> 
(<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45163";>#45163</a>), 
including a CPU path (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45690";>#45690</a>) 
and structured-output guardrails for diffusion decoders (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45468";>#45468</a>).</li>
   <li><strong>WideEP / DeepEP v2</strong>: Integrated <strong>DeepEP 
v2</strong> for expert parallelism (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/41183";>#41183</a>), 
with follow-on robustness fixes (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46404";>#46404</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46432";>#46432</a>).</li>
   <li><strong>Rust frontend matures further</strong>: Added API-key 
authentication (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44321";>#44321</a>), 
CORS (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45753";>#45753</a>), 
<code>/tokenize</code> + <code>/detokenize</code> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44222";>#44222</a>), 
<code>/pause</code> <code>/resume</code> <code>/is_paused</code> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44499";>#44499</a>), 
<code>/abort_requests</code> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44382";>#44382</a>), 
<code>/get_world_size</code> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44801";>#44801</a>), 
<code>thinking_token_budget</code> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46137";>#46137</a>), 
a Python bridge for Rust tool parsers (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44624";>#44624</a>),
  and many new parsers and validation paths.</li>
   <li><strong>Device selection change</strong>: vLLM no longer sets 
<code>CUDA_VISIBLE_DEVICES</code> internally; a new <code>device_ids</code> 
argument is provided instead (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45026";>#45026</a>). 
On ROCm, a deprecation window for <code>CUDA_VISIBLE_DEVICES</code> has begun 
(<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46636";>#46636</a>).</li>
   </ul>
   <h3>Model Support</h3>
   <ul>
   <li><strong>New models</strong>: MiniMax-M3 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45381";>#45381</a>), 
DiffusionGemma (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45163";>#45163</a>) + 
Gemma Diffusion on CPU (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45690";>#45690</a>), 
Hierarchical Reasoning Model — Text / HrmTextForCausalLM (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43098";>#43098</a>), 
OpenMOSS (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44124";>#44124</a>).</li>
   <li><strong>Gemma 4</strong>: Unified FlashAttention (FA4) across all layers 
+ <code>mm_prefix</code> support (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42175";>#42175</a>); 
many parser/serving fixes — forced-JSON skip for required/named tool choice (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45795";>#45795</a>), 
parsing with thinking disabled (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45832";>#45832</a>), 
streaming reasoning-state init (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45852";>#45852</a>), 
reasoning rendering on assistant turns (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45867";>#45867</a>), 
offline-parser truncation/token-leak fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45553";>#45553</a>); 
legacy Gemma4 parsers replaced with an engine-based implementation (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45588";>#45588</a>).</li>
   <li><strong>DeepSeek-V4</strong>: OOM fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44914";>#44914</a>), 
MTP projection prefixing (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44821";>#44821</a>), 
supported KV-cache dtypes (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44892";>#44892</a>).</li>
   <li><strong>Qwen / multimodal</strong>: Qwen3-VL video loader (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44412";>#44412</a>), 
Qwen2-VL/Qwen2.5-VL processor-mapped video loader (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45555";>#45555</a>), 
Qwen3-VL multi-video processing optimization (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46026";>#46026</a>) 
and multi-video crash fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46305";>#46305</a>), 
Qwen3-Omni VIT cu_seqlens device fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44264";>#44264</a>), 
fused qk-rmsnorm-rope-gate for Qwen3.5 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44176";>#44176</a>), 
Qwen3.5 EP weight-loading fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45002";>#45002</a>).</li>
   <li><strong>ViT full CUDA graph</strong>: GLM-4.1V (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/40576";>#40576</a>), 
DeepSeek-OCR dual-path (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43586";>#43586</a>), 
Kimi-VL (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/41992";>#41992</a>), 
mllama4 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/40660";>#40660</a>), 
Lfm2VL encoder (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44930";>#44930</a>).</li>
   <li><strong>Other model fixes</strong>: Llama4 weight loading (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45047";>#45047</a>) 
and streamed loading to avoid host-OOM (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44645";>#44645</a>), 
MiMo v2.x QKV TP sharding + FP4 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45200";>#45200</a>), 
ColQwen3.5 retrieval correctness (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46108";>#46108</a>), 
EXAONE-4.5 vision encoder (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45073";>#45073</a>), 
MiDashengLM TP&gt;1 audio-encoder crash (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44408";>#44408</a>), 
MiniCPM-o/V device-placement and image-size fixes (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43844";>#43844</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42332";>#42332</a>, 
<a href="https://redirect.github.com/vllm-project/vll
 m/issues/44980">#44980</a>, <a 
href="https://redirect.github.com/vllm-project/vllm/issues/45244";>#45244</a>), 
Cohere2 MoE weight loading + parser (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44747";>#44747</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44907";>#44907</a>), 
Nemotron V3 reasoning-as-content (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/39091";>#39091</a>), 
ColBERT AutoWeightsLoader + query/document embedding io processor (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44999";>#44999</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45210";>#45210</a>).</li>
   <li><strong>Kernels</strong>: GLM-5 TRT-LLM ragged MLA prefill dimensions 
(<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43525";>#43525</a>), 
GLM-5 router GEMM (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46385";>#46385</a>).</li>
   </ul>
   <h3>Engine Core</h3>
   <ul>
   <li><strong>Model Runner V2</strong>: Quantized models by default (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44446";>#44446</a>), 
GraniteMoE default (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45461";>#45461</a>), 
Qwen/DSv2 MoE migration (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42667";>#42667</a>), 
DFlash (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44586";>#44586</a>), 
simplified async output handling (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45442";>#45442</a>), 
attention-group split on <code>num_heads_q</code> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45564";>#45564</a>), 
LoRA warmup fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/35536";>#35536</a>), 
more accurate FP32 Gumbel sampling (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45996";>#45996</a>), 
<code>min_tokens</code> off-by-one fix in the V2 GPU sampler (<a 
href="https://re
 direct.github.com/vllm-project/vllm/issues/46243">#46243</a>), plus assorted 
model/config compatibility fixes (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45868";>#45868</a>).</li>
   <li><strong>Speculative decoding</strong>: Dynamic SD (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/32374";>#32374</a>); 
DFlash with FlashInfer (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43081";>#43081</a>), 
mixed KV page sizes (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45181";>#45181</a>), 
and Qwen3Next targets (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45319";>#45319</a>); 
EAGLE3 support for Qwen3 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43132";>#43132</a>); 
reduced TP communication for large-vocab drafts (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/39419";>#39419</a>); 
race fix in async accepted counts (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45100";>#45100</a>); 
EAGLE multimodal encoder cache fixes (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46315";>#46315</a>).</li>
   <li><strong>KV cache &amp; scheduler</strong>: KV-cache watermark to reduce 
preemptions (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44594";>#44594</a>), 
two-phase allocation for cross-group prefix-cache hits (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44409";>#44409</a>), 
Marconi-style admission policy for hybrid cache (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/37898";>#37898</a>), 
prefix-cache retention for Mamba/linear attention (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45845";>#45845</a>), 
DS Mamba tail-copy for MTP align mode (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45473";>#45473</a>), 
reduced scheduler copy overhead (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45840";>#45840</a>).</li>
   <li><strong>Attention</strong>: Re-enabled cross-layer KV cache layout for 
MLA via stride-aware kernels (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45111";>#45111</a>), 
MLA prefill FA4 fp8 output (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43050";>#43050</a>), 
FlexAttention custom mask mods made fully cudagraphable (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45232";>#45232</a>), 
triton diff-kv backend for MiMo (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/41797";>#41797</a>), 
FlashMLA sparse accuracy fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/36616";>#36616</a>).</li>
   <li><strong>Weight loading &amp; core</strong>: fastsafetensors 
<code>ParallelLoader</code> for weight loading (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/40183";>#40183</a>), 
release of cached device memory under pressure on UMA GPUs (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45179";>#45179</a>), 
structured outputs for beam search (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/35022";>#35022</a>), 
<code>device_ids</code> arg / no internal <code>CUDA_VISIBLE_DEVICES</code> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45026";>#45026</a>), 
graceful fallback when <code>numactl --membind</code> is blocked (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45438";>#45438</a>), 
config-class registration before tokenizer init (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/40299";>#40299</a>), 
async scheduling with prompt embeds for multimodal models (<a 
href="https://redirect.github.com/vllm-pr
 oject/vllm/issues/45673">#45673</a>).</li>
   </ul>
   <h3>Large Scale Serving &amp; Distributed</h3>
   <ul>
   <li><strong>Expert parallel</strong>: DeepEP v2 integration (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/41183";>#41183</a>) 
with token-bound and topk-index fixes (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46404";>#46404</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46432";>#46432</a>); 
NIXL EP — DBO with NIXL EP (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45275";>#45275</a>), 
top-k index dtype query (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45298";>#45298</a>), 
NVFP4 post-receive quantization skip (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45606";>#45606</a>), 
elastic-EP communicator (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45013";>#45013</a>); 
reject NCCL-based EPLB with async EPLB (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44978";>#44978</a>).</li>
   <li><strong>KV connectors / disaggregated serving</strong>: KV push from 
prefill to decode via NIXL (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/35264";>#35264</a>); 
per-region KV transfer classification for mixed full-attn + MLA groups (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44583";>#44583</a>); 
Mooncake pipeline-parallel PD support (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44528";>#44528</a>), 
async lookup (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45659";>#45659</a>), 
compact chunk-hash zero-copy lookup (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45969";>#45969</a>), 
SWA-block skipping (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45444";>#45444</a>); 
P/D fixes with DP supervisor (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46628";>#46628</a>) 
and DSV4 disaggregation (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45831";>#45831</a>); 
re
 moved <code>P2pNcclConnector</code> (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44854";>#44854</a>).</li>
   <li><strong>KV offloading</strong>: Multi-tier async batched lookup (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44193";>#44193</a>), 
packed HMA KV-cache layout (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46205";>#46205</a>, 
gated <a 
href="https://redirect.github.com/vllm-project/vllm/issues/46252";>#46252</a>), 
parallel-agnostic fs-tier cache (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44733";>#44733</a>), 
offloading-manager stats (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/35669";>#35669</a>) 
and labeled/CPU-usage metrics (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45957";>#45957</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45737";>#45737</a>), 
self-describing KV events (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43468";>#43468</a>), 
non-blocking idle flush (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45595";>#45595</a>), 
and numerous co
 rrectness/race fixes (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44784";>#44784</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45823";>#45823</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46231";>#46231</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46278";>#46278</a>).</li>
   <li><strong>Distributed core</strong>: Prefill step cadence for better 
non-PD DP balancing (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44558";>#44558</a>), 
KV-event map encoding (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42892";>#42892</a>), 
one-shot fused all-reduce PDL NaN fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45448";>#45448</a>).</li>
   </ul>
   <h3>Hardware &amp; Performance</h3>
   <ul>
   <li><strong>NVIDIA / kernels</strong>: SM90 CUTLASS FP8 mm odd-M support via 
swap_ab (180–290% kernel speedup) (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44572";>#44572</a>), 
tuned <code>fused_moe</code> FP8 for Qwen3-Next-80B on H100 (+25%) (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44830";>#44830</a>), 
native DSA indexer decode on SM100 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45322";>#45322</a>), 
cluster-cooperative topK for DeepSeek low-latency (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43008";>#43008</a>), 
PDL support for DeepGEMM (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46006";>#46006</a>), 
FlashInfer cutedsl NVFP4 GEMM (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42235";>#42235</a>) 
and cute-dsl MXFP8 linear kernel (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46393";>#46393</a>), 
new Helion kernels for FP8/RMSNorm quant (<a href="https://red
 irect.github.com/vllm-project/vllm/issues/36902">#36902</a>, <a 
href="https://redirect.github.com/vllm-project/vllm/issues/33790";>#33790</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/36895";>#36895</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/34432";>#34432</a>).</li>
   <li><strong>torch stable ABI</strong>: Continued (and completed) migration 
of kernels to the libtorch stable ABI — MoE [10c/n] (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44565";>#44565</a>), 
Marlin [11a/n] (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45176";>#45176</a>), 
Machete [11b/n] (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45304";>#45304</a>), 
final <code>_C</code> library migration [12/n] (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45415";>#45415</a>).</li>
   <li><strong>AMD ROCm</strong>: Torch 2.11 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45362";>#45362</a>); 
fused AR + RMSNorm + per-group FP8 quant (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42864";>#42864</a>), 
fused softplus-sqrt-topk MoE router under AITER (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44945";>#44945</a>), 
DSv4 flash-decode split-K kernel (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44899";>#44899</a>) 
and inverse-RoPE fusion (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45103";>#45103</a>), 
W4A16 FlyDSL MoE (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44400";>#44400</a>), 
A8W4 MoE CDNA4 swizzle gate for gpt-oss (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44804";>#44804</a>); 
deprecation window begun for <code>CUDA_VISIBLE_DEVICES</code> on ROCm (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46636";>#46636</a>).</li>
   <li><strong>Intel XPU</strong>: Sequence-parallel support (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/38608";>#38608</a>), 
torch-xpu 2.12 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42262";>#42262</a>), 
vllm-xpu-kernels v0.1.10 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/40367";>#40367</a>), 
W4A16 int4 group_size=32 MoE (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45136";>#45136</a>), 
DeepSeek-V4 attention/MoE paths (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44144";>#44144</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44517";>#44517</a>, 
<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45240";>#45240</a>), 
top-p sampling correctness fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44470";>#44470</a>).</li>
   <li><strong>CPU &amp; other architectures</strong>: 2.5× faster ASR CPU 
preprocessing via multi-threading (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44612";>#44612</a>), 
CPU W4A16 INT4 MoE (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43409";>#43409</a>), 
cgroup memory-limit-aware KV cache sizing (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45086";>#45086</a>), 
RISC-V oneDNN W8A8 INT8 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44478";>#44478</a>) 
and RVV micro-GEMM for WNA16 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44324";>#44324</a>), 
pinned memory for WSL2 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/41496";>#41496</a>), 
ZenCPU runtime logging (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42726";>#42726</a>).</li>
   <li><strong>TPU</strong>: tpu-inference upgraded to v0.22.1 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45793";>#45793</a>).</li>
   <li><strong>Misc perf</strong>: <code>VLLM_TRITON_FORCE_FIRST_CONFIG</code> 
to skip Triton autotuning (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42425";>#42425</a>), 
Triton recompile detection (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45631";>#45631</a>), 
fused multi-group block-table staged writes (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44944";>#44944</a>).</li>
   </ul>
   <h3>Quantization</h3>
   <ul>
   <li><strong>Online &amp; mixed-precision</strong>: Online FP8 
per-token-per-channel (PTPC) quantization (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/44132";>#44132</a>); 
<code>modelopt_mixed</code> support extended to Ampere/SM80-86 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45306";>#45306</a>) 
and Turing/SM75 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45375";>#45375</a>).</li>
   <li><strong>FP4 / MXFP</strong>: FlashInfer cutedsl NVFP4 GEMM backend (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/42235";>#42235</a>) 
and cute-dsl MXFP8 linear kernel (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46393";>#46393</a>), 
MXFP4 W4A4 MoE CUTLASS E8M0 scale fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/43557";>#43557</a>), 
SwiGLU clamp wired for NVFP4 MoE on non-Blackwell (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/45836";>#45836</a>), 
<code>flashinfer_cutlass</code> allowed as a clamped NVFP4 MoE backend (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46492";>#46492</a>), 
NVFP4/OCP MX MoE emulation fix (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46254";>#46254</a>), 
FP8 MoE re-enabled on NVIDIA Thor (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46339";>#46339</a>).</li>
   </ul>
   <!-- raw HTML omitted -->
   </blockquote>
   <p>... (truncated)</p>
   </details>
   <details>
   <summary>Commits</summary>
   <ul>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/ee0da84ab9e04ac7610e28580af62c365e898389";><code>ee0da84</code></a>
 [KV-Offloading] Fix tensors_per_block stride (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46888";>#46888</a>)</li>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/217c64a976869883fcf0c52a8cf8bc3c954d285b";><code>217c64a</code></a>
 [CI] Raise gsm8k startup timeout for MoE Refactor Qwen3 NVFP4 configs (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46882";>#46882</a>)</li>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/cfe8a4d06326b49aceed56f8faee85cd831c42d8";><code>cfe8a4d</code></a>
 [CI] Raise gsm8k startup timeout for Qwen3 NVFP4 trtllm configs (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46881";>#46881</a>)</li>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/6d37570a1c6d8f96e9f2b2b72594ab9fd3ae0993";><code>6d37570</code></a>
 Fix P/D with DP Supervisor (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46628";>#46628</a>)</li>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/f85a9f112ab6dbb1ad006c7caf93aa76e07c6ba3";><code>f85a9f1</code></a>
 [Bugfix] FLASHINFER_MLA_SPARSE_SM120 compatibility with GLM-5 NVFP4 (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46506";>#46506</a>)</li>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/836b5acb1b8e90878d9845bb7da4b7c75f043289";><code>836b5ac</code></a>
 [ROCm] Begin Deprecation Window for CUDA_VISIBLE_DEVICES on ROCm (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46636";>#46636</a>)</li>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/b36db10f27ef2ca81b4cc270ed6bed20f1d74b82";><code>b36db10</code></a>
 [KV Offload] Gate packed HMA KV cache on cross-layer config (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46252";>#46252</a>)</li>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/b70c13ea47a85964fe162337b16c17410a131f5e";><code>b70c13e</code></a>
 [Bug] Fix `IndentationError: expected an indented block after 'with' 
statemen...</li>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/6829a6d55f6b0eedd142150a354b68882ad25693";><code>6829a6d</code></a>
 [Bugfix] Re-enable FP8 MoE on NVIDIA Thor (<a 
href="https://redirect.github.com/vllm-project/vllm/issues/46339";>#46339</a>)</li>
   <li><a 
href="https://github.com/vllm-project/vllm/commit/6ed56e04ff8aacffc60276817e3a49aa2ee23899";><code>6ed56e0</code></a>
 [Bugfix] Fix illegal memory access from a forward during a partial wake_up 
(#...</li>
   <li>Additional commits viewable in <a 
href="https://github.com/vllm-project/vllm/compare/v0.10.1.1...v0.24.0";>compare 
view</a></li>
   </ul>
   </details>
   <br />
   
   
   [![Dependabot compatibility 
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=vllm&package-manager=pip&previous-version=0.10.1.1&new-version=0.24.0)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)
   
   Dependabot will resolve any conflicts with this PR as long as you don't 
alter it yourself. You can also trigger a rebase manually by commenting 
`@dependabot rebase`.
   
   [//]: # (dependabot-automerge-start)
   [//]: # (dependabot-automerge-end)
   
   ---
   
   <details>
   <summary>Dependabot commands and options</summary>
   <br />
   
   You can trigger Dependabot actions by commenting on this PR:
   - `@dependabot rebase` will rebase this PR
   - `@dependabot recreate` will recreate this PR, overwriting any edits that 
have been made to it
   - `@dependabot show <dependency name> ignore conditions` will show all of 
the ignore conditions of the specified dependency
   - `@dependabot ignore this major version` will close this PR and stop 
Dependabot creating any more for this major version (unless you reopen the PR 
or upgrade to it yourself)
   - `@dependabot ignore this minor version` will close this PR and stop 
Dependabot creating any more for this minor version (unless you reopen the PR 
or upgrade to it yourself)
   - `@dependabot ignore this dependency` will close this PR and stop 
Dependabot creating any more for this dependency (unless you reopen the PR or 
upgrade to it yourself)
   You can disable automated security fix PRs for this repo from the [Security 
Alerts page](https://github.com/apache/beam/network/alerts).
   
   </details>


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to