solrbot opened a new pull request, #4842: URL: https://github.com/apache/solr/pull/4842
> ℹ️ **Note** > > This PR body was truncated due to platform limits. This PR contains the following updates: | Package | Type | Update | Change | Pending | |---|---|---|---|---| | [org.threeten:threetenbp](https://www.threeten.org/threetenbp) ([source](https://redirect.github.com/ThreeTen/threetenbp)) | dependencies | patch | `1.7.3` → `1.7.4` | | | io.swagger.core.v3.swagger-gradle-plugin | plugin | patch | `2.2.52` → `2.2.54` | `2.2.55` | | [io.swagger.core.v3:swagger-jaxrs2-jakarta](https://redirect.github.com/swagger-api/swagger-core) | dependencies | patch | `2.2.52` → `2.2.54` | `2.2.55` | | [io.swagger.core.v3:swagger-annotations-jakarta](https://redirect.github.com/swagger-api/swagger-core) | dependencies | patch | `2.2.52` → `2.2.54` | `2.2.55` | | [com.github.spotbugs:spotbugs-annotations](https://spotbugs.github.io/) ([source](https://redirect.github.com/spotbugs/spotbugs)) | dependencies | patch | `4.10.2` → `4.10.4` | | | org.openapi.generator | plugin | minor | `7.23.0` → `7.25.0` | | | [com.microsoft.onnxruntime:onnxruntime](https://microsoft.github.io/onnxruntime/) ([source](https://redirect.github.com/microsoft/onnxruntime)) | dependencies | minor | `1.26.0` → `1.29.0` | | | [io.nlopez.compose.rules:ktlint](https://redirect.github.com/mrmans0n/compose-rules) | dependencies | patch | `0.6.2` → `0.6.4` | | | [no.nav.security:mock-oauth2-server](https://redirect.github.com/navikt/mock-oauth2-server) | dependencies | patch | `5.0.1` → `5.0.2` | | | net.ltgt.errorprone | plugin | patch | `5.1.0` → `5.1.1` | | | [dev.logchange](https://redirect.github.com/logchange/logchange) | plugin | patch | `1.19.15` → `1.19.16` | | | nl.littlerobots.version-catalog-update | plugin | patch | `1.1.0` → `1.1.1` | | | [dev.langchain4j:langchain4j-bom](https://redirect.github.com/langchain4j/langchain4j/tree/main/langchain4j-bom) ([source](https://redirect.github.com/langchain4j/langchain4j/tree/HEAD/langchain4j-bom)) | dependencies | minor | `1.17.0` → `1.19.0` | | | [joda-time:joda-time](https://www.joda.org/joda-time/) ([source](https://redirect.github.com/JodaOrg/joda-time)) | dependencies | patch | `2.14.2` → `2.14.3` | | | [org.jctools:jctools-core](https://redirect.github.com/JCTools) ([source](https://redirect.github.com/JCTools/JCTools)) | dependencies | patch | `4.0.6` → `4.0.7` | | | [com.google.guava:guava](https://redirect.github.com/google/guava) | dependencies | minor | `33.6.0-jre` → `33.7.1-jre` | | | [org.eclipse.jgit:org.eclipse.jgit](https://eclipse.gerrithub.io/admin/repos/eclipse-jgit/jgit) | dependencies | patch | `7.7.0.202606012155-r` → `7.7.1.202607240634-r` | | | com.diffplug.spotless | plugin | minor | `8.7.0` → `8.10.0` | `8.10.1` | | [com.nvidia.cuvs:cuvs-java](https://rapids.ai) ([source](https://redirect.github.com/rapidsai/cuvs)) | dependencies | minor | `26.06.0` → `26.08.1` | | | [commons-codec:commons-codec](https://commons.apache.org/proper/commons-codec/) ([source](https://redirect.github.com/apache/commons-codec)) | dependencies | patch | `1.22.0` → `1.22.1` | | | [org.checkerframework:checker-qual](https://checkerframework.org/) ([source](https://redirect.github.com/typetools/checker-framework)) | dependencies | patch | `4.2.0` → `4.2.2` | | | [com.carrotsearch:hppc](https://redirect.github.com/carrotsearch/hppc) | dependencies | minor | `0.10.0` → `0.11.1` | | | [org.bouncycastle:bcprov-jdk18on](https://www.bouncycastle.org/download/bouncy-castle-java/) ([source](https://redirect.github.com/bcgit/bc-java)) | dependencies | minor | `1.84` → `1.85.2` | | | [org.bouncycastle:bcpkix-jdk18on](https://www.bouncycastle.org/download/bouncy-castle-java/) ([source](https://redirect.github.com/bcgit/bc-java)) | dependencies | minor | `1.84` → `1.85` | | | com.github.ben-manes.versions | plugin | minor | `0.54.0` → `0.61.0` | | | [org.apache.tika:tika-core](https://tika.apache.org/) ([source](https://redirect.github.com/apache/tika)) | dependencies | patch | `3.3.1` → `3.3.2` | | | [org.apache.opennlp:opennlp-tools](https://www.apache.org/) ([source](https://redirect.github.com/apache/opennlp)) | dependencies | patch | `2.5.10` → `2.5.11` | | | [org.apache.opennlp:opennlp-dl](https://www.apache.org/) ([source](https://redirect.github.com/apache/opennlp)) | dependencies | patch | `2.5.10` → `2.5.11` | | | [org.apache.commons:commons-collections4](https://commons.apache.org/proper/commons-collections/) ([source](https://gitbox.apache.org/repos/asf?p=commons-collections.git)) | dependencies | minor | `4.5.0` → `4.6.0` | | | [com.adobe.testing:s3mock-testcontainers](https://redirect.github.com/adobe/S3Mock) | dependencies | minor | `5.1.0` → `5.2.0` | | --- ### Release Notes <details> <summary>ThreeTen/threetenbp (org.threeten:threetenbp)</summary> ### [`v1.7.4`](https://redirect.github.com/ThreeTen/threetenbp/releases/tag/v1.7.4) See the [change notes](https://www.threeten.org/threetenbp/changes-report.html) for more information. </details> <details> <summary>swagger-api/swagger-core (io.swagger.core.v3:swagger-jaxrs2-jakarta)</summary> ### [`v2.2.54`](https://redirect.github.com/swagger-api/swagger-core/blob/HEAD/CHANGELOG.md#2254---2026-08-18) ##### Fixed - Java 8 date/time types (`OffsetTime`, `Duration`, `LocalTime`) now map by default to the correct OpenAPI Formats Registry strings (`"time"`, `"duration"`, `"time-local"`) instead of an unusable expanded object. ([#​5172](https://redirect.github.com/swagger-api/swagger-core/issues/5172)) - `LocalDateTime` deserialization from an existing OpenAPI spec now correctly round-trips through the new `TimeSchema`/`DurationSchema`/`DateTimeLocalSchema`/ `TimeLocalSchema` classes instead of falling back to a generic `StringSchema`. ##### Added - `PrimitiveType.enableJava8Formats()` — opt-in to map `LocalDateTime` to the registry-compliant `"date-time-local"` format (default remains `"date-time"` for backward compatibility). ##### Deprecated - `PrimitiveType.enablePartialTime()` — prefer the new default `"time-local"` mapping for `LocalTime`; kept for callers who specifically need the non-registry `"partial-time"` format. ### [`v2.2.53`](https://redirect.github.com/swagger-api/swagger-core/releases/tag/v2.2.53): Swagger-core 2.2.53 released! - chore: update Jackson to 2.22.1 ([#​5258](https://redirect.github.com/swagger-api/swagger-core/issues/5258)) - refactor: replace `writer(new DefaultPrettyPrinter())` with `writerWithDefaultPrettyPrinter()` ([#​5252](https://redirect.github.com/swagger-api/swagger-core/issues/5252)) - fix: Stabilize CI Maven and Gradle builds ([#​5238](https://redirect.github.com/swagger-api/swagger-core/issues/5238)) - test: remove system.out.println from tests ([#​5236](https://redirect.github.com/swagger-api/swagger-core/issues/5236)) - chore: bump dependencies ([#​5229](https://redirect.github.com/swagger-api/swagger-core/issues/5229)) - refactor: simplify type handling in ModelDeserializer ([#​5227](https://redirect.github.com/swagger-api/swagger-core/issues/5227)) - Revert "hotfix: temporarily allow Release workflow to skip mvn deploy to recover 2.2.52 ([#​5220](https://redirect.github.com/swagger-api/swagger-core/issues/5220))" ([#​5223](https://redirect.github.com/swagger-api/swagger-core/issues/5223)) - chore(deps-dev): bump org.codehaus.groovy:groovy from 3.0.23 to 3.0.25 ([#​5216](https://redirect.github.com/swagger-api/swagger-core/issues/5216)) - chore(deps): bump commons-cli:commons-cli from 1.9.0 to 1.11.0 ([#​5209](https://redirect.github.com/swagger-api/swagger-core/issues/5209)) - chore(deps-dev): bump commons-codec:commons-codec from 1.17.2 to 1.22.0 ([#​5208](https://redirect.github.com/swagger-api/swagger-core/issues/5208)) - fix: emit $ref for array items when cycle guard suppresses implementation processing ([#​5205](https://redirect.github.com/swagger-api/swagger-core/issues/5205)) - Restore inner property name from map key in handleUnwrapped ([#​5193](https://redirect.github.com/swagger-api/swagger-core/issues/5193)) - Honor PropertyNamingStrategy for get/is-prefixed property names ([#​5192](https://redirect.github.com/swagger-api/swagger-core/issues/5192)) - fix: let explicit [@​Schema](https://redirect.github.com/Schema)(format) override type-derived format ([#​5185](https://redirect.github.com/swagger-api/swagger-core/issues/5185)) ([#​5186](https://redirect.github.com/swagger-api/swagger-core/issues/5186)) - docs: update format of javadoc to produce a functional link ([#​5182](https://redirect.github.com/swagger-api/swagger-core/issues/5182)) - fix: exclude overridable annotation values when parsing composed annotations ([#​5179](https://redirect.github.com/swagger-api/swagger-core/issues/5179)) - fix: negative and positive validation annotations uses relevant OAS 3.1 syntax ( [#​5170](https://redirect.github.com/swagger-api/swagger-core/issues/5170)) ([#​5171](https://redirect.github.com/swagger-api/swagger-core/issues/5171)) </details> <details> <summary>spotbugs/spotbugs (com.github.spotbugs:spotbugs-annotations)</summary> ### [`v4.10.4`](https://redirect.github.com/spotbugs/spotbugs/blob/HEAD/CHANGELOG.md#4104---2026-08-19) [Compare Source](https://redirect.github.com/spotbugs/spotbugs/compare/4.10.3...4.10.4) ##### Fixed - Fix `NN_NAKED_NOTIFY` false negatives when a field read is stored in a local variable before `notify()` or `notifyAll()` ([#​3884](https://redirect.github.com/spotbugs/spotbugs/issues/3884)) - Fix `ASE_ASSERTION_WITH_SIDE_EFFECT` and `ASE_ASSERTION_WITH_SIDE_EFFECT_METHOD` false positives in every method analysed after a method that reads `$assertionsDisabled` without throwing an `AssertionError` ([#​3483](https://redirect.github.com/spotbugs/spotbugs/issues/3483)) - Fix `INT_BAD_COMPARISON_WITH_SIGNED_BYTE` false positive for meaningful comparisons of a signed byte with `127` (`b < 127`, `b >= 127`) ([#​4201](https://redirect.github.com/spotbugs/spotbugs/pull/4201)) - Fix `EI_EXPOSE_REP` false negative for public getters in anonymous classes ([#​4237](https://redirect.github.com/spotbugs/spotbugs/pull/4237)) - Fix missing class report for `java.util.Collections$EmptyNavigableSet` and `java.util.Collections$EmptyNavigableMap` when the result of `Collections.emptySortedSet()`, `emptyNavigableSet()`, `emptySortedMap()` or `emptyNavigableMap()` is stored ([#​4244](https://redirect.github.com/spotbugs/spotbugs/pull/4244)) - Fix `URF_UNREAD_FIELD` false negative for unread instance fields declared in enums ([#​4246](https://redirect.github.com/spotbugs/spotbugs/issues/4246)) - Stop publishing global dependency-management constraints to consumer POMs. ([#​4223](https://redirect.github.com/spotbugs/spotbugs/pull/4223)) ### [`v4.10.3`](https://redirect.github.com/spotbugs/spotbugs/blob/HEAD/CHANGELOG.md#4103---2026-07-12) [Compare Source](https://redirect.github.com/spotbugs/spotbugs/compare/4.10.2...4.10.3) ##### Fixed - Fix `LI_LAZY_INIT_STATIC` false negative when the null guard is written in yoda-style (`null == field`) ([#​4144](https://redirect.github.com/spotbugs/spotbugs/pull/4144)) - Fix `DC_DOUBLECHECK`, `NP_SYNC_AND_NULL_CHECK_FIELD` and `SP_SPIN_ON_FIELD` false negatives when the null guard is written in yoda-style (`null == field`) ([#​4144](https://redirect.github.com/spotbugs/spotbugs/pull/4144)) - Fix message for `UNS_UNSAFE_CALL` bug pattern - Restore CLI plugin loading by fixing DetectorFactoryCollection bootstrap ordering ([#​4191](https://redirect.github.com/spotbugs/spotbugs/pull/4191)) - Fix `UWF_NULL_FIELD` false negative for fields initialized with cast null values ([#​4034](https://redirect.github.com/spotbugs/spotbugs/issues/4034)) - Fix `UMAC_UNCALLABLE_METHOD_OF_ANONYMOUS_CLASS` false positive for methods reached only through method references ([#​4059](https://redirect.github.com/spotbugs/spotbugs/pull/4059)) ##### Changed - Ant `FindBugsViewerTask`: use default look and feel by default. ([#​4165](https://redirect.github.com/spotbugs/spotbugs/pull/4165)) ##### Refactor - Ant `FindBugsViewerTask`: extend `AbstractFindBugsTask` to reduce duplicate code. ([#​4165](https://redirect.github.com/spotbugs/spotbugs/pull/4165)) </details> <details> <summary>microsoft/onnxruntime (com.microsoft.onnxruntime:onnxruntime)</summary> ### [`v1.29.0`](https://redirect.github.com/microsoft/onnxruntime/releases/tag/v1.29.0): ONNX Runtime v1.29.0 #### Announcements & Breaking Changes - onnxruntime-web has announced the deprecation of WebGL and JSEP. The native WebGPU EP is the recommended path going forward. See the deprecation and migration plans for details ([#​29716](https://redirect.github.com/microsoft/onnxruntime/pull/29716), [#​31683](https://redirect.github.com/microsoft/onnxruntime/pull/31683)). - POSIX telemetry is now available on Linux, macOS, Android, and iOS when ONNX Runtime is built with telemetry enabled. It does not change the public ABI, WebAssembly remains telemetry-free, and setting `ORT_DISABLE_TELEMETRY=1` before initialization disables non-Windows telemetry for the process ([#​27379](https://redirect.github.com/microsoft/onnxruntime/pull/27379), [#​29872](https://redirect.github.com/microsoft/onnxruntime/pull/29872)). - The unused internal `onnxruntime/python/tools/tensorrt` dashboard tooling was removed. This does not affect the TensorRT Execution Provider APIs ([#​29395](https://redirect.github.com/microsoft/onnxruntime/pull/29395)). #### Security Fixes ##### Path, bounds, and input validation - Fixed a path traversal vulnerability in TensorRT and NvTensorRTRTX engine refitting by making external-data path validation unconditional ([#​29396](https://redirect.github.com/microsoft/onnxruntime/pull/29396)). - Validated the CPU MoE `k` attribute against the number of experts and fixed a CPU `TensorScatter` security issue ([#​29907](https://redirect.github.com/microsoft/onnxruntime/pull/29907), [#​29916](https://redirect.github.com/microsoft/onnxruntime/pull/29916)). - Added missing rank, shape, and parameter validation for pooling, LSTM and DynamicQuantizeLSTM, Sampling, FeatureVectorizer, SkipLayerNorm, QLinearConv, Whisper decoding, RNN activations, GridSample, contrib `Range`, and `CropAndResize` ([#​29254](https://redirect.github.com/microsoft/onnxruntime/pull/29254), [#​29255](https://redirect.github.com/microsoft/onnxruntime/pull/29255), [#​29265](https://redirect.github.com/microsoft/onnxruntime/pull/29265), [#​29579](https://redirect.github.com/microsoft/onnxruntime/pull/29579), [#​29595](https://redirect.github.com/microsoft/onnxruntime/pull/29595), [#​29605](https://redirect.github.com/microsoft/onnxruntime/pull/29605), [#​29871](https://redirect.github.com/microsoft/onnxruntime/pull/29871), [#​31636](https://redirect.github.com/microsoft/onnxruntime/pull/31636), [#​31671](https://redirect.github.com/microsoft/onnxruntime/pull/31671), [#​31675](https://redirect.github.com/m icrosoft/onnxruntime/pull/31675), [#​31676](https://redirect.github.com/microsoft/onnxruntime/pull/31676), [#​31684](https://redirect.github.com/microsoft/onnxruntime/pull/31684)). - Hardened CUDA indexing and buffer handling in GridSample, transpose, GatherBlockQuantized, InstanceNormalization, LayerNorm/RMSNorm, BeamSearch, DeformConv, AveragePool, and MaxPool ([#​29581](https://redirect.github.com/microsoft/onnxruntime/pull/29581), [#​29631](https://redirect.github.com/microsoft/onnxruntime/pull/29631), [#​29638](https://redirect.github.com/microsoft/onnxruntime/pull/29638), [#​31640](https://redirect.github.com/microsoft/onnxruntime/pull/31640), [#​31642](https://redirect.github.com/microsoft/onnxruntime/pull/31642), [#​31644](https://redirect.github.com/microsoft/onnxruntime/pull/31644), [#​31645](https://redirect.github.com/microsoft/onnxruntime/pull/31645), [#​31647](https://redirect.github.com/microsoft/onnxruntime/pull/31647), [#​31650](https://redirect.github.com/microsoft/onnxruntime/pull/31650)). - Fixed packed sub-byte tensor over-copying in `OrtApi::GetValue` and validated DML constant tensor byte sizes ([#​29157](https://redirect.github.com/microsoft/onnxruntime/pull/29157), [#​31665](https://redirect.github.com/microsoft/onnxruntime/pull/31665)). ##### Supply chain and tooling - Updated npm lockfiles, refreshed the Next.js end-to-end fixture lockfile for security advisories, and upgraded `adm-zip` for `onnxruntime-node` ([#​29827](https://redirect.github.com/microsoft/onnxruntime/pull/29827), [#​29926](https://redirect.github.com/microsoft/onnxruntime/pull/29926), [#​31192](https://redirect.github.com/microsoft/onnxruntime/pull/31192)). #### New Features ##### Core APIs & Runtime - Default intra-op and inter-op thread-pool sizes can now be set with `ORT_INTRA_OP_NUM_THREADS` and `ORT_INTER_OP_NUM_THREADS`. Explicit thread settings still take precedence, and `0` preserves machine-sized defaults ([#​29688](https://redirect.github.com/microsoft/onnxruntime/pull/29688)). - Added weightless-model support for all initializer types, allowed zero-input `EpContext` nodes, and wired maximum-shape inference into workspace estimation ([#​29607](https://redirect.github.com/microsoft/onnxruntime/pull/29607), [#​29799](https://redirect.github.com/microsoft/onnxruntime/pull/29799), [#​31613](https://redirect.github.com/microsoft/onnxruntime/pull/31613)). - Added ONNX-domain support for rotary embedding and a fused `MRotaryEmbedding` contrib operator for Qwen mRoPE variants ([#​29261](https://redirect.github.com/microsoft/onnxruntime/pull/29261), [#​31728](https://redirect.github.com/microsoft/onnxruntime/pull/31728)). - Added multi-shape profiling to `onnxruntime_perf_test` through `--data_shape`, plus verbose graph-transformer tracing and broader inference-session error-path coverage ([#​29555](https://redirect.github.com/microsoft/onnxruntime/pull/29555), [#​29558](https://redirect.github.com/microsoft/onnxruntime/pull/29558), [#​29569](https://redirect.github.com/microsoft/onnxruntime/pull/29569), [#​29571](https://redirect.github.com/microsoft/onnxruntime/pull/29571)). ##### Execution Provider ABI & Plugin EPs - WebGPU now supports device-free compile-only sessions for offline graph transformation ([#​29681](https://redirect.github.com/microsoft/onnxruntime/pull/29681)). - Expanded CUDA plugin EP packaging and testing, including Windows ARM64 package and size options, updated package outputs, and aligned architecture selections across Python, C API, TensorRT, Node.js, and plugin packages ([#​31635](https://redirect.github.com/microsoft/onnxruntime/pull/31635), [#​31722](https://redirect.github.com/microsoft/onnxruntime/pull/31722), [#​31992](https://redirect.github.com/microsoft/onnxruntime/pull/31992)). - Improved plugin lifecycle handling by unloading failed EP library loads and fixing allocator-deleter lifetime ([#​29634](https://redirect.github.com/microsoft/onnxruntime/pull/29634), [#​29770](https://redirect.github.com/microsoft/onnxruntime/pull/29770)). #### Execution Provider Updates ##### NVIDIA CUDA EP ##### Attention and decoding - Added `PagedAttention` with quantized KV cache, XQA decode, MLA, QK-Norm, and head-sink support ([#​29912](https://redirect.github.com/microsoft/onnxruntime/pull/29912)). - Extended quantized KV-cache support with attention sinks, independent and per-channel scales, sliding-window cache support, and a fused K/V dequantization launch ([#​29900](https://redirect.github.com/microsoft/onnxruntime/pull/29900), [#​29904](https://redirect.github.com/microsoft/onnxruntime/pull/29904), [#​31480](https://redirect.github.com/microsoft/onnxruntime/pull/31480)). - Added a cuDNN SDPA decode tier to the standard ONNX `Attention` CUDA kernel and enabled cuDNN SDPA for contrib `Attention` ([#​29715](https://redirect.github.com/microsoft/onnxruntime/pull/29715), [#​29717](https://redirect.github.com/microsoft/onnxruntime/pull/29717)). - Added `attention_bias` support to the GroupQueryAttention unfused path and `state_window` support to LinearAttention and CausalConvWithState for MTP ([#​29525](https://redirect.github.com/microsoft/onnxruntime/pull/29525), [#​31157](https://redirect.github.com/microsoft/onnxruntime/pull/31157)). - Fixed LinearAttention on GPUs with limited shared memory ([#​31982](https://redirect.github.com/microsoft/onnxruntime/pull/31982)). ##### MoE and quantized GEMM - Added NVFP4 QMoE, including native FP4xFP4 prefill on SM120, faster decode GEMV, fused routing/finalization paths, and reduced activation and weight-dequantization overhead ([#​29697](https://redirect.github.com/microsoft/onnxruntime/pull/29697), [#​29824](https://redirect.github.com/microsoft/onnxruntime/pull/29824), [#​29887](https://redirect.github.com/microsoft/onnxruntime/pull/29887), [#​29919](https://redirect.github.com/microsoft/onnxruntime/pull/29919), [#​31156](https://redirect.github.com/microsoft/onnxruntime/pull/31156), [#​31159](https://redirect.github.com/microsoft/onnxruntime/pull/31159), [#​31349](https://redirect.github.com/microsoft/onnxruntime/pull/31349), [#​31479](https://redirect.github.com/microsoft/onnxruntime/pull/31479)). - Added `MatMulBlockQuantizedFp4Weight` and `MatMulBlockQuantizedFp8Weight`, plus block-scaled tensor-core/GEMV decode paths, packed FP4 decode, M-tiling, and folded W8A8 activation QDQ ([#​29818](https://redirect.github.com/microsoft/onnxruntime/pull/29818), [#​29850](https://redirect.github.com/microsoft/onnxruntime/pull/29850), [#​29896](https://redirect.github.com/microsoft/onnxruntime/pull/29896), [#​31155](https://redirect.github.com/microsoft/onnxruntime/pull/31155), [#​31481](https://redirect.github.com/microsoft/onnxruntime/pull/31481)). - Improved MatMulNBits and QMoE robustness and efficiency by optimizing 8-bit dequantization, releasing raw MXFP4 initializers after prepack, and fixing subgraph prepacking and mixed FP8/FP4 build failures ([#​29852](https://redirect.github.com/microsoft/onnxruntime/pull/29852), [#​31141](https://redirect.github.com/microsoft/onnxruntime/pull/31141), [#​31154](https://redirect.github.com/microsoft/onnxruntime/pull/31154), [#​31350](https://redirect.github.com/microsoft/onnxruntime/pull/31350)). ##### Operators and collectives - Added `LinearAttentionGate`, `GatedRMSNorm`, and `GatedAdd` contrib operators ([#​31158](https://redirect.github.com/microsoft/onnxruntime/pull/31158), [#​31835](https://redirect.github.com/microsoft/onnxruntime/pull/31835)). - Added bfloat16 support to `AllReduce`, `AllGather`, and `AllToAll` ([#​31571](https://redirect.github.com/microsoft/onnxruntime/pull/31571)). - Fixed the default zero point in CUDA `GatherBlockQuantized` ([#​31693](https://redirect.github.com/microsoft/onnxruntime/pull/31693)). ##### WebGPU EP - Added DFT, HardSwish, Max/Min, Trilu, GRU, PRelu, MatMulBnb4, and MRotaryEmbedding support ([#​29454](https://redirect.github.com/microsoft/onnxruntime/pull/29454), [#​29587](https://redirect.github.com/microsoft/onnxruntime/pull/29587), [#​29828](https://redirect.github.com/microsoft/onnxruntime/pull/29828), [#​29833](https://redirect.github.com/microsoft/onnxruntime/pull/29833), [#​29840](https://redirect.github.com/microsoft/onnxruntime/pull/29840), [#​29845](https://redirect.github.com/microsoft/onnxruntime/pull/29845), [#​30512](https://redirect.github.com/microsoft/onnxruntime/pull/30512), [#​31976](https://redirect.github.com/microsoft/onnxruntime/pull/31976)). - Expanded integer support across Clip, Reshape, Cast, Add, Tile, Concat, Expand, Gather, CumSum, Max, and Min ([#​29830](https://redirect.github.com/microsoft/onnxruntime/pull/29830), [#​29834](https://redirect.github.com/microsoft/onnxruntime/pull/29834), [#​29835](https://redirect.github.com/microsoft/onnxruntime/pull/29835), [#​29839](https://redirect.github.com/microsoft/onnxruntime/pull/29839), [#​29844](https://redirect.github.com/microsoft/onnxruntime/pull/29844), [#​29847](https://redirect.github.com/microsoft/onnxruntime/pull/29847), [#​29854](https://redirect.github.com/microsoft/onnxruntime/pull/29854), [#​29861](https://redirect.github.com/microsoft/onnxruntime/pull/29861), [#​29897](https://redirect.github.com/microsoft/onnxruntime/pull/29897), [#​29918](https://redirect.github.com/microsoft/onnxruntime/pull/29918), [#​31049](https://redirect.github.com/microsoft/onnxruntime/pull/31049), [#​31702 ](https://redirect.github.com/microsoft/onnxruntime/pull/31702), [#​31709](https://redirect.github.com/microsoft/onnxruntime/pull/31709)). - Added the initial WebGPU PagedAttention implementation and moved Softmax and non-flash Attention to online algorithms ([#​29694](https://redirect.github.com/microsoft/onnxruntime/pull/29694), [#​29724](https://redirect.github.com/microsoft/onnxruntime/pull/29724), [#​31611](https://redirect.github.com/microsoft/onnxruntime/pull/31611)). - Improved MatMulNBits wide-tile accumulation precision ([#​29611](https://redirect.github.com/microsoft/onnxruntime/pull/29611)). - Added and extended Intel subgroup-matrix MatMul/Gemm kernels, including f16, batched-B, and odd-N support ([#​29592](https://redirect.github.com/microsoft/onnxruntime/pull/29592), [#​29749](https://redirect.github.com/microsoft/onnxruntime/pull/29749), [#​29813](https://redirect.github.com/microsoft/onnxruntime/pull/29813), [#​29893](https://redirect.github.com/microsoft/onnxruntime/pull/29893)). - Reduced cold-start and upload overhead with deferred dispatch and staging-buffer improvements; tuned FlashAttention, subgroup Gemm/MatMul, Split-K on Panther Lake, and Xe im2col-matmul ([#​29271](https://redirect.github.com/microsoft/onnxruntime/pull/29271), [#​29505](https://redirect.github.com/microsoft/onnxruntime/pull/29505), [#​29557](https://redirect.github.com/microsoft/onnxruntime/pull/29557), [#​29586](https://redirect.github.com/microsoft/onnxruntime/pull/29586), [#​29846](https://redirect.github.com/microsoft/onnxruntime/pull/29846), [#​30514](https://redirect.github.com/microsoft/onnxruntime/pull/30514)). - Upgraded Dawn and improved reliability by avoiding exceptions in Dawn callbacks, fixing a Linux adapter-failure self-deadlock, correcting Windows x86 transfer callbacks, and fixing TurboQuant batched sequence lengths ([#​29389](https://redirect.github.com/microsoft/onnxruntime/pull/29389), [#​29591](https://redirect.github.com/microsoft/onnxruntime/pull/29591), [#​29625](https://redirect.github.com/microsoft/onnxruntime/pull/29625), [#​29752](https://redirect.github.com/microsoft/onnxruntime/pull/29752), [#​31568](https://redirect.github.com/microsoft/onnxruntime/pull/31568)). ##### WebNN EP - Added uint8-packed 4-bit `GatherBlockQuantized` and `LpNormalization`, reused the shared WASM loader for Blob-backed external data, and fixed per-axis QDQ and MatMulNBits edge cases ([#​29475](https://redirect.github.com/microsoft/onnxruntime/pull/29475), [#​29801](https://redirect.github.com/microsoft/onnxruntime/pull/29801), [#​31151](https://redirect.github.com/microsoft/onnxruntime/pull/31151), [#​31152](https://redirect.github.com/microsoft/onnxruntime/pull/31152), [#​31197](https://redirect.github.com/microsoft/onnxruntime/pull/31197)). ##### OpenVINO / QNN / DML / XNNPACK / TensorRT - OpenVINO fixed float16 constant-output corruption and output-name routing, added dot-separated KV-cache names to the stateful transform, and corrected raw-data-backed float initializer handling ([#​29729](https://redirect.github.com/microsoft/onnxruntime/pull/29729), [#​29882](https://redirect.github.com/microsoft/onnxruntime/pull/29882), [#​29895](https://redirect.github.com/microsoft/onnxruntime/pull/29895), [#​31138](https://redirect.github.com/microsoft/onnxruntime/pull/31138)). - QNN added a reshape handler for split-axis reshapes ([#​29660](https://redirect.github.com/microsoft/onnxruntime/pull/29660)). - DML fixed wide-string handling and made fused graph kernels own their model paths ([#​31656](https://redirect.github.com/microsoft/onnxruntime/pull/31656), [#​31664](https://redirect.github.com/microsoft/onnxruntime/pull/31664)). - XNNPACK now reads dynamic Gemm `M` from the input tensor at compute time ([#​31189](https://redirect.github.com/microsoft/onnxruntime/pull/31189)). - TensorRT deduplicated context-path handling and added a build option for fused-attention cubins ([#​29640](https://redirect.github.com/microsoft/onnxruntime/pull/29640), [#​31632](https://redirect.github.com/microsoft/onnxruntime/pull/31632)). #### CPU & Core Optimizations ##### MLAS - Added Arm64 half-precision GEMM and convolution support through KleidiAI, including FP16 MatMul/Gemm/Conv paths and asymmetric Q4 and SME2 MatMulNBits kernels ([#​28786](https://redirect.github.com/microsoft/onnxruntime/pull/28786), [#​29654](https://redirect.github.com/microsoft/onnxruntime/pull/29654), [#​29709](https://redirect.github.com/microsoft/onnxruntime/pull/29709), [#​29898](https://redirect.github.com/microsoft/onnxruntime/pull/29898)). - Added a RISC-V RVV QNBitGemm backend, an Arm64 NEON fp32 RoPE kernel, portable SVE elementwise kernels with FEXPA exp, and Arm64 UDOT routing for S8U8 QGEMM ([#​29537](https://redirect.github.com/microsoft/onnxruntime/pull/29537), [#​29787](https://redirect.github.com/microsoft/onnxruntime/pull/29787), [#​29836](https://redirect.github.com/microsoft/onnxruntime/pull/29836), [#​31145](https://redirect.github.com/microsoft/onnxruntime/pull/31145)). - Added AVX2/VNNI 2-bit weight kernels and vectorized 2-bit dequantization, and improved fp16 MatMulNBits paths by avoiding fp32 temporaries and writing fp16 output directly across 2-, 4-, and 8-bit paths ([#​29619](https://redirect.github.com/microsoft/onnxruntime/pull/29619), [#​29766](https://redirect.github.com/microsoft/onnxruntime/pull/29766), [#​29791](https://redirect.github.com/microsoft/onnxruntime/pull/29791), [#​29842](https://redirect.github.com/microsoft/onnxruntime/pull/29842), [#​29864](https://redirect.github.com/microsoft/onnxruntime/pull/29864), [#​29901](https://redirect.github.com/microsoft/onnxruntime/pull/29901)). ##### CPU Attention & Kernels - Improved masked Attention performance, enabled CPU FlashAttention on Linux Arm64 through L2-cache detection, and added FP16 GQA with quantized KV cache ([#​29621](https://redirect.github.com/microsoft/onnxruntime/pull/29621), [#​29719](https://redirect.github.com/microsoft/onnxruntime/pull/29719), [#​29825](https://redirect.github.com/microsoft/onnxruntime/pull/29825)). - Added double support to CPU `Cos` and int32 support to CPU `Trilu`, and fixed int8 QLinearSoftmax saturation and AvgPool `ceil_mode`/`count_include_pad` behavior ([#​28975](https://redirect.github.com/microsoft/onnxruntime/pull/28975), [#​29476](https://redirect.github.com/microsoft/onnxruntime/pull/29476), [#​29629](https://redirect.github.com/microsoft/onnxruntime/pull/29629), [#​29728](https://redirect.github.com/microsoft/onnxruntime/pull/29728)). - Fixed `TfIdfVectorizer` weight indexing and skipped MinLength logits-processor construction when `eos_token_id` is negative ([#​29604](https://redirect.github.com/microsoft/onnxruntime/pull/29604), [#​31649](https://redirect.github.com/microsoft/onnxruntime/pull/31649)). - Tightened K/V and cache-indirection shape contracts in CPU Attention and MultiHeadAttention, and fixed LinearAttention output shape inference for grouped-query attention ([#​29892](https://redirect.github.com/microsoft/onnxruntime/pull/29892), [#​31190](https://redirect.github.com/microsoft/onnxruntime/pull/31190), [#​31634](https://redirect.github.com/microsoft/onnxruntime/pull/31634)). ##### Graph, Optimizer, and Runtime - Extended reshape fusion, fixed double recursion in subgraph type/shape inference, and made constant-folding output deterministic ([#​29027](https://redirect.github.com/microsoft/onnxruntime/pull/29027), [#​29617](https://redirect.github.com/microsoft/onnxruntime/pull/29617), [#​29789](https://redirect.github.com/microsoft/onnxruntime/pull/29789)). - Fixed in-memory external initializer loading, memory-pattern allocation stream selection, and a leak in `GetOverridableInitializerNames()` ([#​29349](https://redirect.github.com/microsoft/onnxruntime/pull/29349), [#​29589](https://redirect.github.com/microsoft/onnxruntime/pull/29589), [#​29616](https://redirect.github.com/microsoft/onnxruntime/pull/29616)). - Reduced small MatMul batch allocations and redundant LUT initialization ([#​29085](https://redirect.github.com/microsoft/onnxruntime/pull/29085), [#​29690](https://redirect.github.com/microsoft/onnxruntime/pull/29690)). - Fixed static-initialization-order crashes when importing ONNX Runtime and reduced eager runtime initialization ([#​29880](https://redirect.github.com/microsoft/onnxruntime/pull/29880), [#​31964](https://redirect.github.com/microsoft/onnxruntime/pull/31964)). - Negative CPU `Split` axes now produce an error instead of being accepted ([#​31149](https://redirect.github.com/microsoft/onnxruntime/pull/31149)). #### Web & JavaScript - Added on-demand loading of Blob-backed external data in JSPI builds ([#​29477](https://redirect.github.com/microsoft/onnxruntime/pull/29477)). - Fixed JSEP pooling output shape for `ceil_mode`, allowed DFT to ignore excess input data, and fixed a webpack/Terser release-build crash ([#​29627](https://redirect.github.com/microsoft/onnxruntime/pull/29627), [#​29680](https://redirect.github.com/microsoft/onnxruntime/pull/29680), [#​31652](https://redirect.github.com/microsoft/onnxruntime/pull/31652)). #### Build, Packaging & CI - CUDA package architecture selections are now aligned across plugin EP, Python, C API, TensorRT, and Node.js pipelines. Windows arm64 is only available in CUDA plugin EP ([#​31992](https://redirect.github.com/microsoft/onnxruntime/pull/31992)): | OS | CUDA | CUDA architectures (all in `-real` form) | | ------------- | ---- | ---------------------------------------- | | Linux x64 | 12.8 | 60;70;75;80;86;89;90a;120a | | Linux x64 | 13.x | 75;80;86;89;90a;120a | | Linux aarch64 | 13.x | 89;90a;120a;121a | | Windows x64 | 12.8 | 61;75;86;89;120a | | Windows x64 | 13.x | 75;80;86;89;120a | | Windows arm64 | 13.x | 120a;121a | - Reduced CUDA compilation time and memory usage by splitting generated SM80 MoE, fpA\_intB, and MatMulNBits translation units and adding two-level workspace estimation ([#​29614](https://redirect.github.com/microsoft/onnxruntime/pull/29614), [#​29699](https://redirect.github.com/microsoft/onnxruntime/pull/29699), [#​29811](https://redirect.github.com/microsoft/onnxruntime/pull/29811), [#​31834](https://redirect.github.com/microsoft/onnxruntime/pull/31834), [#​31837](https://redirect.github.com/microsoft/onnxruntime/pull/31837)). - Fixed CUDA 13 plugin and packaging builds on Windows, Linux, and Windows ARM64, including MSVC/TMA compatibility and CI memory limits ([#​31608](https://redirect.github.com/microsoft/onnxruntime/pull/31608), [#​31609](https://redirect.github.com/microsoft/onnxruntime/pull/31609), [#​31615](https://redirect.github.com/microsoft/onnxruntime/pull/31615), [#​31616](https://redirect.github.com/microsoft/onnxruntime/pull/31616), [#​31617](https://redirect.github.com/microsoft/onnxruntime/pull/31617), [#​31622](https://redirect.github.com/microsoft/onnxruntime/pull/31622), [#​31729](https://redirect.github.com/microsoft/onnxruntime/pull/31729), [#​31748](https://redirect.github.com/microsoft/onnxruntime/pull/31748)). - Fixed MLAS AVX2 builds on toolchains without AVX-VNNI assembler support, GCC 15 `-Werror` builds, and an MSVC C1001 issue in the W2 AVX-512-VNNI dispatch path ([#​28767](https://redirect.github.com/microsoft/onnxruntime/pull/28767), [#​29679](https://redirect.github.com/microsoft/onnxruntime/pull/29679), [#​29885](https://redirect.github.com/microsoft/onnxruntime/pull/29885)). - Improved Windows compatibility by skipping DXGI discovery when Win32k system calls are unavailable and delay-loading `shell32` ([#​29755](https://redirect.github.com/microsoft/onnxruntime/pull/29755), [#​30889](https://redirect.github.com/microsoft/onnxruntime/pull/30889)). - Fixed Dawn parallel-build races, and GPU discovery in build/test environments ([#​29858](https://redirect.github.com/microsoft/onnxruntime/pull/29858), [#​29866](https://redirect.github.com/microsoft/onnxruntime/pull/29866)). #### Contributors Thanks to our 63 contributors for this release! [@​adrastogi](https://redirect.github.com/adrastogi), [@​ahsan-ca](https://redirect.github.com/ahsan-ca), [@​AngelGalindo7](https://redirect.github.com/AngelGalindo7), [@​ankitm3k](https://redirect.github.com/ankitm3k), [@​apsonawane](https://redirect.github.com/apsonawane), [@​blazingphoenix7](https://redirect.github.com/blazingphoenix7), [@​bmehta001](https://redirect.github.com/bmehta001), [@​chilo-ms](https://redirect.github.com/chilo-ms), [@​claude](https://redirect.github.com/claude), [@​daijh](https://redirect.github.com/daijh), [@​ducviet00](https://redirect.github.com/ducviet00), [@​edgchen1](https://redirect.github.com/edgchen1), [@​elwhyjay](https://redirect.github.com/elwhyjay), [@​eserscor](https://redirect.github.com/eserscor), [@​GopalakrishnanN](https://redirect.github.com/GopalakrishnanN), [@​guptaishaan](https://redirect.github.com/guptaishaan), [@​hariharans29]( https://redirect.github.com/hariharans29), [@​Honry](https://redirect.github.com/Honry), [@​huningxin](https://redirect.github.com/huningxin), [@​jchen10](https://redirect.github.com/jchen10), [@​jiafatom](https://redirect.github.com/jiafatom), [@​jiangzhuo](https://redirect.github.com/jiangzhuo), [@​Jiawei-Shao](https://redirect.github.com/Jiawei-Shao), [@​JonathanC-ARM](https://redirect.github.com/JonathanC-ARM), [@​justinchuby](https://redirect.github.com/justinchuby), [@​kjg0724](https://redirect.github.com/kjg0724), [@​kunal-vaishnavi](https://redirect.github.com/kunal-vaishnavi), [@​kylo5aby](https://redirect.github.com/kylo5aby), [@​Laan33](https://redirect.github.com/Laan33), [@​martin-klacer-arm](https://redirect.github.com/martin-klacer-arm), [@​mastryukov1990](https://redirect.github.com/mastryukov1990), [@​mcollinswisc](https://redirect.github.com/mcollinswisc), [@​miaobin](ht tps://redirect.github.com/miaobin), [@​mingmingtasd](https://redirect.github.com/mingmingtasd), [@​mirounga](https://redirect.github.com/mirounga), [@​mustjab](https://redirect.github.com/mustjab), [@​n1harika](https://redirect.github.com/n1harika), [@​namgyu-youn](https://redirect.github.com/namgyu-youn), [@​neilmsft](https://redirect.github.com/neilmsft), [@​nenad1002](https://redirect.github.com/nenad1002), [@​nicholascelestin](https://redirect.github.com/nicholascelestin), [@​OscarFree](https://redirect.github.com/OscarFree), [@​prathikr](https://redirect.github.com/prathikr), [@​qjia7](https://redirect.github.com/qjia7), [@​quic-muchhsu](https://redirect.github.com/quic-muchhsu), [@​Sammy-Dabbas](https://redirect.github.com/Sammy-Dabbas), [@​sanaa-hamel-microsoft](https://redirect.github.com/sanaa-hamel-microsoft), [@​shiyi9801](https://redirect.github.com/shiyi9801), [@​skottmckay]( https://redirect.github.com/skottmckay), [@​tairenpiao](https://redirect.github.com/tairenpiao), [@​TedThemistokleous](https://redirect.github.com/TedThemistokleous), [@​the0cp](https://redirect.github.com/the0cp), [@​tianleiwu](https://redirect.github.com/tianleiwu), [@​titaiwangms](https://redirect.github.com/titaiwangms), [@​velonica0](https://redirect.github.com/velonica0), [@​wangw-1991](https://redirect.github.com/wangw-1991), [@​wuisabel-gif](https://redirect.github.com/wuisabel-gif), [@​xadupre](https://redirect.github.com/xadupre), [@​xhcao](https://redirect.github.com/xhcao), [@​xiaofeihan1](https://redirect.github.com/xiaofeihan1), [@​xiaoyu-work](https://redirect.github.com/xiaoyu-work), [@​yen-shi](https://redirect.github.com/yen-shi), [@​zlma7001](https://redirect.github.com/zlma7001) Full Changelog: [v1.28.0...v1.29.0](https://redirect.github.com/microsoft/onnxruntime/compare/v1.28.0...v1.29.0) ### [`v1.28.0`](https://redirect.github.com/microsoft/onnxruntime/releases/tag/v1.28.0): ONNX Runtime v1.28.0 #### Announcements & Breaking Changes - Upgraded to **ONNX 1.22.0** and protobuf 6.33.5 ([#​28754](https://redirect.github.com/microsoft/onnxruntime/pull/28754), [#​29606](https://redirect.github.com/microsoft/onnxruntime/pull/29606), [#​28967](https://redirect.github.com/microsoft/onnxruntime/pull/28967)). Graph optimizer opset version checks were updated accordingly ([#​28966](https://redirect.github.com/microsoft/onnxruntime/pull/28966)). - **cuDNN and cuFFT are now optional at runtime** for the CUDA EP, and `nvrtc` is no longer linked, which significantly reduces the required CUDA redistributable footprint ([#​29252](https://redirect.github.com/microsoft/onnxruntime/pull/29252), [#​29808](https://redirect.github.com/microsoft/onnxruntime/pull/29808), [#​29705](https://redirect.github.com/microsoft/onnxruntime/pull/29705), [#​29620](https://redirect.github.com/microsoft/onnxruntime/pull/29620)). - An **experimental C/C++ API surface** was introduced. `OrtModelPackageApi` now lives in the experimental C API and may change in future releases ([#​28746](https://redirect.github.com/microsoft/onnxruntime/pull/28746), [#​29142](https://redirect.github.com/microsoft/onnxruntime/pull/29142), [#​28990](https://redirect.github.com/microsoft/onnxruntime/pull/28990)). - **Deprecated / removed:** - SkipLayerNorm strict mode is deprecated ([#​29388](https://redirect.github.com/microsoft/onnxruntime/pull/29388)). - The TensorRT fused causal attention kernels were removed from the CUDA EP ([#​29143](https://redirect.github.com/microsoft/onnxruntime/pull/29143)). - The dynamic WGSL generator (duktape/Node) path was removed in favor of the Python `wgsl-gen` implementation ([#​29141](https://redirect.github.com/microsoft/onnxruntime/pull/29141), [#​28355](https://redirect.github.com/microsoft/onnxruntime/pull/28355)). - `CUDA_QUANT_PREPROCESS` is off by default ([#​29687](https://redirect.github.com/microsoft/onnxruntime/pull/29687)). - NPM packages are now published from the CUDA 13 pipeline ([#​28773](https://redirect.github.com/microsoft/onnxruntime/pull/28773)). - The CUDA 12.8 package architecture list was refreshed for this release ([#​29711](https://redirect.github.com/microsoft/onnxruntime/pull/29711)). #### Security Fixes ##### Memory safety & input validation - Hardened the ORT FlatBuffer model loader against malformed buffers, and removed now-redundant table offset validation ([#​28186](https://redirect.github.com/microsoft/onnxruntime/pull/28186), [#​29068](https://redirect.github.com/microsoft/onnxruntime/pull/29068)) - Fixed type confusion in raw-pointer `bind_input` causing an out-of-bounds write ([#​28839](https://redirect.github.com/microsoft/onnxruntime/pull/28839)) - Fixed out-of-bounds pointer in `TensorAt` for sub-byte packed types ([#​28973](https://redirect.github.com/microsoft/onnxruntime/pull/28973)) - Fixed arbitrary memory read, out-of-bounds dereference, and other OOB accesses in kernels ([#​28991](https://redirect.github.com/microsoft/onnxruntime/pull/28991), [#​29011](https://redirect.github.com/microsoft/onnxruntime/pull/29011), [#​29012](https://redirect.github.com/microsoft/onnxruntime/pull/29012), [#​29014](https://redirect.github.com/microsoft/onnxruntime/pull/29014)) - Validated `Col2Im` inputs to prevent heap over-read ([#​28706](https://redirect.github.com/microsoft/onnxruntime/pull/28706)) - Hardened `CropAndResize` against malformed `crop_size` tensors ([#​28766](https://redirect.github.com/microsoft/onnxruntime/pull/28766)) - Validated `BeamSearch` `vocab_size` against logits width ([#​28774](https://redirect.github.com/microsoft/onnxruntime/pull/28774)) - Fixed bounds in `WhisperDecoderSubgraph::CreateInitialFeeds` ([#​29239](https://redirect.github.com/microsoft/onnxruntime/pull/29239)) - Validated `SparseAttention` CSR indices/key lengths and rejected zero-dimension `block_row_indices` ([#​29015](https://redirect.github.com/microsoft/onnxruntime/pull/29015), [#​29242](https://redirect.github.com/microsoft/onnxruntime/pull/29242)) - Clamped derived sequence lengths and KV-cache index in CUDA GroupQueryAttention, and fixed a CPU GQA out-of-bounds read in the past-KV buffer ([#​29240](https://redirect.github.com/microsoft/onnxruntime/pull/29240), [#​29447](https://redirect.github.com/microsoft/onnxruntime/pull/29447)) - Clamped 1D attention `mask_index` to valid bounds ([#​29449](https://redirect.github.com/microsoft/onnxruntime/pull/29449)) - Validated `MaxpoolWithMask` kernel rank against input spatial rank ([#​29253](https://redirect.github.com/microsoft/onnxruntime/pull/29253)) - Rejected CUDA BERT `EmbedLayerNorm`/`SkipLayerNorm` shapes exceeding 32-bit output indexing ([#​29264](https://redirect.github.com/microsoft/onnxruntime/pull/29264)) - Fixed the optional-output guard in `DecoderAttention`/`MultiHeadAttention` shape inference and negative-axis handling in `ExpandDims` shape inference ([#​29268](https://redirect.github.com/microsoft/onnxruntime/pull/29268), [#​29448](https://redirect.github.com/microsoft/onnxruntime/pull/29448)) - Fixed `TreeEnsemble` target id validation and added input validation to `LinearClassifier` ([#​29293](https://redirect.github.com/microsoft/onnxruntime/pull/29293), [#​29060](https://redirect.github.com/microsoft/onnxruntime/pull/29060)) - Fixed `DynamicQuantizeLSTM` zero-point/scale validation typos ([#​29462](https://redirect.github.com/microsoft/onnxruntime/pull/29462)) - Handled non-trivially-copyable types in `Loop`/`Scan` output concatenation ([#​29397](https://redirect.github.com/microsoft/onnxruntime/pull/29397)) - Normalized bool tensor `raw_data` to `{0, 1}` on unpack ([#​29238](https://redirect.github.com/microsoft/onnxruntime/pull/29238)) - Addressed hardening gaps in `Resize`, `PadFusion`, and LoRA handling ([#​28779](https://redirect.github.com/microsoft/onnxruntime/pull/28779), [#​28780](https://redirect.github.com/microsoft/onnxruntime/pull/28780), [#​28801](https://redirect.github.com/microsoft/onnxruntime/pull/28801)) - Fixed unbounded lifetime on `WithOutputTensor` in the Rust bindings ([#​29251](https://redirect.github.com/microsoft/onnxruntime/pull/29251)) ##### Integer overflow & allocation size - Guarded `MlasConvPrepare` working-buffer products and `ConvTranspose` pad computation with SafeInt ([#​29444](https://redirect.github.com/microsoft/onnxruntime/pull/29444), [#​29446](https://redirect.github.com/microsoft/onnxruntime/pull/29446)) - Fixed signed-int overflow in `SamplingState::Init` that could cause a heap buffer overflow ([#​29443](https://redirect.github.com/microsoft/onnxruntime/pull/29443)) - Hardened QMoE against integer overflow and partial K tiles ([#​29067](https://redirect.github.com/microsoft/onnxruntime/pull/29067)) - Validated `B`/scales/zero-points shape in `MatMulNBits::PrePack` ([#​29445](https://redirect.github.com/microsoft/onnxruntime/pull/29445)) - Pre-checked `ConstantOfShape` output size against the input initializer before constant folding ([#​28751](https://redirect.github.com/microsoft/onnxruntime/pull/28751)) - Fixed integer overflow in RKNPU implicit bias allocation ([#​29249](https://redirect.github.com/microsoft/onnxruntime/pull/29249)) - Fixed WebGPU out-of-bounds reads in `Pad` (int64/int32 truncation), `Slice`, and `GatherBlockQuantized` ([#​28721](https://redirect.github.com/microsoft/onnxruntime/pull/28721), [#​28704](https://redirect.github.com/microsoft/onnxruntime/pull/28704), [#​28718](https://redirect.github.com/microsoft/onnxruntime/pull/28718)) ##### Supply chain & tooling - Updated protobuf to mitigate CVE-2026-0994 and bumped ONNX/protobuf to fix additional CVEs ([#​28967](https://redirect.github.com/microsoft/onnxruntime/pull/28967), [#​29606](https://redirect.github.com/microsoft/onnxruntime/pull/29606)) - Avoided shell injection in the training helper and switched Triton compile helpers to `subprocess` ([#​28776](https://redirect.github.com/microsoft/onnxruntime/pull/28776), [#​28775](https://redirect.github.com/microsoft/onnxruntime/pull/28775)) - Validated archive extraction paths in the transformers tooling ([#​28777](https://redirect.github.com/microsoft/onnxruntime/pull/28777)) - Validated and inlined external data in node tensor attributes during session initialization ([#​29250](https://redirect.github.com/microsoft/onnxruntime/pull/29250)) - Enabled Spectre-mitigated MSVC libraries for BinSkim builds ([#​29624](https://redirect.github.com/microsoft/onnxruntime/pull/29624)) - Bumped npm dependencies: `shell-quote`, `esbuild`, `tmp`, `ws`, `protobufjs`, `js-yaml`, `tar`, `markdown-it`, `@babel/core` ([#​29022](https://redirect.github.com/microsoft/onnxruntime/pull/29022), [#​29044](https://redirect.github.com/microsoft/onnxruntime/pull/29044), [#​29055](https://redirect.github.com/microsoft/onnxruntime/pull/29055), [#​29057](https://redirect.github.com/microsoft/onnxruntime/pull/29057), [#​29061](https://redirect.github.com/microsoft/onnxruntime/pull/29061), [#​29062](https://redirect.github.com/microsoft/onnxruntime/pull/29062), [#​29063](https://redirect.github.com/microsoft/onnxruntime/pull/29063), [#​29079](https://redirect.github.com/microsoft/onnxruntime/pull/29079), [#​29090](https://redirect.github.com/microsoft/onnxruntime/pull/29090), [#​29156](https://redirect.github.com/microsoft/onnxruntime/pull/29156)) #### New Features ##### Execution Provider ABI & Plugin EPs - Model Package support Phase 2, plus authoring tools, schema versioning, and folding `external_data` into session options ([#​28271](https://redirect.github.com/microsoft/onnxruntime/pull/28271), [#​28989](https://redirect.github.com/microsoft/onnxruntime/pull/28989), [#​29501](https://redirect.github.com/microsoft/onnxruntime/pull/29501)) - Added an API to select the best compiled-model compatibility info from candidate strings ([#​28387](https://redirect.github.com/microsoft/onnxruntime/pull/28387)) - Added crypto support: applications can supply I/O callbacks to an EP, with callback and fallback helpers ([#​28624](https://redirect.github.com/microsoft/onnxruntime/pull/28624)) - Implemented name-based partitioning with accompanying documentation ([#​28903](https://redirect.github.com/microsoft/onnxruntime/pull/28903)) - Added Linux NPU discovery through sysfs accel devices ([#​28703](https://redirect.github.com/microsoft/onnxruntime/pull/28703)) - Relaxed `CompileModel` validation to accept zero-input `OrtModel` graphs ([#​28771](https://redirect.github.com/microsoft/onnxruntime/pull/28771)) - CUDA plugin EP: user compute stream with CUDA graph, kernel sync stream exposed for scratch allocation, and Windows ARM64 packages ([#​29221](https://redirect.github.com/microsoft/onnxruntime/pull/29221), [#​29244](https://redirect.github.com/microsoft/onnxruntime/pull/29244), [#​28896](https://redirect.github.com/microsoft/onnxruntime/pull/28896), [#​28789](https://redirect.github.com/microsoft/onnxruntime/pull/28789)) - WebGPU plugin EP version bumped to 0.3.0 ([#​29056](https://redirect.github.com/microsoft/onnxruntime/pull/29056)) ##### Core APIs & Runtime - Added `OrtErrorCode` documentation, single-sourced the values so `StatusCode` stays in sync, and added `OrtErrorCode::ORT_DEVICE_RESET` ([#​29018](https://redirect.github.com/microsoft/onnxruntime/pull/29018), [#​29065](https://redirect.github.com/microsoft/onnxruntime/pull/29065), [#​29748](https://redirect.github.com/microsoft/onnxruntime/pull/29748)) - Added memory statistics to profiling output ([#​29058](https://redirect.github.com/microsoft/onnxruntime/pull/29058)) - Added EP version logging on inference failure, in the `EpDeviceUsage` event, and ORT version logging ([#​28794](https://redirect.github.com/microsoft/onnxruntime/pull/28794)) - User-supplied external initializers are now used in place when already on the planned device ([#​29013](https://redirect.github.com/microsoft/onnxruntime/pull/29013)) - `model_external_initializers_file_folder_path` is now honored for file-path model loads ([#​29459](https://redirect.github.com/microsoft/onnxruntime/pull/29459)) - Added a Python API for `HOST_ACCESSIBLE` `OrtValue` allocation ([#​28038](https://redirect.github.com/microsoft/onnxruntime/pull/28038)) ##### Quantization Tooling - Added `CudaQuantizer` to `onnxruntime.quantization` ([#​29509](https://redirect.github.com/microsoft/onnxruntime/pull/29509)) - Registered `Flatten` as a Direct8Bit op in the Python QDQ static quantizer ([#​28340](https://redirect.github.com/microsoft/onnxruntime/pull/28340)) - Skipped `MaxPool` during FP8 static quantization and fixed the FP8 (`FLOAT8E4M3FN`) scale reference distribution ([#​28488](https://redirect.github.com/microsoft/onnxruntime/pull/28488), [#​29350](https://redirect.github.com/microsoft/onnxruntime/pull/29350)) - Added Float16/BFloat16/Float8 support in the `TensorArray` custom op ([#​28335](https://redirect.github.com/microsoft/onnxruntime/pull/28335)) - Clarified CPU parameter recommendations in the quantization docs ([#​28415](https://redirect.github.com/microsoft/onnxruntime/pull/28415)) #### Execution Provider Updates ##### NVIDIA CUDA EP **Attention & LLM decode** - Enabled XQA by default for FP16/BF16 GroupQueryAttention, and extended XQA decode with attention sink, sliding window, and QK-Norm support ([#​29046](https://redirect.github.com/microsoft/onnxruntime/pull/29046), [#​29162](https://redirect.github.com/microsoft/onnxruntime/pull/29162), [#​29177](https://redirect.github.com/microsoft/onnxruntime/pull/29177), [#​29186](https://redirect.github.com/microsoft/onnxruntime/pull/29186)) - Upgraded `cudnn_frontend` to 1.24 and enabled cuDNN SDPA for MHA/GQA ([#​28849](https://redirect.github.com/microsoft/onnxruntime/pull/28849)) - Added decode-optimized LinearAttention (GatedDeltaNet) kernels ([#​28985](https://redirect.github.com/microsoft/onnxruntime/pull/28985)) - Optimized FlashDecode split planning for local-window GQA and fixed Flash/Lean attention split heuristics ([#​29161](https://redirect.github.com/microsoft/onnxruntime/pull/29161), [#​29554](https://redirect.github.com/microsoft/onnxruntime/pull/29554)) - Updated the GroupQueryAttention contrib op documentation ([#​29173](https://redirect.github.com/microsoft/onnxruntime/pull/29173)) **MoE & quantized GEMM** - Prepacked int4/int8 QMoE expert weights in the `PrePack` hook, symmetric with `MatMulNBits`, and fix > ✂ **Note** > > PR body was truncated to here. </details> --- ### Configuration 📅 **Schedule**: (UTC) - Branch creation - "before 9am on the first day of the month" - Automerge - At any time (no schedule defined) 🚦 **Automerge**: Disabled by config. Please merge this manually once you are satisfied. ♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox. 👻 **Immortal**: This PR will be recreated if closed unmerged. Get [config help](https://redirect.github.com/renovatebot/renovate/discussions) if that's undesired. --- - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this box --- This PR has been generated by [Renovate Bot](https://redirect.github.com/solrbot/renovate-github-action) <!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0My4xNzAuMTgiLCJ1cGRhdGVkSW5WZXIiOiI0My4xNzAuMTgiLCJ0YXJnZXRCcmFuY2giOiJtYWluIiwibGFiZWxzIjpbImV4ZW1wdC1zdGFsZSJdfQ==--> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
