This is an automated email from the ASF dual-hosted git repository.

github-merge-queue[bot] pushed a commit to branch 
gh-readonly-queue/main/pr-6910-1b9a8c6ee84f96e9c70712171f99b2a209d8709b
in repository https://gitbox.apache.org/repos/asf/texera.git

commit 589f601d48025ada64836bbeb3adfc183282bb5c
Author: Mend Renovate <[email protected]>
AuthorDate: Tue Jul 28 00:34:24 2026 +0100

    fix(deps, pyamber): update dependency torch to v2.13.0 [security] (#6910)
    
    > ℹ️ **Note**
    >
    > This PR body was truncated due to platform limits.
    
    This PR contains the following updates:
    
    | Package | Change |
    [Age](https://docs.renovatebot.com/merge-confidence/) |
    [Confidence](https://docs.renovatebot.com/merge-confidence/) |
    |---|---|---|---|
    | [torch](https://redirect.github.com/pytorch/pytorch) | `==2.12.1` →
    `==2.13.0` |
    
![age](https://developer.mend.io/api/mc/badges/age/pypi/torch/2.13.0?slim=true)
    |
    
![confidence](https://developer.mend.io/api/mc/badges/confidence/pypi/torch/2.12.1/2.13.0?slim=true)
    |
    | [torch](https://redirect.github.com/pytorch/pytorch) | `==2.12.1+cpu`
    → `==2.13.0` |
    
![age](https://developer.mend.io/api/mc/badges/age/pypi/torch/2.13.0?slim=true)
    |
    
![confidence](https://developer.mend.io/api/mc/badges/confidence/pypi/torch/2.12.1+cpu/2.13.0?slim=true)
    |
    
    ---
    
    > [!WARNING]
    > Some dependencies could not be looked up. Check the [Dependency
    Dashboard](../issues/6912) for more information.
    
    ---
    
    ### PyTorch is vulnerable to memory corruption through its
    torch.jit.script function
    BIT-pytorch-2025-3000 /
    [CVE-2025-3000](https://nvd.nist.gov/vuln/detail/CVE-2025-3000) /
    
[GHSA-rrmf-rvhw-rf47](https://redirect.github.com/advisories/GHSA-rrmf-rvhw-rf47)
    / PYSEC-2025-194
    
    <details>
    <summary>More information</summary>
    
    #### Details
    A vulnerability classified as critical has been found in PyTorch 2.6.0.
    This affects the function torch.jit.script. The manipulation leads to
    memory corruption. It is possible to launch the attack on the local
    host. The exploit has been disclosed to the public and may be used.
    
    #### Severity
    - CVSS Score: 1.9 / 10 (Low)
    - Vector String:
    `CVSS:4.0/AV:L/AC:L/AT:N/PR:L/UI:N/VC:L/VI:L/VA:L/SC:N/SI:N/SA:N/E:P`
    
    #### References
    -
    
[https://nvd.nist.gov/vuln/detail/CVE-2025-3000](https://nvd.nist.gov/vuln/detail/CVE-2025-3000)
    -
    
[https://github.com/pytorch/pytorch/issues/149623](https://redirect.github.com/pytorch/pytorch/issues/149623)
    -
    
[https://github.com/pytorch/pytorch/issues/149623#issue-2935703015](https://redirect.github.com/pytorch/pytorch/issues/149623#issue-2935703015)
    -
    
[https://github.com/pytorch/pytorch/commit/b90c94991cdf8b87c8f7439f79518e0ef2c4ca4f](https://redirect.github.com/pytorch/pytorch/commit/b90c94991cdf8b87c8f7439f79518e0ef2c4ca4f)
    -
    
[https://github.com/pypa/advisory-database/tree/main/vulns/torch/PYSEC-2025-194.yaml](https://redirect.github.com/pypa/advisory-database/tree/main/vulns/torch/PYSEC-2025-194.yaml)
    -
    
[https://github.com/pytorch/pytorch](https://redirect.github.com/pytorch/pytorch)
    - [https://vuldb.com/?ctiid.302049](https://vuldb.com/?ctiid.302049)
    - [https://vuldb.com/?id.302049](https://vuldb.com/?id.302049)
    - [https://vuldb.com/?submit.524197](https://vuldb.com/?submit.524197)
    
    This data is provided by
    [OSV](https://osv.dev/vulnerability/GHSA-rrmf-rvhw-rf47) and the [GitHub
    Advisory Database](https://redirect.github.com/github/advisory-database)
    ([CC-BY
    
4.0](https://redirect.github.com/github/advisory-database/blob/main/LICENSE.md)).
    </details>
    
    ---
    
    ### Release Notes
    
    <details>
    <summary>pytorch/pytorch (torch)</summary>
    
    ###
    
[`v2.13.0`](https://redirect.github.com/pytorch/pytorch/releases/tag/v2.13.0):
    PyTorch 2.13.0 Release
    
    ### PyTorch 2.13.0 Release Notes
    
    - [Highlights](#highlights)
    - [Backwards Incompatible Changes](#backwards-incompatible-changes)
    - [Deprecations](#deprecations)
    - [New Features](#new-features)
    - [Improvements](#improvements)
    - [Bug fixes](#bug-fixes)
    - [Performance](#performance)
    - [Documentation](#documentation)
    - [Developers](#developers)
    
    ### Highlights
    
    <table>
    <tr><td><strong>FlexAttention</strong> lands on Apple Silicon (MPS),
    with up to ~12x speedup over SDPA on sparse patterns, and gains a
    deterministic backward path on CUDA for reproducible gradient
    computation.</td></tr>
    <tr><td><strong>CuTeDSL "Native DSL" backend</strong> gives Inductor a
    second high-performance code path (alongside Triton) for key GPU
    operations, with faster compilation. [Prototype]</td></tr>
    <tr><td><strong><code>nn.LinearCrossEntropyLoss</code></strong> combines
    the final prediction and loss computation to cut peak GPU memory by up
    to 4x for large-vocabulary language model training.</td></tr>
    <tr><td><strong>torchcomms</strong>, a new communications backend for
    PyTorch Distributed, improves fault tolerance, scalability, and
    debuggability for large-cluster training.</td></tr>
    <tr><td><strong>FSDP2</strong> now overlaps reduce-scatter and
    all-gather communications via a dedicated process group (opt-in),
    increasing distributed training throughput.</td></tr>
    <tr><td><strong>Python 3.15 wheel support</strong> for PyTorch on Linux
    via the pytorch repository index, including builds compatible with
    free-threaded 3.15t.</td></tr>
    <tr><td><strong>Broader platform support</strong>: ROCm gains AOTriton
    0.12b with native HIP CMake, Arm adds Armv9-A <code>torch.compile</code>
    targeting, and Intel XPU exposes new device telemetry APIs.</td></tr>
    </table>
    
    For more details about these highlighted features, you can look at the
    release blogpost. Below are the full release notes for this release.
    
    ### Tracked Regressions
    
    ##### ROCm wheels break `torch.compile` on CPU in environments without a
    GPU
    
    Running a `torch==2.13.0+rocm7.2` wheel in an environment where no GPU
    is available (`torch.cuda.is_available()` is `False`) breaks
    `torch.compile` on the CPU path: the first compile raises `RuntimeError:
    Can't detect vectorized ISA for CPU`
    
([#&#8203;189194](https://redirect.github.com/pytorch/pytorch/issues/189194)).
    This is a regression from `torch==2.12.1+rocm7.2`, which compiles CPU
    code fine (detecting e.g. `VecAVX2`) in the same setup. The 2.13 ROCm
    wheel appears to rely on something present in the ROCm builder image to
    detect the CPU vectorized ISA, so it works when run on a ROCm image but
    fails on a plain CPU-only image.
    
    Workaround: run the `+rocm` wheel on a ROCm image, or install a standard
    CPU/CUDA build for GPU-less environments.
    
    ### Backwards Incompatible Changes
    
    - Stop building CPython 3.13t (free-threaded) binaries
    
([#&#8203;182951](https://redirect.github.com/pytorch/pytorch/issues/182951))
    
    Upstream `pypa/manylinux` removed CPython 3.13t (free-threaded) on
    2026-05-07, because 3.13t
    was experimental and has been superseded by the now-non-experimental
    CPython 3.14t. As a result,
    PyTorch 2.13 no longer ships `cp313t` wheels (Linux, Triton, and related
    artifacts). Users on the
      free-threaded interpreter should move to Python 3.14t.
    
      PyTorch 2.12:
    
      ```bash
      # cp313t (free-threaded 3.13) wheels were available
      python3.13t -m pip install torch
      ```
    
      PyTorch 2.13:
    
      ```bash
      # Use free-threaded Python 3.14t instead
      python3.14t -m pip install torch
      ```
    
    - Bare `PyObject` is no longer allowed in operator schemas
    
([#&#8203;184209](https://redirect.github.com/pytorch/pytorch/issues/184209))
    
    Bare `PyObject` was accidentally accepted in operator schema strings in
    PyTorch 2.12. This was undocumented and is now rejected, since
    `torch.compile`
      does not support arbitrary `PyObject` inputs to custom ops. If
    you parse or register a schema with a bare `PyObject` argument or return
    type,
      you will now get a schema parse error.
    
      PyTorch 2.12:
    
      ```python
      >>> from torch._C import parse_schema
      >>> parse_schema("foo(PyObject x) -> ()")  # accepted
      ```
    
      PyTorch 2.13:
    
      ```python
      >>> from torch._C import parse_schema
    >>> parse_schema("foo(PyObject x) -> ()") # raises a schema parse error
      ```
    
    - Remove Bazel build support
    
([#&#8203;180883](https://redirect.github.com/pytorch/pytorch/issues/180883))
    
    The Bazel build was never broadly adopted and still depended on the
    antiquated Bazel 6,
    while the wider ecosystem has since moved to Bazel 9. All Bazel build
    files and CI jobs have
    been removed. Users building PyTorch with Bazel should migrate to the
    supported CMake/`pip install`
      build flow.
    
      PyTorch 2.12:
    
      ```bash
      # Build PyTorch with Bazel
      bazel build //:torch
      ```
    
      PyTorch 2.13:
    
      ```bash
    # Bazel build files have been removed; build from source with pip
    instead
      pip install --no-build-isolation -e .
      ```
    
    - Enforce C++20 minimum in header guards
    
([#&#8203;178150](https://redirect.github.com/pytorch/pytorch/issues/178150)).
    In PyTorch 2.13, C++20 is now required to import PyTorch headers.
    
    - `StorageImpl`'s built-in copy-on-write (COW) materialization is
    replaced by a pluggable materializer hook
    
([#&#8203;179063](https://redirect.github.com/pytorch/pytorch/issues/179063))
    
    `StorageImpl` no longer knows about COW directly. Its internal COW entry
    points
    `StorageImpl::is_cow()`, `StorageImpl::maybe_materialize_cow()`, and the
    friend
    `cow::materialize_cow_storage()` have been removed in favor of a single
    pluggable
    `MaterializeFn` hook (`void(*)(StorageImpl*)`) that a backend registers
    to run once,
    on the first mutable data-pointer access. COW is now just one consumer
    of this hook
    (`c10::impl::cow::materialize_cow`), and all COW behavior (lazy clone,
    refcounted
    shared data, copy-on-write) is unchanged. This also gives accelerator
    backends and
    eager-mode graph compilers a zero-fast-path-cost place to commit
    deferred allocations
      or materialize symbolic buffers on first mutation.
    
    This is a C++-only change. It affects out-of-tree backends/extensions
    that called the
    removed `StorageImpl` COW symbols directly; they will fail to compile
    against 2.13
    with errors such as `no member named 'is_cow' in 'c10::StorageImpl'`.
    Migrate to the
    new hook API (`set_materializer()` / `has_materializer()` /
    `clear_materializer()`).
    
      PyTorch 2.12:
    
      ```cpp
      // Detect a COW storage and force it to materialize.
      if (storage.is_cow()) {
        storage.maybe_materialize_cow();
      }
      ```
    
      PyTorch 2.13:
    
      ```cpp
    // Register a one-shot materializer; it runs on the next mutable-data
    access
    // and then clears itself. COW registers c10::impl::cow::materialize_cow
    this way.
    storage.set_materializer(&my_backend_materialize); // void(StorageImpl*)
    
    // `has_materializer()` replaces `is_cow()` for "is a deferred
    materialization pending?"
      if (storage.has_materializer()) { /* ... */ }
      ```
    
    - Convert `shared_ptr<Node>` to `intrusive_ptr<Node>`
    
([#&#8203;181139](https://redirect.github.com/pytorch/pytorch/issues/181139)).
    This changes the signature of `Tensor::grad_fn`. Accesses to
    `Tensor.grad_fn()` should change from `std::shared_ptr<Node>` to
    `c10::intrusive_ptr<Node>`. Similarly, construction of a C++ autograd
    function should change:
    
      PyTorch 2.12:
    
      ```cpp
    std::shared_ptr<CustomCppNode> node(new CustomCppNode(),
    torch::autograd::deleteNode);
      ```
    
      PyTorch 2.13:
    
      ```cpp
      auto node = c10::make_intrusive<CustomCppNode>();
      ```
    
    - The minimum supported NCCL version when building from source is now
    2.23
    
([#&#8203;186292](https://redirect.github.com/pytorch/pytorch/issues/186292))
    
    PyTorch now requires NCCL >= 2.23 at compile time, and the
    preprocessor/runtime gates that guarded NCCL features introduced in 2.23
    or earlier have been removed. Users who build PyTorch from source
    against a system NCCL older than 2.23 will hit compile errors against
    the dropped gates. Upgrade the NCCL installation to >= 2.23 to build.
    The prebuilt PyTorch wheels already bundle a compatible NCCL, so
    pip/conda users are unaffected.
    
    - Remove named tensors
    
([#&#8203;173895](https://redirect.github.com/pytorch/pytorch/issues/173895))
    
    The named tensor feature (a long-deprecated prototype) has been fully
    removed to reduce overhead and code bloat. All associated Python and C++
    APIs are gone, including `Tensor.names`, `Tensor.rename()`,
    `Tensor.refine_names()`, `Tensor.align_to()`, `Tensor.align_as()`,
    `torch.align_tensors()`, the `names=` keyword on factory functions (e.g.
    `torch.zeros`, `torch.empty`, `torch.ones`), and the C++ `Dimname` /
    `DimnameList` APIs. Code that previously relied on named dimensions must
    track dimension order positionally and avoid usage of any of these
    now-removed APIs or op overloads.
    
    - The `onednn::qconv2d_pointwise.binary` and `.binary_tensor` operators
    no longer alias their input but rather return fresh tensors. Previously
    these ops mutated the `qaccum` input buffer and returned it directly,
    violating the PyTorch invariant that custom operator outputs must not
    alias inputs. This silently bypassed aliasing checks via the old `->
    Tensor(a!)` schema and would become a hard error in a future PyTorch
    version (as mentioned in
    
[#&#8203;182063](https://redirect.github.com/pytorch/pytorch/issues/182063)),
    so the schema and implementation were corrected to return a fresh
    output. Most users are unaffected, only code that calls these ops
    directly and relies on the in-place mutation of qaccum must now read the
    returned tensor instead.
    
([#&#8203;177171](https://redirect.github.com/pytorch/pytorch/issues/177171))
    
    ### Deprecations
    
    - Custom operators that return an output aliasing one of their inputs
    are deprecated
    
([#&#8203;182063](https://redirect.github.com/pytorch/pytorch/issues/182063))
    
    When a custom operator returns an output that is the same tensor as (or
    otherwise aliases) one of its inputs under `torch.compile`, PyTorch now
    emits a `UserWarning` stating that this is deprecated and will become an
    error in a future version of PyTorch. Previously the warning stated the
    change would land in PyTorch 2.12; that timeline has been pushed back.
    To update your code, return a clone of the offending output instead of
    the input, or refactor the operator so it does not return the aliased
    tensor.
    
      Deprecated:
    
      ```python
      @&#8203;torch.library.custom_op("mylib::foo", mutates_args=())
      def foo(x: torch.Tensor) -> torch.Tensor:
          return x  # output aliases the input -- deprecated
      ```
    
      Updated:
    
      ```python
      @&#8203;torch.library.custom_op("mylib::foo", mutates_args=())
      def foo(x: torch.Tensor) -> torch.Tensor:
          return x.clone()  # return a clone instead
      ```
    
    - Creating tensors with the quantized dtypes `quint8`, `qint8`, and
    `qint32` is now deprecated and emits a warning. This covers both Python
    and C++ call sites; see
    [#&#8203;184982](https://redirect.github.com/pytorch/pytorch/issues/184982)
    for migration guidance
    
([#&#8203;184984](https://redirect.github.com/pytorch/pytorch/issues/184984))
    
      PyTorch 2.12:
    
      ```python
    >>> x = torch.quantize_per_tensor(torch.randn(3), 0.1, 0, torch.quint8)
      ```
    
      PyTorch 2.13:
    
      ```python
    >>> x = torch.quantize_per_tensor(torch.randn(3), 0.1, 0, torch.quint8)
    UserWarning: Creating tensors with quantized dtypes (quint8, qint8,
    qint32) is deprecated
      ```
    
    - Rename distributed collective ops to the `_single` naming scheme and
    deprecate the old names
    
([#&#8203;186123](https://redirect.github.com/pytorch/pytorch/issues/186123),
    [#&#8203;186124](https://redirect.github.com/pytorch/pytorch/issues/186124),
    [#&#8203;186125](https://redirect.github.com/pytorch/pytorch/issues/186125),
    [#&#8203;186134](https://redirect.github.com/pytorch/pytorch/issues/186134),
    [#&#8203;186135](https://redirect.github.com/pytorch/pytorch/issues/186135),
    [#&#8203;186144](https://redirect.github.com/pytorch/pytorch/issues/186144))
    
    To align the public `torch.distributed` collective APIs with the naming
    used by torchcomms' `TorchCommBackend`, `all_gather_into_tensor` is
    renamed to `all_gather_single` and `reduce_scatter_tensor` to
    `reduce_scatter_single`. The previous names continue to work as thin
    wrappers that delegate to the new functions, but now emit a
    `FutureWarning`.
    
      PyTorch 2.12:
    
      ```python
      dist.all_gather_into_tensor(output, input)
      dist.reduce_scatter_tensor(output, input)
      ```
    
      PyTorch 2.13:
    
      ```python
      dist.all_gather_single(output, input)
      dist.reduce_scatter_single(output, input)
      ```
    
    ### New Features
    
    #### Python Frontend
    
    - Add two new operator tags, `torch.Tag.inplace` and `torch.Tag.out`,
    that let an operator declare how it writes its result: `inplace` means
    it mutates a tensor in place, and `out` means it writes into a
    caller-provided output tensor. Native PyTorch operators are tagged
    automatically, and custom operators defined with `torch.library` can opt
    in by adding the tag. To be tagged `inplace`, an operator must take the
    tensor it mutates as its first positional argument (declared as
    `Tensor(a!)`, and the only mutable argument) and return that same
    tensor. Tagging a custom operator this way improves its behavior under
    `torch.compile`: `inplace` ops now go through `auto_functionalize`, so
    the reinplacing pass can analyze clones and skip unnecessary copies, and
    both `inplace` and `out` ops get their fake/meta kernels generated for
    free. See the [Python custom operators
    
tutorial](https://docs.pytorch.org/tutorials/advanced/python_custom_ops.html)
    for how to author and tag custom operators.
    
([#&#8203;181100](https://redirect.github.com/pytorch/pytorch/issues/181100),
    [#&#8203;181099](https://redirect.github.com/pytorch/pytorch/issues/181099),
    [#&#8203;184199](https://redirect.github.com/pytorch/pytorch/issues/184199),
    [#&#8203;184200](https://redirect.github.com/pytorch/pytorch/issues/184200),
    [#&#8203;184201](https://redirect.github.com/pytorch/pytorch/issues/184201),
    [#&#8203;184202](https://redirect.github.com/pytorch/pytorch/issues/184202),
    [#&#8203;184203](https://redirect.github.com/pytorch/pytorch/issues/184203),
    [#&#8203;180851](https://redirect.github.com/pytorch/pytorch/issues/180851),
    [#&#8203;180852](https://redirect.github.com/pytorch/pytorch/issues/180852))
    - Add `const_data_ptr()` Python binding to `torch.Tensor` for read-only
    data pointer access
    
([#&#8203;180382](https://redirect.github.com/pytorch/pytorch/issues/180382))
    - Add an `abbr` property to `torch.dtype` that returns a dtype's short
    string abbreviation (e.g. `torch.float32.abbr` returns `"f32"`)
    
([#&#8203;177296](https://redirect.github.com/pytorch/pytorch/issues/177296))
    - Allow positional arguments to be passed as keyword arguments to
    autograd custom `Function`s
    
([#&#8203;182206](https://redirect.github.com/pytorch/pytorch/issues/182206))
    - Expose `rearrange` in the `torch.func` namespace for einops-style
    tensor reshaping
    
([#&#8203;173183](https://redirect.github.com/pytorch/pytorch/issues/173183))
    
    #### torch.nn
    
    - Add `nn.LinearCrossEntropyLoss`, a fused linear-projection plus
    cross-entropy loss module that avoids materializing the full logits
    tensor
    
([#&#8203;181573](https://redirect.github.com/pytorch/pytorch/issues/181573),
    [#&#8203;185852](https://redirect.github.com/pytorch/pytorch/issues/185852),
    [#&#8203;172286](https://redirect.github.com/pytorch/pytorch/issues/172286),
    [#&#8203;186113](https://redirect.github.com/pytorch/pytorch/issues/186113))
    
    #### Autograd
    
    - Add `torch.autograd.graph.region_activation_memory_budget`
    
([#&#8203;185979](https://redirect.github.com/pytorch/pytorch/issues/185979))
    - Support passing gradient inputs as a `dict` to `torch.autograd.grad`
    and `torch.autograd.backward`
    
([#&#8203;178140](https://redirect.github.com/pytorch/pytorch/issues/178140))
    
    #### Distributed
    
    - Add a registration API for symmetric memory arguments
    (`lib.register_symm_mem_args()`), letting operators (including
    out-of-tree ops) declare which arguments require symmetric-memory
    allocation
    
([#&#8203;173513](https://redirect.github.com/pytorch/pytorch/issues/173513))
    
    - Remove `NCCLSymmetricMemory`'s explicit dependency on
    `ProcessGroupNCCL`, enabling symmetric memory to work with out-of-tree
    backends such as torchcomms
    
([#&#8203;184260](https://redirect.github.com/pytorch/pytorch/issues/184260))
    
    - Support accessing the `ReduceOp.PREMUL_SUM` factor from Python when
    implementing process group backends in Python
    
([#&#8203;185863](https://redirect.github.com/pytorch/pytorch/issues/185863))
    
    - Expose the NCCL 2.30 `maxP2pPeers` config binding
    
([#&#8203;181686](https://redirect.github.com/pytorch/pytorch/issues/181686))
    
    - Add rocSHMEM Triton integration for symmetric memory on ROCm
    
([#&#8203;178658](https://redirect.github.com/pytorch/pytorch/issues/178658))
    
    - Support passing extra keyword arguments to the loss function in
    pipeline schedules via a new `loss_kwargs` parameter to `step()`,
    enabling loss functions that require arguments beyond `(output, target)`
    (such as chunked cross-entropy needing token counts for scaling)
    
([#&#8203;181057](https://redirect.github.com/pytorch/pytorch/issues/181057))
    
    #### Distributed FSDP2
    
    - Add `FSDPModule.set_separate_reduce_scatter_group` to give
    reduce-scatter its own NCCL communicator, enabling opt-in overlap of
    all-gather and reduce-scatter
    
([#&#8203;186335](https://redirect.github.com/pytorch/pytorch/issues/186335))
    - Add `set_reduce_scatter_max_input_buffers` to keep multiple
    reduce-scatter input buffers in flight, so backward compute no longer
    stalls waiting to recycle a single reduce-scatter buffer
    
([#&#8203;186000](https://redirect.github.com/pytorch/pytorch/issues/186000))
    
    #### Profiler
    
    - Profiler/Kineto now emits channel metadata on CUDA backends
    
([#&#8203;185968](https://redirect.github.com/pytorch/pytorch/issues/185968))
    
    #### Dynamo
    
    - Add `torch.compiler.set_default_backend` to override the default
    `torch.compile` backend globally, so out-of-tree backend authors don't
    need to pass `backend=` at every call site (following the pattern of
    `torch.set_default_dtype`/`torch.set_default_device`). Explicit
    `backend=` arguments still take precedence
    
([#&#8203;178944](https://redirect.github.com/pytorch/pytorch/issues/178944))
    - Add `torch.compile(f, isolate_recompiles=True)` to give each
    `torch.compile` call its own isolated cache bucket, preventing
    cross-compile interference in cache lookups and recompile-limit checks
    when multiple `torch.compile` calls target the same function
    
([#&#8203;178351](https://redirect.github.com/pytorch/pytorch/issues/178351))
    - Add `register_multi_grad_hook` support to `@leaf_function`, allowing a
    backward hook to fire once per backward pass when all `requires_grad`
    inputs have their gradients computed
    
([#&#8203;179609](https://redirect.github.com/pytorch/pytorch/issues/179609))
    
    #### Inductor
    
    - Add flash-decoding support to the CPU `FlexAttention` template (chosen
    when query length is 1) with a new configurable `PARTITION_SIZE` kernel
    option
    
([#&#8203;159835](https://redirect.github.com/pytorch/pytorch/issues/159835))
    - Add Triton convolution backward kernels (input and weight gradients)
    as an autotuning backend in place of the ATen-only fallback
    
([#&#8203;178945](https://redirect.github.com/pytorch/pytorch/issues/178945))
    - Add an Inductor FX pass (`decomp_comms`) that eliminates `all_gather`
    for Gram-matrix optimizer patterns (Muon/Shampoo) under FSDP, gated by
    `config.aten_distributed_optimizations.allow_comms_decompositions`,
    yielding 1.25-1.95x training speedups
    
([#&#8203;184370](https://redirect.github.com/pytorch/pytorch/issues/184370))
    
    #### Ahead-Of-Time Inductor (AOTI)
    
    - Generated C shims for the AOTI stable ABI are now versioned and gated
    by `TORCH_TARGET_VERSION`, so shims introduced in newer releases are
    only exposed when the target version supports them
    
([#&#8203;181916](https://redirect.github.com/pytorch/pytorch/issues/181916))
    - Triton CPU AOTI models now work end-to-end through the public
    `torch._inductor.aoti_compile_and_package` / `aoti_load_package` API,
    including packaging and loading of the multiple `.so` files emitted per
    kernel
    
([#&#8203;182251](https://redirect.github.com/pytorch/pytorch/issues/182251))
    - Added stable C shim functions (`torch_exception_get_what`,
    `torch_exception_get_what_without_backtrace`, and
    `STABLE_TORCH_ERROR_CODE_CHECK`) so extensions built against the stable
    ABI can retrieve the original error message across the C API boundary
    (target version 2.13+)
    
([#&#8203;180135](https://redirect.github.com/pytorch/pytorch/issues/180135))
    - Added a stable AOTI stream shim `aoti_torch_stream_native_handle` and
    `torch::stable::accelerator::Stream::nativeHandle()`, gated behind
    `TORCH_FEATURE_VERSION >= 2.13`, for retrieving a native stream handle
    from the stable ABI
    
([#&#8203;183930](https://redirect.github.com/pytorch/pytorch/issues/183930))
    
    #### Release Engineering
    
    - Add Python 3.15 wheel builds across Linux (CPU/CUDA), Triton, ROCm,
    and XPU
    
([#&#8203;182954](https://redirect.github.com/pytorch/pytorch/issues/182954),
    [#&#8203;184600](https://redirect.github.com/pytorch/pytorch/issues/184600),
    [#&#8203;185409](https://redirect.github.com/pytorch/pytorch/issues/185409),
    [#&#8203;184829](https://redirect.github.com/pytorch/pytorch/issues/184829),
    [#&#8203;184891](https://redirect.github.com/pytorch/pytorch/issues/184891),
    [#&#8203;184906](https://redirect.github.com/pytorch/pytorch/issues/184906),
    [#&#8203;185094](https://redirect.github.com/pytorch/pytorch/issues/185094))
    
    #### CUDA
    
    - Add `CUDAGraph.get_graph_data()` for graph topology introspection
    
([#&#8203;183165](https://redirect.github.com/pytorch/pytorch/issues/183165))
    - Lightweight API to get private pool reserved memory bytes
    
([#&#8203;178240](https://redirect.github.com/pytorch/pytorch/issues/178240))
    
    #### MPS
    
    - Add FlexAttention support for MPS
    
([#&#8203;182552](https://redirect.github.com/pytorch/pytorch/issues/182552),
    [#&#8203;186215](https://redirect.github.com/pytorch/pytorch/issues/186215))
    - Add support for `torch.distributions.Dirichlet` on MPS by adding
    `_sample_dirichlet` and `_dirichlet_grad` Metal implementations
    
([#&#8203;185458](https://redirect.github.com/pytorch/pytorch/issues/185458),
    [#&#8203;185854](https://redirect.github.com/pytorch/pytorch/issues/185854))
    - Add `grid_sampler_2d` backward support on MPS
    
([#&#8203;179756](https://redirect.github.com/pytorch/pytorch/issues/179756))
    - Add `grid_sampler_3d` backward support on MPS
    
([#&#8203;179388](https://redirect.github.com/pytorch/pytorch/issues/179388))
    - Add `lcm` support on MPS via a new Metal kernel
    
([#&#8203;186279](https://redirect.github.com/pytorch/pytorch/issues/186279))
    - Add complex support to `c10/metal/reduction_utils.h`
    
([#&#8203;180708](https://redirect.github.com/pytorch/pytorch/issues/180708))
    and a complex->bool specialization
    
([#&#8203;185938](https://redirect.github.com/pytorch/pytorch/issues/185938))
    
    #### ROCm
    
    - Enable external events in CUDA graphs
    
([#&#8203;178264](https://redirect.github.com/pytorch/pytorch/issues/178264))
    - Enable GPU Address Sanitizer build
    
([#&#8203;183792](https://redirect.github.com/pytorch/pytorch/issues/183792),
    [#&#8203;176461](https://redirect.github.com/pytorch/pytorch/issues/176461))
    - Improve Inductor GEMM search space performance using the Origami
    project
    
([#&#8203;172512](https://redirect.github.com/pytorch/pytorch/issues/172512))
    - Use CMake native HIP language support, `enable_language(HIP)`
    
([#&#8203;180485](https://redirect.github.com/pytorch/pytorch/issues/180485))
    - New Inductor benchmarker based on Torch Profiler
    
([#&#8203;175097](https://redirect.github.com/pytorch/pytorch/issues/175097))
    
    #### XPU
    
    - Add XPU device telemetry APIs for temperature, frequency, power draw,
    engine utilization, memory bandwidth usage, and used device memory
    through `torch.xpu.*`
    
([#&#8203;181082](https://redirect.github.com/pytorch/pytorch/issues/181082),
    [#&#8203;183427](https://redirect.github.com/pytorch/pytorch/issues/183427),
    [#&#8203;183428](https://redirect.github.com/pytorch/pytorch/issues/183428),
    [#&#8203;183429](https://redirect.github.com/pytorch/pytorch/issues/183429),
    [#&#8203;183430](https://redirect.github.com/pytorch/pytorch/issues/183430),
    [#&#8203;183431](https://redirect.github.com/pytorch/pytorch/issues/183431))
    - Add FP8 blockwise scaling support for `scaled_mm` on XPU
    
([#&#8203;173630](https://redirect.github.com/pytorch/pytorch/issues/173630),
    [#&#8203;176043](https://redirect.github.com/pytorch/pytorch/issues/176043))
    
    ### Improvements
    
    #### Python Frontend
    
    - Make it possible to load safetensors with `torch.load`
    
([#&#8203;170592](https://redirect.github.com/pytorch/pytorch/issues/170592))
    - Make `Storage.pin_memory` / `Storage.is_pinned` device-agnostic
    
([#&#8203;186223](https://redirect.github.com/pytorch/pytorch/issues/186223))
    - Add `op_overloads` to `OpOverloadPacket` to enumerate an operator's
    overloads
    
([#&#8203;182993](https://redirect.github.com/pytorch/pytorch/issues/182993))
    
    #### torch.nn
    
    - Expose `num_splits` in FlashAttention-2 and bump the flash-attention
    submodule
    
([#&#8203;179760](https://redirect.github.com/pytorch/pytorch/issues/179760))
    - Support `linear_bias` in `linear_cross_entropy` on the reference and
    chunked paths
    
([#&#8203;185129](https://redirect.github.com/pytorch/pytorch/issues/185129),
    [#&#8203;185276](https://redirect.github.com/pytorch/pytorch/issues/185276))
    
    #### Optimizer
    
    - Fix `SequentialLR` wrong learning rate initialization when
    `milestones` contain 0
    
([#&#8203;185986](https://redirect.github.com/pytorch/pytorch/issues/185986))
    
    #### Autograd
    
    - Implement autograd derivatives for `torch.nextafter`
    
([#&#8203;148820](https://redirect.github.com/pytorch/pytorch/issues/148820))
    - Add `torch.autograd.enforce_grad_layout_policy` to control the memory
    layout policy for accumulated gradients
    
([#&#8203;180552](https://redirect.github.com/pytorch/pytorch/issues/180552))
    
    #### Distributed
    
    - When TorchComms is enabled, route `new_group` through `split_group`
    for subgroup creation, raising `NotImplementedError` for arguments
    `split_group` cannot honor (e.g. `use_local_synchronization=True`,
    `sort_ranks=False`) instead of silently falling back
    
([#&#8203;185416](https://redirect.github.com/pytorch/pytorch/issues/185416))
    
    - Delegate `dist.new_group` to custom process group subclasses
    
([#&#8203;184262](https://redirect.github.com/pytorch/pytorch/issues/184262))
    
    - Surface started-work metadata in NCCL watchdog timeouts
    
([#&#8203;183656](https://redirect.github.com/pytorch/pytorch/issues/183656))
    
    - Add a health check endpoint to the distributed debug server
    
([#&#8203;179326](https://redirect.github.com/pytorch/pytorch/issues/179326))
    
    - Make the `DeviceMesh` non-overlapping check stricter
    
([#&#8203;172343](https://redirect.github.com/pytorch/pytorch/issues/172343))
    
    - Allow `elastic_launch`/`launch_agent` to accept a pre-created
    torchelastic health check server, so it can be started before rendezvous
    
([#&#8203;180543](https://redirect.github.com/pytorch/pytorch/issues/180543))
    
    - Add an `overlap_pp_comm` flag to pipeline schedules (default `True`)
    that, when set to `False`, defers each pipeline RECV op to immediately
    before the compute op that consumes it, using rank-parity P2P ordering
    to avoid deadlock (helps platforms such as AMD ROCm where a pending RECV
    blocks unrelated compute)
    
([#&#8203;178815](https://redirect.github.com/pytorch/pytorch/issues/178815))
    
    #### DTensor
    
    - Migrate embedding and random ops to single-dim sharding strategies and
    increase op coverage
    
([#&#8203;180281](https://redirect.github.com/pytorch/pytorch/issues/180281),
    [#&#8203;180503](https://redirect.github.com/pytorch/pytorch/issues/180503))
    - Add auto-infrastructure that derives single-dim sharding strategies
    for autogenerated op variants (`.out`, inplace, functional, and
    `foreach`), expanding strategy coverage to hundreds of additional ops
    
([#&#8203;185386](https://redirect.github.com/pytorch/pytorch/issues/185386))
    - Register sharding strategies for additional ops: `scatter`,
    upsample/interpolation backward, anti-aliased upsample, batch norm
    backward, and `aten.detach_.default`
    
([#&#8203;186149](https://redirect.github.com/pytorch/pytorch/issues/186149),
    [#&#8203;180311](https://redirect.github.com/pytorch/pytorch/issues/180311),
    [#&#8203;184626](https://redirect.github.com/pytorch/pytorch/issues/184626),
    [#&#8203;182743](https://redirect.github.com/pytorch/pytorch/issues/182743),
    [#&#8203;181876](https://redirect.github.com/pytorch/pytorch/issues/181876))
    
    #### Distributed FSDP2
    
    - Support forward-mode automatic differentiation (`torch.func.jvp`) on
    models wrapped with `fully_shard` or `replicate`, including with mixed
    precision
    
([#&#8203;182732](https://redirect.github.com/pytorch/pytorch/issues/182732))
    
    #### Linear Algebra Frontend
    
    - Add `Half` and `BFloat16` dispatch support for `torch.trace` on CPU
    
([#&#8203;184874](https://redirect.github.com/pytorch/pytorch/issues/184874))
    - Improve heuristics for the cuSOLVER vs cuBLAS backend switch in
    `torch.linalg.lu`
    
([#&#8203;185344](https://redirect.github.com/pytorch/pytorch/issues/185344))
    
    #### Profiler
    
    - The memory viz tool now more accurately represents GPU footprint when
    impacted by fragmentation
    
([#&#8203;180515](https://redirect.github.com/pytorch/pytorch/issues/180515))
    - The memory viz tool now aggregates stripes per-pool to improve
    visualization for large snapshots
    
([#&#8203;180613](https://redirect.github.com/pytorch/pytorch/issues/180613))
    - Profiler now also exposes CUDA occupancy metadata as a nested
    dictionary in the `.events()` output
    
([#&#8203;180275](https://redirect.github.com/pytorch/pytorch/issues/180275))
    
    #### FX
    
    - `split_module` now supports `torch.Size` crossing graph split
    boundaries by decomposing `size()` calls into per-dimension `sym_size`
    nodes, and builds submodules lazily for faster inference graph splitting
    
([#&#8203;179839](https://redirect.github.com/pytorch/pytorch/issues/179839))
    
    - `CapabilityBasedPartitioner` can now opt out of horizontal fusion via
    `skip_horizontal_fusion=True`, partitioning only through direct data
    dependencies
    
([#&#8203;184904](https://redirect.github.com/pytorch/pytorch/issues/184904))
    
    - Enable rewriting of FX traces containing complex tensors during
    compilation
    
([#&#8203;169832](https://redirect.github.com/pytorch/pytorch/issues/169832))
    
    #### Dynamo
    
    - Implement additional Python operators in Dynamo: bitwise and
    
([#&#8203;184788](https://redirect.github.com/pytorch/pytorch/issues/184788)),
    bitwise xor
    
([#&#8203;184789](https://redirect.github.com/pytorch/pytorch/issues/184789)),
    left/right shift
    
([#&#8203;183462](https://redirect.github.com/pytorch/pytorch/issues/183462)),
    floor division
    
([#&#8203;185652](https://redirect.github.com/pytorch/pytorch/issues/185652)),
    true division
    
([#&#8203;185653](https://redirect.github.com/pytorch/pytorch/issues/185653)),
    remainder
    
([#&#8203;185654](https://redirect.github.com/pytorch/pytorch/issues/185654)),
    and `divmod`
    
([#&#8203;185655](https://redirect.github.com/pytorch/pytorch/issues/185655))
    - Support tracing more constructs in Dynamo: `einops` 0.8.2
    
([#&#8203;185619](https://redirect.github.com/pytorch/pytorch/issues/185619)),
    `record_function` as a decorator
    
([#&#8203;184703](https://redirect.github.com/pytorch/pytorch/issues/184703)),
    `inference_mode` retracing helpers
    
([#&#8203;185066](https://redirect.github.com/pytorch/pytorch/issues/185066)),
    `mark_dirty` in the autograd Function HOP
    
([#&#8203;184267](https://redirect.github.com/pytorch/pytorch/issues/184267)),
    `warn_only` deterministic toggles
    
([#&#8203;180373](https://redirect.github.com/pytorch/pytorch/issues/180373)),
    and the `_maybe_view_chunk_cat` functional collective
    
([#&#8203;180389](https://redirect.github.com/pytorch/pytorch/issues/180389))
    - Support item assignment and deletion (`__setitem__`/`__delitem__`) on
    more container types in Dynamo via `sq_ass_item`/`mp_ass_subscript`
    slots
    
([#&#8203;182862](https://redirect.github.com/pytorch/pytorch/issues/182862),
    [#&#8203;182996](https://redirect.github.com/pytorch/pytorch/issues/182996))
    - Support `torch.accelerator.device_index` and `torch.xpu.device` in the
    device context manager
    
([#&#8203;181846](https://redirect.github.com/pytorch/pytorch/issues/181846),
    [#&#8203;181847](https://redirect.github.com/pytorch/pytorch/issues/181847))
    - Improve Triton support under `torch.compile`: accept `tl.constexpr`
    values as kernel arguments
    
([#&#8203;181783](https://redirect.github.com/pytorch/pytorch/issues/181783))
    and handle `capture_triton` as a no-op during tracing
    
([#&#8203;183555](https://redirect.github.com/pytorch/pytorch/issues/183555))
    - Improve dynamic shape specification: reduce verbosity in shape specs
    for the common case
    
([#&#8203;184271](https://redirect.github.com/pytorch/pytorch/issues/184271)),
    add `SeqSpec` for list/tuple specs with better walk-spec errors
    
([#&#8203;185327](https://redirect.github.com/pytorch/pytorch/issues/185327)),
    add `ObjectSpec`
    
([#&#8203;182764](https://redirect.github.com/pytorch/pytorch/issues/182764)),
    pipe dynamic spec through `torch.compile`
    
([#&#8203;184501](https://redirect.github.com/pytorch/pytorch/issues/184501)),
    and revisit guarding in `mark_dynamic` APIs
    
([#&#8203;181469](https://redirect.github.com/pytorch/pytorch/issues/181469))
    - Improve `torch.compile` device mismatch errors with a dedicated
    `FakeTensorDeviceMismatchError` and actionable guidance to place inputs,
    parameters, and buffers on the same device
    
([#&#8203;185412](https://redirect.github.com/pytorch/pytorch/issues/185412))
    - Improve error messages and diagnostics: clearer data-dependent errors
    for `.any()`/`.all()`
    
([#&#8203;180406](https://redirect.github.com/pytorch/pytorch/issues/180406)),
    clearer `torch._check` tensor predicate errors
    
([#&#8203;185777](https://redirect.github.com/pytorch/pytorch/issues/185777)),
    user-friendly reasons for skipped frames
    
([#&#8203;183596](https://redirect.github.com/pytorch/pytorch/issues/183596)),
    carets in stack traces
    
([#&#8203;182393](https://redirect.github.com/pytorch/pytorch/issues/182393)),
    and reporting why a symbol was created dynamically in `symbolic_shapes`
    logs
    
([#&#8203;168331](https://redirect.github.com/pytorch/pytorch/issues/168331))
    - Make Dynamo exceptions pickleable
    
([#&#8203;185725](https://redirect.github.com/pytorch/pytorch/issues/185725))
    - Inline decomposed quantization helpers in Dynamo
    
([#&#8203;185628](https://redirect.github.com/pytorch/pytorch/issues/185628))
    - Make Dynamo debug/repro utilities device-agnostic
    
([#&#8203;184851](https://redirect.github.com/pytorch/pytorch/issues/184851))
    
    #### Inductor
    
    - Support `pin_memory` for `torch.ones`, `torch.zeros` and `torch.full`
    
([#&#8203;174595](https://redirect.github.com/pytorch/pytorch/issues/174595))
    - Add a ROCm config flag to disable the `pointer_range_32` optimization
    
([#&#8203;179604](https://redirect.github.com/pytorch/pytorch/issues/179604))
    - Add a peak memory threshold config for combo kernels
    
([#&#8203;180578](https://redirect.github.com/pytorch/pytorch/issues/180578))
    - Enable autotuning and a fast compensation path for CPU static/dynamic
    smooth-quant qlinear, with correct handling of 0D `x_scale`/`x_zp` and
    `ReinterpretView` strides
    
([#&#8203;181090](https://redirect.github.com/pytorch/pytorch/issues/181090))
    - Add `aot_inductor.autotune_per_kernel_alloc` config to
    allocate-run-delete tensors per kernel during AOTI autotuning, avoiding
    OOM on large models
    
([#&#8203;181176](https://redirect.github.com/pytorch/pytorch/issues/181176))
    - Bound `AsyncCompile` future waits with the
    `compile_worker_wait_timeout` setting
    
([#&#8203;181293](https://redirect.github.com/pytorch/pytorch/issues/181293))
    - Add `a100_default_flex_config` entries for `head_dim=192`
    
([#&#8203;181835](https://redirect.github.com/pytorch/pytorch/issues/181835))
    - Add an explicit lowering (fallback) for `aten.multinomial` to avoid
    graph breaks
    
([#&#8203;182423](https://redirect.github.com/pytorch/pytorch/issues/182423))
    - Add an Inductor lowering for `_scaled_mm_v2`
    
([#&#8203;182527](https://redirect.github.com/pytorch/pytorch/issues/182527))
    - Enable `combo_kernel_autotune_grouping` by default
    
([#&#8203;182567](https://redirect.github.com/pytorch/pytorch/issues/182567))
    - Add `cudagraph_partition_memory_budget` config for partition
    reordering
    
([#&#8203;183569](https://redirect.github.com/pytorch/pytorch/issues/183569))
    - Add ATen fallbacks for `bincount`, `unique` variants, and AMP scale
    ops so they compile without graph breaks
    
([#&#8203;183590](https://redirect.github.com/pytorch/pytorch/issues/183590))
    - Unfuse the bias add from `addmm` when the bias is a narrowing dtype
    cast (fp32->bf16/fp16) to preserve bias precision in XPU AMP training
    
([#&#8203;183680](https://redirect.github.com/pytorch/pytorch/issues/183680))
    - Emit a clearer diagnostic when a backward CUDAGraph output installed
    as a `.grad` buffer is invalidated on a later run (gradient
    accumulation)
    
([#&#8203;184003](https://redirect.github.com/pytorch/pytorch/issues/184003))
    - Support split online softmax reductions in Inductor
    
([#&#8203;184069](https://redirect.github.com/pytorch/pytorch/issues/184069))
    - Support fusing `index_add`-style atomic scatter mutations into Triton
    template epilogues behind a config flag
    
([#&#8203;184179](https://redirect.github.com/pytorch/pytorch/issues/184179))
    - Add a `cpp.march` Inductor config knob so AOTInductor cpp-only builds
    can override or suppress the default CPU architecture flag
    
([#&#8203;184297](https://redirect.github.com/pytorch/pytorch/issues/184297))
    - Improve the Triton cache directory guidance when loading shared
    objects from a noexec filesystem fails
    
([#&#8203;184362](https://redirect.github.com/pytorch/pytorch/issues/184362))
    - Enable the Bert SDPA pattern rewrite on CUDA while keeping the
    original matmul/softmax math path
    
([#&#8203;184417](https://redirect.github.com/pytorch/pytorch/issues/184417))
    - Improve the Triton launcher argument-mismatch error with a clearer
    message and a cached preflight check
    
([#&#8203;184522](https://redirect.github.com/pytorch/pytorch/issues/184522))
    - Add broader CuteDSL op overrides (`rsqrt`, `exp2`, `log2`, `log10`,
    `tan`, `acos`, `asin`, `atan`, `atan2`, `floor`, `logical_xor`)
    
([#&#8203;184538](https://redirect.github.com/pytorch/pytorch/issues/184538))
    - Add an optional `fake_mode` argument to `standalone_compile` (with
    `dynamic_shapes="from_example_inputs"`) so it can reuse the caller's
    `FakeTensorMode`/`ShapeEnv` instead of always creating a fresh one
    
([#&#8203;184776](https://redirect.github.com/pytorch/pytorch/issues/184776))
    - Support `signbit` on unsigned integer dtypes
    
([#&#8203;185985](https://redirect.github.com/pytorch/pytorch/issues/185985))
    - Add a `keep_static_cubin_raw` config to retain cubin bytes in cached
    kernels so caches restored on another machine avoid recompilation
    
([#&#8203;186404](https://redirect.github.com/pytorch/pytorch/issues/186404))
    - Extend `BatchLinearLHSFusion`'s matcher to also match the inlined
    `torch._C._nn.linear` form so the (opt-in) fusion can fire on
    Dynamo-inlined linear
    
([#&#8203;186632](https://redirect.github.com/pytorch/pytorch/issues/186632))
    - Add a clearer error message with install instructions when a
    compatible Flash Attention package is unavailable for flex attention
    
([#&#8203;186827](https://redirect.github.com/pytorch/pytorch/issues/186827))
    - Add another anchor node to the batch-linear fusion pass so more small
    `torch._C._nn.linear` operations are grouped into a single batched
    kernel
    
([#&#8203;180477](https://redirect.github.com/pytorch/pytorch/issues/180477))
    - Extend the batch fusion pass to support `detach()` method calls
    
([#&#8203;180513](https://redirect.github.com/pytorch/pytorch/issues/180513))
    - Add a deterministic backward for the FlexAttention flash kernel
    
([#&#8203;174813](https://redirect.github.com/pytorch/pytorch/issues/174813))
    - Use rand4x for Inductor Triton random number generation
    
([#&#8203;184377](https://redirect.github.com/pytorch/pytorch/issues/184377))
    - Add a Quack-based CuTeDSL RMSNorm kernel
    
([#&#8203;182108](https://redirect.github.com/pytorch/pytorch/issues/182108))
    
    #### Ahead-Of-Time Inductor (AOTI)
    
    - Use fatbinary for multi-arch CUDA kernels
    
([#&#8203;184456](https://redirect.github.com/pytorch/pytorch/issues/184456))
    - Support mixed-device constants in `update_constant_buffer`
    
([#&#8203;181114](https://redirect.github.com/pytorch/pytorch/issues/181114))
    - Add FP8 header files in the AOTI `shim.h`
    
([#&#8203;178120](https://redirect.github.com/pytorch/pytorch/issues/178120))
    - Add throttled `cudaMemcpy` for AOTI constant loading to reduce peak
    memory usage
    
([#&#8203;184823](https://redirect.github.com/pytorch/pytorch/issues/184823))
    - Preserve AOTI `proxy_executor` error messages
    
([#&#8203;180884](https://redirect.github.com/pytorch/pytorch/issues/180884))
    - Enable Triton kernels in AOTI C++ wrapper on CPU
    
([#&#8203;181068](https://redirect.github.com/pytorch/pytorch/issues/181068))
    - Skip CPU vec ISA setup for device-only `cpp_wrapper`
    
([#&#8203;182089](https://redirect.github.com/pytorch/pytorch/issues/182089))
    - Expose torchbind constants from `AOTIModelPackageLoader`
    
([#&#8203;182149](https://redirect.github.com/pytorch/pytorch/issues/182149))
    - Improve AOTI error for Python custom ops
    
([#&#8203;186305](https://redirect.github.com/pytorch/pytorch/issues/186305))
    
    #### Export
    
    - Support serialization of opaque type constants in `torch.export`
    save/load
    
([#&#8203;181676](https://redirect.github.com/pytorch/pytorch/issues/181676))
    - Make functorch JVP operator `torch.export`-able
    
([#&#8203;179686](https://redirect.github.com/pytorch/pytorch/issues/179686))
    - Add `UpdateConstantBufferFromCpu` for host-to-device copy
    
([#&#8203;181637](https://redirect.github.com/pytorch/pytorch/issues/181637))
    
    #### AOTAutograd
    
    - Support CPU activation offloading in the rematerialization pass,
    including marking recomputed nodes for backward and adding a resize-to-0
    deallocation op so offloaded tensors are freed after their
    host-to-device copy
    
([#&#8203;181937](https://redirect.github.com/pytorch/pytorch/issues/181937),
    [#&#8203;181938](https://redirect.github.com/pytorch/pytorch/issues/181938))
    - Use FX node names in `merge_view_inputs` error messages, so
    non-differentiable view input mutation errors identify the specific
    offending inputs
    
([#&#8203;180424](https://redirect.github.com/pytorch/pytorch/issues/180424))
    
    #### Composability
    
    - Add fake tensor support for `_transformer_encoder_layer_fwd` so it
    traces under `torch.compile`
    
([#&#8203;183916](https://redirect.github.com/pytorch/pytorch/issues/183916))
    - Enable Armv9-A target support for `torch.compile` on AArch64
    
([#&#8203;184555](https://redirect.github.com/pytorch/pytorch/issues/184555))
    - Functionalize in-place `c10d` collectives in standalone compile
    
([#&#8203;181836](https://redirect.github.com/pytorch/pytorch/issues/181836))
    
    #### Foreach
    
    - Fix/add empty check for `_foreach_max`
    
([#&#8203;173483](https://redirect.github.com/pytorch/pytorch/issues/173483))
    
    #### ONNX
    
    - Add `adaptive_max_pool2d` and `adaptive_max_pool3d` decompositions for
    ONNX export
    
([#&#8203;184396](https://redirect.github.com/pytorch/pytorch/issues/184396))
    
    #### C++ Frontend
    
    - Add stable ABI for `set_python_module` on `torch::Library`
    
([#&#8203;182720](https://redirect.github.com/pytorch/pytorch/issues/182720))
    - Add `==` overloads for `HeaderOnlyArrayRef`
    
([#&#8203;185379](https://redirect.github.com/pytorch/pytorch/issues/185379))
    - Add `torch::stable::Generator`
    
([#&#8203;186423](https://redirect.github.com/pytorch/pytorch/issues/186423))
    - Add `c10::layout` typecaster for `torch.layout`
    
([#&#8203;179607](https://redirect.github.com/pytorch/pytorch/issues/179607))
    - Add default-args support to `def_static`
    
([#&#8203;175644](https://redirect.github.com/pytorch/pytorch/issues/175644))
    - Add support for controlling scientific notation in C++-side tensor
    printing
    
([#&#8203;173321](https://redirect.github.com/pytorch/pytorch/issues/173321))
    
    #### Release Engineering
    
    - Move the NCCL pin to 2.30
    
([#&#8203;181313](https://redirect.github.com/pytorch/pytorch/issues/181313))
    - Advance the Triton pin to 3.7.1
    
([#&#8203;181001](https://redirect.github.com/pytorch/pytorch/issues/181001),
    [#&#8203;186792](https://redirect.github.com/pytorch/pytorch/issues/186792))
    - Upgrade the XPU support package to 2026.0
    
([#&#8203;182003](https://redirect.github.com/pytorch/pytorch/issues/182003))
    - Add a configurable threshold to avoid power-of-two rounding for large
    pinned memory allocations
    
([#&#8203;171662](https://redirect.github.com/pytorch/pytorch/issues/171662))
    - Move some pre-build steps from `setup.py` to CMake
    
([#&#8203;177641](https://redirect.github.com/pytorch/pytorch/issues/177641))
    
    #### CUDA
    
    - Debugging tool to verify that external inputs to a CUDA graph are
    alive before replay
    
([#&#8203;174649](https://redirect.github.com/pytorch/pytorch/issues/174649))
    - Add get/set/reset functions for BLAS workspace sizes
    
([#&#8203;177912](https://redirect.github.com/pytorch/pytorch/issues/177912))
    - Cleanup double import in `BinaryDivFloorKernel.cu`
    
([#&#8203;179260](https://redirect.github.com/pytorch/pytorch/issues/179260))
    - Return supported CUDA arch list when no GPU is present but GPU is
    compiled
    
([#&#8203;180356](https://redirect.github.com/pytorch/pytorch/issues/180356))
    - Detect and fix stale stream references in autograd during CUDA graph
    capture
    
([#&#8203;180090](https://redirect.github.com/pytorch/pytorch/issues/180090))
    - Use `opmath_t` in `i1` and `i1e` CUDA kernels
    
([#&#8203;183778](https://redirect.github.com/pytorch/pytorch/issues/183778))
    - Support `resize_` with address hint
    
([#&#8203;178215](https://redirect.github.com/pytorch/pytorch/issues/178215))
    - Support `bfloat16` in `_embedding_bag_per_sample_weights_backward` on
    CUDA
    
([#&#8203;185889](https://redirect.github.com/pytorch/pytorch/issues/185889))
    - Align `parsePerProcessMemoryFraction`'s return type with other parsers
    
([#&#8203;185139](https://redirect.github.com/pytorch/pytorch/issues/185139))
    - Improve error message when `cuda-bindings` version is too old
    
([#&#8203;185990](https://redirect.github.com/pytorch/pytorch/issues/185990))
    - Expose `torch.cuda.current_solver_handle` for cuSOLVER handle sharing
    
([#&#8203;176705](https://redirect.github.com/pytorch/pytorch/issues/176705))
    
    #### cuDNN
    
    - Add flag to select depthwise convolution backend
    
([#&#8203;176500](https://redirect.github.com/pytorch/pytorch/issues/176500))
    - Upgrade `cudnn_frontend` submodule to 1.24
    
([#&#8203;185554](https://redirect.github.com/pytorch/pytorch/issues/185554))
    
    #### MPS
    
    - Migrate many ops from MPSGraph to native Metal kernels for `bernoulli`
    
([#&#8203;182210](https://redirect.github.com/pytorch/pytorch/issues/182210)),
    `native_dropout`
    
([#&#8203;182232](https://redirect.github.com/pytorch/pytorch/issues/182232)),
    `uniform`/`normal`/`randint`
    
([#&#8203;182386](https://redirect.github.com/pytorch/pytorch/issues/182386)),
    `randperm`
    
([#&#8203;182528](https://redirect.github.com/pytorch/pytorch/issues/182528)),
    comparison ops `eq`/`ne`/`lt`/`le`/`gt`/`ge`
    
([#&#8203;183019](https://redirect.github.com/pytorch/pytorch/issues/183019)),
    bitwise ops
    
([#&#8203;182839](https://redirect.github.com/pytorch/pytorch/issues/182839)),
    scatter/gather
    
([#&#8203;184028](https://redirect.github.com/pytorch/pytorch/issues/184028)),
    copy-cast
    
([#&#8203;184740](https://redirect.github.com/pytorch/pytorch/issues/184740)),
    `gelu`/`gelu_backward`
    
([#&#8203;181451](https://redirect.github.com/pytorch/pytorch/issues/181451)),
    replication pad
    
([#&#8203;183065](https://redirect.github.com/pytorch/pytorch/issues/183065)),
    embedding backward
    
([#&#8203;185119](https://redirect.github.com/pytorch/pytorch/issues/185119)),
    `trace`
    
([#&#8203;183627](https://redirect.github.com/pytorch/pytorch/issues/183627)),
    `count_nonzero`
    
([#&#8203;180725](https://redirect.github.com/pytorch/pytorch/issues/180725)),
    `amax`/`amin`/`aminmax`/`all`/`any`
    
([#&#8203;180752](https://redirect.github.com/pytorch/pytorch/issues/180752)),
    and `cumsum`/`cumprod`
    
([#&#8203;185609](https://redirect.github.com/pytorch/pytorch/issues/185609)),
    `native_group_norm`
    
([#&#8203;183830](https://redirect.github.com/pytorch/pytorch/issues/183830),
    `native_group_norm_backward`
    
[#&#8203;184437](https://redirect.github.com/pytorch/pytorch/issues/184437)),
    `topk` and `kthvalue` Metal kernels
    
([#&#8203;184106](https://redirect.github.com/pytorch/pytorch/issues/184106)),
    single-block and multi-block sort, including a stable sort path
    
([#&#8203;180714](https://redirect.github.com/pytorch/pytorch/issues/180714),
    [#&#8203;182242](https://redirect.github.com/pytorch/pytorch/issues/182242),
    [#&#8203;181736](https://redirect.github.com/pytorch/pytorch/issues/181736))
    - Support integer inputs to `histc`
    
([#&#8203;178624](https://redirect.github.com/pytorch/pytorch/issues/178624))
    - Support `out` variants of unary ops
    
([#&#8203;184743](https://redirect.github.com/pytorch/pytorch/issues/184743))
    - Improvements to SDPA Metal kernels: prefill attention kernels
    
([#&#8203;181575](https://redirect.github.com/pytorch/pytorch/issues/181575)),
    GQA support
    
([#&#8203;183280](https://redirect.github.com/pytorch/pytorch/issues/183280)),
    `is_causal` support
    
([#&#8203;181855](https://redirect.github.com/pytorch/pytorch/issues/181855)),
    head dim 256
    
([#&#8203;181852](https://redirect.github.com/pytorch/pytorch/issues/181852)),
    float mask support
    
([#&#8203;183458](https://redirect.github.com/pytorch/pytorch/issues/183458)),
    and a clearer error for causal + attn mask
    
([#&#8203;181856](https://redirect.github.com/pytorch/pytorch/issues/181856))
    - Add an ILP variant for binary tensor iterators and replace unary VEC4
    with a generic ILP-per-thread dense kernel
    
([#&#8203;182155](https://redirect.github.com/pytorch/pytorch/issues/182155),
    [#&#8203;181509](https://redirect.github.com/pytorch/pytorch/issues/181509),
    [#&#8203;183055](https://redirect.github.com/pytorch/pytorch/issues/183055))
    - Make several Metal kernels stride-aware to avoid unnecessary
    `.contiguous()` calls: `Col2Im`
    
([#&#8203;181949](https://redirect.github.com/pytorch/pytorch/issues/181949)),
    `Im2Col`
    
([#&#8203;182709](https://redirect.github.com/pytorch/pytorch/issues/182709)),
    `Repeat`
    
([#&#8203;182718](https://redirect.github.com/pytorch/pytorch/issues/182718)),
    `LossOps`
    
([#&#8203;182714](https://redirect.github.com/pytorch/pytorch/issues/182714)),
    and `HistogramKernel`
    
([#&#8203;181951](https://redirect.github.com/pytorch/pytorch/issues/181951))
    - Enable `NDHWC`+`DHWIO` fast path for `Conv3d` on `channels_last_3d`
    
([#&#8203;184612](https://redirect.github.com/pytorch/pytorch/issues/184612))
    - Clear the MPSGraph cache in `torch.mps.empty_cache()`
    
([#&#8203;181485](https://redirect.github.com/pytorch/pytorch/issues/181485))
    - Add a private API to get the host alias of Metal storage
    
([#&#8203;180961](https://redirect.github.com/pytorch/pytorch/issues/180961))
    - Add input validation: `stride > 0` in pool ops
    
([#&#8203;184875](https://redirect.github.com/pytorch/pytorch/issues/184875)),
    `F.fold`
    
([#&#8203;182067](https://redirect.github.com/pytorch/pytorch/issues/182067)),
    `im2col`
    
([#&#8203;183593](https://redirect.github.com/pytorch/pytorch/issues/183593)),
    and bernoulli probabilities
    
([#&#8203;185065](https://redirect.github.com/pytorch/pytorch/issues/185065))
    - Alert on non-deterministic algorithms on MPS
    
([#&#8203;185061](https://redirect.github.com/pytorch/pytorch/issues/185061))
    - Raise a clear not-implemented error for `dropout_p`
    
([#&#8203;184126](https://redirect.github.com/pytorch/pytorch/issues/184126))
    - Use proper precise log functions in unary kernels
    
([#&#8203;185381](https://redirect.github.com/pytorch/pytorch/issues/185381))
    
    #### ROCm
    
    - Additional `cub::DeviceHistogram` hipify mappings
    
([#&#8203;180433](https://redirect.github.com/pytorch/pytorch/issues/180433))
    - SDPA improvements via AOTriton 0.12b: `head_dim != head_dim_v`,
    `use_deterministic_algorithms`, `gfx1100` and `gfx1151` promoted out of
    experimental, partial FAv3 support on `gfx950`
    
([#&#8203;184288](https://redirect.github.com/pytorch/pytorch/issues/184288))
    
    #### XPU
    
    - Add `last_level_cache_size` and `is_integrated_gpu` to XPU device
    properties
    
([#&#8203;184499](https://redirect.github.com/pytorch/pytorch/issues/184499),
    [#&#8203;182624](https://redirect.github.com/pytorch/pytorch/issues/182624))
    - Add XPU dispatch for `_fused_adagrad_`
    
([#&#8203;185577](https://redirect.github.com/pytorch/pytorch/issues/185577))
    - Support mixed-type operations between Nested and Dense tensors on XPU
    
([#&#8203;182654](https://redirect.github.com/pytorch/pytorch/issues/182654))
    - Support `torch.xpu.device` in Dynamo device management
    
([#&#8203;181847](https://redirect.github.com/pytorch/pytorch/issues/181847))
    - Recognize additional Intel BMG device IDs on XPU
    
([#&#8203;183414](https://redirect.github.com/pytorch/pytorch/issues/183414))
    - Enable XPU device support for sparse Triton ops
    
([#&#8203;179805](https://redirect.github.com/pytorch/pytorch/issues/179805))
    - Enable the `bmm_outer_product` Triton override on XPU
    
([#&#8203;180441](https://redirect.github.com/pytorch/pytorch/issues/180441))
    - Improve test coverage for the XPU backend
    
([#&#8203;174370](https://redirect.github.com/pytorch/pytorch/issues/174370),
    [#&#8203;180881](https://redirect.github.com/pytorch/pytorch/issues/180881),
    [#&#8203;171154](https://redirect.github.com/pytorch/pytorch/issues/171154))
    - Support non-blocking pinned device-to-host copies on XPU
    
([#&#8203;186224](https://redirect.github.com/pytorch/pytorch/issues/186224))
    - Refactor the XPU oneDNN integration from the C API to the C++ API
    
([#&#8203;184486](https://redirect.github.com/pytorch/pytorch/issues/184486))
    
    #### Functorch
    
    - Add `vmap` batching rule for `torch.unbind_copy`
    ([#&#8203;178035](https://redirect.github.com/pyt
    
    > ✂ **Note**
    >
    > PR body was truncated to here.
    
    
    </details>
    
    ---
    
    ### Configuration
    
    📅 **Schedule**: (in timezone Etc/UTC)
    
    - Branch creation
      - At any time (no schedule defined)
    - Automerge
      - At any time (no schedule defined)
    
    🚦 **Automerge**: Disabled by config. Please merge this manually once you
    are satisfied.
    
    ♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the
    rebase/retry checkbox.
    
    🔕 **Ignore**: Close this PR and you won't be reminded about these
    updates again.
    
    ---
    
    - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check
    this box
    
    ---
    
    This PR was generated by [Mend Renovate](https://mend.io/renovate/).
    View the [repository job
    log](https://developer.mend.io/github/apache/texera).
    
    
<!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0My4yODAuMCIsInVwZGF0ZWRJblZlciI6IjQzLjI4MC4wIiwidGFyZ2V0QnJhbmNoIjoibWFpbiIsImxhYmVscyI6WyJkZXBlbmRlbmNpZXMiLCJyZWxlYXNlL3YxLjIiLCJzZWN1cml0eSJdfQ==-->
    
    ---------
    
    Co-authored-by: Xinyuan Lin <[email protected]>
    Co-authored-by: Claude Fable 5 <[email protected]>
    Co-authored-by: Yicong Huang 
<[email protected]>
---
 amber/LICENSE-binary-python     | 2 +-
 amber/operator-requirements.txt | 4 ++--
 2 files changed, 3 insertions(+), 3 deletions(-)

diff --git a/amber/LICENSE-binary-python b/amber/LICENSE-binary-python
index f950686bd3..03dfd037f7 100644
--- a/amber/LICENSE-binary-python
+++ b/amber/LICENSE-binary-python
@@ -318,7 +318,7 @@ Python packages:
   - sympy==1.14.0
   - threadpoolctl==3.6.0
   - tifffile==2026.6.1
-  - torch==2.12.1
+  - torch==2.13.0
   - zstandard==0.25.0
 
 
--------------------------------------------------------------------------------
diff --git a/amber/operator-requirements.txt b/amber/operator-requirements.txt
index bf3e0ff96c..d78fe77fe8 100644
--- a/amber/operator-requirements.txt
+++ b/amber/operator-requirements.txt
@@ -22,8 +22,8 @@ pillow==12.3.0
 
 # Pin torch to the CPU wheel on Linux x86_64 to avoid the NVIDIA CUDA deps.
 --extra-index-url https://download.pytorch.org/whl/cpu
-torch==2.12.1+cpu ; platform_system == "Linux" and platform_machine == "x86_64"
-torch==2.12.1 ; platform_system != "Linux" or platform_machine != "x86_64"
+torch==2.13.0+cpu ; platform_system == "Linux" and platform_machine == "x86_64"
+torch==2.13.0 ; platform_system != "Linux" or platform_machine != "x86_64"
 
 scikit-learn==1.7.2
 transformers==5.5.0

Reply via email to