When amdgpu_vm_ptes_update fails, there might be jobs leak and mappings
might be kept in an on-stack list, which will lead to invalid list
walking later on.

One of the approaches that was attempted was introducing an abort
function to amdgpu_vm_update_funcs. But given the possibility that some
mappings have already been pushed and committed and some may have been
moved to tlb_flush_waitlist, do a partial commit instead and return the
error to amdgpu_vm_update_range caller.

Also, fix amdgpu_vm_clear_freed ignoring the result of
amdgpu_vm_update_range.

I just noticed a different fix for the latter was submitted at [1]. It
re-adds the mapping to the freed list instead of not removing it when
amdgpu_vm_update_range fails.

[1] https://lore.kernel.org/amd-gfx/[email protected]/

Signed-off-by: Thadeu Lima de Souza Cascardo <[email protected]>
---
Thadeu Lima de Souza Cascardo (4):
      drm/amdgpu: make amdgpu_vm_sdma_commit handle NULL or empty job
      drm/amdgpu: make amdgpu_vm_update_funcs commit return void
      drm/amdgpu: don't free mapping when amdgpu_vm_update_range fails
      drm/amdgpu: commit pending work when amdgpu_vm_ptes_update fails

 drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c      | 14 +++++---------
 drivers/gpu/drm/amd/amdgpu/amdgpu_vm.h      |  4 ++--
 drivers/gpu/drm/amd/amdgpu/amdgpu_vm_cpu.c  |  8 +++-----
 drivers/gpu/drm/amd/amdgpu/amdgpu_vm_pt.c   |  2 +-
 drivers/gpu/drm/amd/amdgpu/amdgpu_vm_sdma.c | 26 ++++++++++++++++----------
 5 files changed, 27 insertions(+), 27 deletions(-)
---
base-commit: 59ced288fcba9e91bd38e61a972ad782c4edb7d0
change-id: 20260914-amdgpu_vm_fixes-d64d1478657e

Best regards,
--  
Thadeu Lima de Souza Cascardo <[email protected]>

Reply via email to