On Wed, Sep 23, 2026 at 09:53:54AM +0200, Christian König wrote:
> Hi Greg,
> 
> On 9/23/26 09:41, Greg KH wrote:
> > On Tue, Sep 22, 2026 at 06:33:34PM +0300, Vadim Nikitushkin wrote:
> >> commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream.
> >>
> >> ttm_tt_swapout() returns the number of pages swapped out on success and
> >> a negative error code on failure; for a populated ttm it never returns
> >> zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU
> >> walk on swapout failure") moved the bulk_move bookkeeping in
> >> ttm_bo_swapout_cb() under "if (!ret)", so the
> >> ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail()
> >> pair is now skipped on every successful swapout. The equivalent change
> >> for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink()
> >> infinite LRU walk on backup failure") tests "lret > 0", which is what
> >> was intended here as well.
> >>
> >> Before b2ed01e7ad3d the resource was taken off the bulk_move before the
> >> swapout; since then a swapped-out resource stays inside its BO's
> >> bulk_move range (and on the manager LRU) although it is unevictable.
> >> When it is later freed or the BO leaves the bulk_move
> >> (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()),
> >> ttm_resource_del_bulk_move() skips it because of its
> >> !ttm_resource_unevictable() guard, so a range endpoint in pos->first /
> >> pos->last is left pointing at freed memory. The next
> >> ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor
> >> is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(),
> >> "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL
> >> dereference in ttm_resource_manager_next() -- minutes to hours after a
> >> hibernation, or at process exit / reboot following one. Samuel
> >> Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the
> >> dangling cursor; the missing removal at swapout time is the reason it
> >> dangles.
> >>
> >> Testing the condition for success restores the removal. On an AMD
> >> Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on
> >> a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug
> >> crashed 5 of 18 hibernation cycles; a function profile of one
> >> hibernation showed 336 ttm_tt_swapout() calls and zero
> >> ttm_resource_del_bulk_move_unevictable() calls. With this change the
> >> removal happens for every swapped-out resource and 12 further cycles
> >> were clean.
> >>
> >> Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on 
> >> swapout failure")
> >> Cc: [email protected] # v7.1+
> >> Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387
> >> Link: 
> >> https://lore.kernel.org/dri-devel/cahyinpa6avacjoloje-qz1gyyx-9p0tn4nup8d_esl+ujee...@mail.gmail.com/
> >> Signed-off-by: Vadim Nikitushkin <[email protected]>
> >> Reviewed-by: Thomas Hellström <[email protected]>
> >> Reviewed-by: Christian König <[email protected]>
> >> Signed-off-by: Christian König <[email protected]>
> >> Link: https://lore.kernel.org/r/[email protected]
> >> [ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed
> > 
> > Do not squash, send a patch series instead please.
> 
> in this particular case that squashing is the correct approach.
> 
> AMDs mail servers mangled the initial patch so badly that I had trouble 
> applying it and ended up pushing a broken patch upstream. The squash is 
> basically fixing that up.

Ok, I'll take this then, thanks for the response.

greg k-h

Reply via email to