* [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range [not found] <2026092253-dimness-unethical-2515@gregkh> @ 2026-09-22 15:33 ` Vadim Nikitushkin 2026-09-23 7:41 ` Greg KH 0 siblings, 1 reply; 4+ messages in thread From: Vadim Nikitushkin @ 2026-09-22 15:33 UTC (permalink / raw) To: stable Cc: gregkh, christian.koenig, thomas.hellstrom, dri-devel, skainsworth, Vadim Nikitushkin commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream. ttm_tt_swapout() returns the number of pages swapped out on success and a negative error code on failure; for a populated ttm it never returns zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure") moved the bulk_move bookkeeping in ttm_bo_swapout_cb() under "if (!ret)", so the ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail() pair is now skipped on every successful swapout. The equivalent change for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink() infinite LRU walk on backup failure") tests "lret > 0", which is what was intended here as well. Before b2ed01e7ad3d the resource was taken off the bulk_move before the swapout; since then a swapped-out resource stays inside its BO's bulk_move range (and on the manager LRU) although it is unevictable. When it is later freed or the BO leaves the bulk_move (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()), ttm_resource_del_bulk_move() skips it because of its !ttm_resource_unevictable() guard, so a range endpoint in pos->first / pos->last is left pointing at freed memory. The next ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(), "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL dereference in ttm_resource_manager_next() -- minutes to hours after a hibernation, or at process exit / reboot following one. Samuel Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the dangling cursor; the missing removal at swapout time is the reason it dangles. Testing the condition for success restores the removal. On an AMD Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug crashed 5 of 18 hibernation cycles; a function profile of one hibernation showed 336 ttm_tt_swapout() calls and zero ttm_resource_del_bulk_move_unevictable() calls. With this change the removal happens for every swapped-out resource and 12 further cycles were clean. Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure") Cc: stable@vger.kernel.org # v7.1+ Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387 Link: https://lore.kernel.org/dri-devel/CAHYiNPa6aVacJoLOje-qZ1GyYx-9p0tN4NuP8D_eSL+UJeevXw@mail.gmail.com/ Signed-off-by: Vadim Nikitushkin <bub4z0r@gmail.com> Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com> Reviewed-by: Christian König <christian.koenig@amd.com> Signed-off-by: Christian König <christian.koenig@amd.com> Link: https://lore.kernel.org/r/20260909205028.13799-1-bub4z0r@gmail.com [ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed ("drm/ttm: apply the swapout bulk_move fix to the intended condition"): upstream 3db7d7d58341 was applied to the "if (ret)" after ttm_resource_try_charge() in ttm_bo_alloc_at_place() instead of the "if (!ret)" after ttm_tt_swapout() in ttm_bo_swapout_cb(), and fcfe64715b42 restored the former and changed the latter. 7.2.y has no ttm_resource_try_charge() (dmem cgroup charging is 7.3 material), so neither commit applies on its own; the net effect of the two is the single hunk below, which is the change the changelog describes. ] --- drivers/gpu/drm/ttm/ttm_bo.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/gpu/drm/ttm/ttm_bo.c b/drivers/gpu/drm/ttm/ttm_bo.c index bcd76f6..ca6fc5c 100644 --- a/drivers/gpu/drm/ttm/ttm_bo.c +++ b/drivers/gpu/drm/ttm/ttm_bo.c @@ -1178,7 +1178,7 @@ ttm_bo_swapout_cb(struct ttm_lru_walk *walk, struct ttm_buffer_object *bo) if (ttm_tt_is_populated(tt)) { ret = ttm_tt_swapout(bdev, tt, swapout_walk->gfp_flags); - if (!ret) { + if (ret > 0) { spin_lock(&bdev->lru_lock); ttm_resource_del_bulk_move_unevictable(bo->resource, bo); ttm_resource_move_to_lru_tail(bo->resource); -- 2.53.0 ^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range 2026-09-22 15:33 ` [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range Vadim Nikitushkin @ 2026-09-23 7:41 ` Greg KH 2026-09-23 7:53 ` Christian König 0 siblings, 1 reply; 4+ messages in thread From: Greg KH @ 2026-09-23 7:41 UTC (permalink / raw) To: Vadim Nikitushkin Cc: stable, christian.koenig, thomas.hellstrom, dri-devel, skainsworth On Tue, Sep 22, 2026 at 06:33:34PM +0300, Vadim Nikitushkin wrote: > commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream. > > ttm_tt_swapout() returns the number of pages swapped out on success and > a negative error code on failure; for a populated ttm it never returns > zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU > walk on swapout failure") moved the bulk_move bookkeeping in > ttm_bo_swapout_cb() under "if (!ret)", so the > ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail() > pair is now skipped on every successful swapout. The equivalent change > for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink() > infinite LRU walk on backup failure") tests "lret > 0", which is what > was intended here as well. > > Before b2ed01e7ad3d the resource was taken off the bulk_move before the > swapout; since then a swapped-out resource stays inside its BO's > bulk_move range (and on the manager LRU) although it is unevictable. > When it is later freed or the BO leaves the bulk_move > (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()), > ttm_resource_del_bulk_move() skips it because of its > !ttm_resource_unevictable() guard, so a range endpoint in pos->first / > pos->last is left pointing at freed memory. The next > ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor > is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(), > "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL > dereference in ttm_resource_manager_next() -- minutes to hours after a > hibernation, or at process exit / reboot following one. Samuel > Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the > dangling cursor; the missing removal at swapout time is the reason it > dangles. > > Testing the condition for success restores the removal. On an AMD > Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on > a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug > crashed 5 of 18 hibernation cycles; a function profile of one > hibernation showed 336 ttm_tt_swapout() calls and zero > ttm_resource_del_bulk_move_unevictable() calls. With this change the > removal happens for every swapped-out resource and 12 further cycles > were clean. > > Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure") > Cc: stable@vger.kernel.org # v7.1+ > Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387 > Link: https://lore.kernel.org/dri-devel/CAHYiNPa6aVacJoLOje-qZ1GyYx-9p0tN4NuP8D_eSL+UJeevXw@mail.gmail.com/ > Signed-off-by: Vadim Nikitushkin <bub4z0r@gmail.com> > Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com> > Reviewed-by: Christian König <christian.koenig@amd.com> > Signed-off-by: Christian König <christian.koenig@amd.com> > Link: https://lore.kernel.org/r/20260909205028.13799-1-bub4z0r@gmail.com > [ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed Do not squash, send a patch series instead please. thanks, greg k-h ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range 2026-09-23 7:41 ` Greg KH @ 2026-09-23 7:53 ` Christian König 2026-09-23 9:29 ` Greg KH 0 siblings, 1 reply; 4+ messages in thread From: Christian König @ 2026-09-23 7:53 UTC (permalink / raw) To: Greg KH, Vadim Nikitushkin Cc: stable, thomas.hellstrom, dri-devel, skainsworth Hi Greg, On 9/23/26 09:41, Greg KH wrote: > On Tue, Sep 22, 2026 at 06:33:34PM +0300, Vadim Nikitushkin wrote: >> commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream. >> >> ttm_tt_swapout() returns the number of pages swapped out on success and >> a negative error code on failure; for a populated ttm it never returns >> zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU >> walk on swapout failure") moved the bulk_move bookkeeping in >> ttm_bo_swapout_cb() under "if (!ret)", so the >> ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail() >> pair is now skipped on every successful swapout. The equivalent change >> for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink() >> infinite LRU walk on backup failure") tests "lret > 0", which is what >> was intended here as well. >> >> Before b2ed01e7ad3d the resource was taken off the bulk_move before the >> swapout; since then a swapped-out resource stays inside its BO's >> bulk_move range (and on the manager LRU) although it is unevictable. >> When it is later freed or the BO leaves the bulk_move >> (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()), >> ttm_resource_del_bulk_move() skips it because of its >> !ttm_resource_unevictable() guard, so a range endpoint in pos->first / >> pos->last is left pointing at freed memory. The next >> ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor >> is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(), >> "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL >> dereference in ttm_resource_manager_next() -- minutes to hours after a >> hibernation, or at process exit / reboot following one. Samuel >> Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the >> dangling cursor; the missing removal at swapout time is the reason it >> dangles. >> >> Testing the condition for success restores the removal. On an AMD >> Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on >> a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug >> crashed 5 of 18 hibernation cycles; a function profile of one >> hibernation showed 336 ttm_tt_swapout() calls and zero >> ttm_resource_del_bulk_move_unevictable() calls. With this change the >> removal happens for every swapped-out resource and 12 further cycles >> were clean. >> >> Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure") >> Cc: stable@vger.kernel.org # v7.1+ >> Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387 >> Link: https://lore.kernel.org/dri-devel/CAHYiNPa6aVacJoLOje-qZ1GyYx-9p0tN4NuP8D_eSL+UJeevXw@mail.gmail.com/ >> Signed-off-by: Vadim Nikitushkin <bub4z0r@gmail.com> >> Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com> >> Reviewed-by: Christian König <christian.koenig@amd.com> >> Signed-off-by: Christian König <christian.koenig@amd.com> >> Link: https://lore.kernel.org/r/20260909205028.13799-1-bub4z0r@gmail.com >> [ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed > > Do not squash, send a patch series instead please. in this particular case that squashing is the correct approach. AMDs mail servers mangled the initial patch so badly that I had trouble applying it and ended up pushing a broken patch upstream. The squash is basically fixing that up. Back-porting each patch individually doesn't make much sense, you would just end up with a broken tree in between. Sorry for that, Christian. > > thanks, > > greg k-h ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range 2026-09-23 7:53 ` Christian König @ 2026-09-23 9:29 ` Greg KH 0 siblings, 0 replies; 4+ messages in thread From: Greg KH @ 2026-09-23 9:29 UTC (permalink / raw) To: Christian König Cc: Vadim Nikitushkin, stable, thomas.hellstrom, dri-devel, skainsworth On Wed, Sep 23, 2026 at 09:53:54AM +0200, Christian König wrote: > Hi Greg, > > On 9/23/26 09:41, Greg KH wrote: > > On Tue, Sep 22, 2026 at 06:33:34PM +0300, Vadim Nikitushkin wrote: > >> commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream. > >> > >> ttm_tt_swapout() returns the number of pages swapped out on success and > >> a negative error code on failure; for a populated ttm it never returns > >> zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU > >> walk on swapout failure") moved the bulk_move bookkeeping in > >> ttm_bo_swapout_cb() under "if (!ret)", so the > >> ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail() > >> pair is now skipped on every successful swapout. The equivalent change > >> for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink() > >> infinite LRU walk on backup failure") tests "lret > 0", which is what > >> was intended here as well. > >> > >> Before b2ed01e7ad3d the resource was taken off the bulk_move before the > >> swapout; since then a swapped-out resource stays inside its BO's > >> bulk_move range (and on the manager LRU) although it is unevictable. > >> When it is later freed or the BO leaves the bulk_move > >> (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()), > >> ttm_resource_del_bulk_move() skips it because of its > >> !ttm_resource_unevictable() guard, so a range endpoint in pos->first / > >> pos->last is left pointing at freed memory. The next > >> ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor > >> is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(), > >> "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL > >> dereference in ttm_resource_manager_next() -- minutes to hours after a > >> hibernation, or at process exit / reboot following one. Samuel > >> Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the > >> dangling cursor; the missing removal at swapout time is the reason it > >> dangles. > >> > >> Testing the condition for success restores the removal. On an AMD > >> Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on > >> a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug > >> crashed 5 of 18 hibernation cycles; a function profile of one > >> hibernation showed 336 ttm_tt_swapout() calls and zero > >> ttm_resource_del_bulk_move_unevictable() calls. With this change the > >> removal happens for every swapped-out resource and 12 further cycles > >> were clean. > >> > >> Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure") > >> Cc: stable@vger.kernel.org # v7.1+ > >> Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387 > >> Link: https://lore.kernel.org/dri-devel/CAHYiNPa6aVacJoLOje-qZ1GyYx-9p0tN4NuP8D_eSL+UJeevXw@mail.gmail.com/ > >> Signed-off-by: Vadim Nikitushkin <bub4z0r@gmail.com> > >> Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com> > >> Reviewed-by: Christian König <christian.koenig@amd.com> > >> Signed-off-by: Christian König <christian.koenig@amd.com> > >> Link: https://lore.kernel.org/r/20260909205028.13799-1-bub4z0r@gmail.com > >> [ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed > > > > Do not squash, send a patch series instead please. > > in this particular case that squashing is the correct approach. > > AMDs mail servers mangled the initial patch so badly that I had trouble applying it and ended up pushing a broken patch upstream. The squash is basically fixing that up. Ok, I'll take this then, thanks for the response. greg k-h ^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-23 9:32 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <2026092253-dimness-unethical-2515@gregkh>
2026-09-22 15:33 ` [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range Vadim Nikitushkin
2026-09-23 7:41 ` Greg KH
2026-09-23 7:53 ` Christian König
2026-09-23 9:29 ` Greg KH
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox