dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range
       [not found] <2026092253-dimness-unethical-2515@gregkh>
@ 2026-09-22 15:33 ` Vadim Nikitushkin
  2026-09-23  7:41   ` Greg KH
  0 siblings, 1 reply; 4+ messages in thread
From: Vadim Nikitushkin @ 2026-09-22 15:33 UTC (permalink / raw)
  To: stable
  Cc: gregkh, christian.koenig, thomas.hellstrom, dri-devel,
	skainsworth, Vadim Nikitushkin

commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream.

ttm_tt_swapout() returns the number of pages swapped out on success and
a negative error code on failure; for a populated ttm it never returns
zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU
walk on swapout failure") moved the bulk_move bookkeeping in
ttm_bo_swapout_cb() under "if (!ret)", so the
ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail()
pair is now skipped on every successful swapout. The equivalent change
for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink()
infinite LRU walk on backup failure") tests "lret > 0", which is what
was intended here as well.

Before b2ed01e7ad3d the resource was taken off the bulk_move before the
swapout; since then a swapped-out resource stays inside its BO's
bulk_move range (and on the manager LRU) although it is unevictable.
When it is later freed or the BO leaves the bulk_move
(ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()),
ttm_resource_del_bulk_move() skips it because of its
!ttm_resource_unevictable() guard, so a range endpoint in pos->first /
pos->last is left pointing at freed memory. The next
ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor
is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(),
"list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL
dereference in ttm_resource_manager_next() -- minutes to hours after a
hibernation, or at process exit / reboot following one. Samuel
Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the
dangling cursor; the missing removal at swapout time is the reason it
dangles.

Testing the condition for success restores the removal. On an AMD
Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on
a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug
crashed 5 of 18 hibernation cycles; a function profile of one
hibernation showed 336 ttm_tt_swapout() calls and zero
ttm_resource_del_bulk_move_unevictable() calls. With this change the
removal happens for every swapped-out resource and 12 further cycles
were clean.

Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure")
Cc: stable@vger.kernel.org # v7.1+
Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387
Link: https://lore.kernel.org/dri-devel/CAHYiNPa6aVacJoLOje-qZ1GyYx-9p0tN4NuP8D_eSL+UJeevXw@mail.gmail.com/
Signed-off-by: Vadim Nikitushkin <bub4z0r@gmail.com>
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Christian König <christian.koenig@amd.com>
Link: https://lore.kernel.org/r/20260909205028.13799-1-bub4z0r@gmail.com
[ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed
  ("drm/ttm: apply the swapout bulk_move fix to the intended condition"):
  upstream 3db7d7d58341 was applied to the "if (ret)" after
  ttm_resource_try_charge() in ttm_bo_alloc_at_place() instead of the
  "if (!ret)" after ttm_tt_swapout() in ttm_bo_swapout_cb(), and
  fcfe64715b42 restored the former and changed the latter. 7.2.y has no
  ttm_resource_try_charge() (dmem cgroup charging is 7.3 material), so
  neither commit applies on its own; the net effect of the two is the
  single hunk below, which is the change the changelog describes. ]
---
 drivers/gpu/drm/ttm/ttm_bo.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/ttm/ttm_bo.c b/drivers/gpu/drm/ttm/ttm_bo.c
index bcd76f6..ca6fc5c 100644
--- a/drivers/gpu/drm/ttm/ttm_bo.c
+++ b/drivers/gpu/drm/ttm/ttm_bo.c
@@ -1178,7 +1178,7 @@ ttm_bo_swapout_cb(struct ttm_lru_walk *walk, struct ttm_buffer_object *bo)
 
 	if (ttm_tt_is_populated(tt)) {
 		ret = ttm_tt_swapout(bdev, tt, swapout_walk->gfp_flags);
-		if (!ret) {
+		if (ret > 0) {
 			spin_lock(&bdev->lru_lock);
 			ttm_resource_del_bulk_move_unevictable(bo->resource, bo);
 			ttm_resource_move_to_lru_tail(bo->resource);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 4+ messages in thread

* Re: [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range
  2026-09-22 15:33 ` [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range Vadim Nikitushkin
@ 2026-09-23  7:41   ` Greg KH
  2026-09-23  7:53     ` Christian König
  0 siblings, 1 reply; 4+ messages in thread
From: Greg KH @ 2026-09-23  7:41 UTC (permalink / raw)
  To: Vadim Nikitushkin
  Cc: stable, christian.koenig, thomas.hellstrom, dri-devel,
	skainsworth

On Tue, Sep 22, 2026 at 06:33:34PM +0300, Vadim Nikitushkin wrote:
> commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream.
> 
> ttm_tt_swapout() returns the number of pages swapped out on success and
> a negative error code on failure; for a populated ttm it never returns
> zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU
> walk on swapout failure") moved the bulk_move bookkeeping in
> ttm_bo_swapout_cb() under "if (!ret)", so the
> ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail()
> pair is now skipped on every successful swapout. The equivalent change
> for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink()
> infinite LRU walk on backup failure") tests "lret > 0", which is what
> was intended here as well.
> 
> Before b2ed01e7ad3d the resource was taken off the bulk_move before the
> swapout; since then a swapped-out resource stays inside its BO's
> bulk_move range (and on the manager LRU) although it is unevictable.
> When it is later freed or the BO leaves the bulk_move
> (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()),
> ttm_resource_del_bulk_move() skips it because of its
> !ttm_resource_unevictable() guard, so a range endpoint in pos->first /
> pos->last is left pointing at freed memory. The next
> ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor
> is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(),
> "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL
> dereference in ttm_resource_manager_next() -- minutes to hours after a
> hibernation, or at process exit / reboot following one. Samuel
> Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the
> dangling cursor; the missing removal at swapout time is the reason it
> dangles.
> 
> Testing the condition for success restores the removal. On an AMD
> Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on
> a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug
> crashed 5 of 18 hibernation cycles; a function profile of one
> hibernation showed 336 ttm_tt_swapout() calls and zero
> ttm_resource_del_bulk_move_unevictable() calls. With this change the
> removal happens for every swapped-out resource and 12 further cycles
> were clean.
> 
> Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure")
> Cc: stable@vger.kernel.org # v7.1+
> Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387
> Link: https://lore.kernel.org/dri-devel/CAHYiNPa6aVacJoLOje-qZ1GyYx-9p0tN4NuP8D_eSL+UJeevXw@mail.gmail.com/
> Signed-off-by: Vadim Nikitushkin <bub4z0r@gmail.com>
> Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> Reviewed-by: Christian König <christian.koenig@amd.com>
> Signed-off-by: Christian König <christian.koenig@amd.com>
> Link: https://lore.kernel.org/r/20260909205028.13799-1-bub4z0r@gmail.com
> [ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed

Do not squash, send a patch series instead please.

thanks,

greg k-h

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range
  2026-09-23  7:41   ` Greg KH
@ 2026-09-23  7:53     ` Christian König
  2026-09-23  9:29       ` Greg KH
  0 siblings, 1 reply; 4+ messages in thread
From: Christian König @ 2026-09-23  7:53 UTC (permalink / raw)
  To: Greg KH, Vadim Nikitushkin
  Cc: stable, thomas.hellstrom, dri-devel, skainsworth

Hi Greg,

On 9/23/26 09:41, Greg KH wrote:
> On Tue, Sep 22, 2026 at 06:33:34PM +0300, Vadim Nikitushkin wrote:
>> commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream.
>>
>> ttm_tt_swapout() returns the number of pages swapped out on success and
>> a negative error code on failure; for a populated ttm it never returns
>> zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU
>> walk on swapout failure") moved the bulk_move bookkeeping in
>> ttm_bo_swapout_cb() under "if (!ret)", so the
>> ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail()
>> pair is now skipped on every successful swapout. The equivalent change
>> for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink()
>> infinite LRU walk on backup failure") tests "lret > 0", which is what
>> was intended here as well.
>>
>> Before b2ed01e7ad3d the resource was taken off the bulk_move before the
>> swapout; since then a swapped-out resource stays inside its BO's
>> bulk_move range (and on the manager LRU) although it is unevictable.
>> When it is later freed or the BO leaves the bulk_move
>> (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()),
>> ttm_resource_del_bulk_move() skips it because of its
>> !ttm_resource_unevictable() guard, so a range endpoint in pos->first /
>> pos->last is left pointing at freed memory. The next
>> ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor
>> is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(),
>> "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL
>> dereference in ttm_resource_manager_next() -- minutes to hours after a
>> hibernation, or at process exit / reboot following one. Samuel
>> Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the
>> dangling cursor; the missing removal at swapout time is the reason it
>> dangles.
>>
>> Testing the condition for success restores the removal. On an AMD
>> Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on
>> a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug
>> crashed 5 of 18 hibernation cycles; a function profile of one
>> hibernation showed 336 ttm_tt_swapout() calls and zero
>> ttm_resource_del_bulk_move_unevictable() calls. With this change the
>> removal happens for every swapped-out resource and 12 further cycles
>> were clean.
>>
>> Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure")
>> Cc: stable@vger.kernel.org # v7.1+
>> Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387
>> Link: https://lore.kernel.org/dri-devel/CAHYiNPa6aVacJoLOje-qZ1GyYx-9p0tN4NuP8D_eSL+UJeevXw@mail.gmail.com/
>> Signed-off-by: Vadim Nikitushkin <bub4z0r@gmail.com>
>> Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
>> Reviewed-by: Christian König <christian.koenig@amd.com>
>> Signed-off-by: Christian König <christian.koenig@amd.com>
>> Link: https://lore.kernel.org/r/20260909205028.13799-1-bub4z0r@gmail.com
>> [ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed
> 
> Do not squash, send a patch series instead please.

in this particular case that squashing is the correct approach.

AMDs mail servers mangled the initial patch so badly that I had trouble applying it and ended up pushing a broken patch upstream. The squash is basically fixing that up.

Back-porting each patch individually doesn't make much sense, you would just end up with a broken tree in between.

Sorry for that,
Christian.

> 
> thanks,
> 
> greg k-h


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range
  2026-09-23  7:53     ` Christian König
@ 2026-09-23  9:29       ` Greg KH
  0 siblings, 0 replies; 4+ messages in thread
From: Greg KH @ 2026-09-23  9:29 UTC (permalink / raw)
  To: Christian König
  Cc: Vadim Nikitushkin, stable, thomas.hellstrom, dri-devel,
	skainsworth

On Wed, Sep 23, 2026 at 09:53:54AM +0200, Christian König wrote:
> Hi Greg,
> 
> On 9/23/26 09:41, Greg KH wrote:
> > On Tue, Sep 22, 2026 at 06:33:34PM +0300, Vadim Nikitushkin wrote:
> >> commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream.
> >>
> >> ttm_tt_swapout() returns the number of pages swapped out on success and
> >> a negative error code on failure; for a populated ttm it never returns
> >> zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU
> >> walk on swapout failure") moved the bulk_move bookkeeping in
> >> ttm_bo_swapout_cb() under "if (!ret)", so the
> >> ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail()
> >> pair is now skipped on every successful swapout. The equivalent change
> >> for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink()
> >> infinite LRU walk on backup failure") tests "lret > 0", which is what
> >> was intended here as well.
> >>
> >> Before b2ed01e7ad3d the resource was taken off the bulk_move before the
> >> swapout; since then a swapped-out resource stays inside its BO's
> >> bulk_move range (and on the manager LRU) although it is unevictable.
> >> When it is later freed or the BO leaves the bulk_move
> >> (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()),
> >> ttm_resource_del_bulk_move() skips it because of its
> >> !ttm_resource_unevictable() guard, so a range endpoint in pos->first /
> >> pos->last is left pointing at freed memory. The next
> >> ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor
> >> is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(),
> >> "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL
> >> dereference in ttm_resource_manager_next() -- minutes to hours after a
> >> hibernation, or at process exit / reboot following one. Samuel
> >> Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the
> >> dangling cursor; the missing removal at swapout time is the reason it
> >> dangles.
> >>
> >> Testing the condition for success restores the removal. On an AMD
> >> Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on
> >> a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug
> >> crashed 5 of 18 hibernation cycles; a function profile of one
> >> hibernation showed 336 ttm_tt_swapout() calls and zero
> >> ttm_resource_del_bulk_move_unevictable() calls. With this change the
> >> removal happens for every swapped-out resource and 12 further cycles
> >> were clean.
> >>
> >> Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure")
> >> Cc: stable@vger.kernel.org # v7.1+
> >> Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387
> >> Link: https://lore.kernel.org/dri-devel/CAHYiNPa6aVacJoLOje-qZ1GyYx-9p0tN4NuP8D_eSL+UJeevXw@mail.gmail.com/
> >> Signed-off-by: Vadim Nikitushkin <bub4z0r@gmail.com>
> >> Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> >> Reviewed-by: Christian König <christian.koenig@amd.com>
> >> Signed-off-by: Christian König <christian.koenig@amd.com>
> >> Link: https://lore.kernel.org/r/20260909205028.13799-1-bub4z0r@gmail.com
> >> [ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed
> > 
> > Do not squash, send a patch series instead please.
> 
> in this particular case that squashing is the correct approach.
> 
> AMDs mail servers mangled the initial patch so badly that I had trouble applying it and ended up pushing a broken patch upstream. The squash is basically fixing that up.

Ok, I'll take this then, thanks for the response.

greg k-h

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-09-23  9:32 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
     [not found] <2026092253-dimness-unethical-2515@gregkh>
2026-09-22 15:33 ` [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range Vadim Nikitushkin
2026-09-23  7:41   ` Greg KH
2026-09-23  7:53     ` Christian König
2026-09-23  9:29       ` Greg KH

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox