dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Greg KH <gregkh@linuxfoundation.org>
To: "Christian König" <christian.koenig@amd.com>
Cc: Vadim Nikitushkin <bub4z0r@gmail.com>,
	stable@vger.kernel.org, thomas.hellstrom@linux.intel.com,
	dri-devel@lists.freedesktop.org, skainsworth@gmail.com
Subject: Re: [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range
Date: Wed, 23 Sep 2026 11:29:44 +0200	[thread overview]
Message-ID: <2026092332-alkaline-overthrow-494f@gregkh> (raw)
In-Reply-To: <32446a55-070d-4739-bd3a-d56c4b628e29@amd.com>

On Wed, Sep 23, 2026 at 09:53:54AM +0200, Christian König wrote:
> Hi Greg,
> 
> On 9/23/26 09:41, Greg KH wrote:
> > On Tue, Sep 22, 2026 at 06:33:34PM +0300, Vadim Nikitushkin wrote:
> >> commit 3db7d7d583419f7b1f2e141e36418802dbb25cf8 upstream.
> >>
> >> ttm_tt_swapout() returns the number of pages swapped out on success and
> >> a negative error code on failure; for a populated ttm it never returns
> >> zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU
> >> walk on swapout failure") moved the bulk_move bookkeeping in
> >> ttm_bo_swapout_cb() under "if (!ret)", so the
> >> ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail()
> >> pair is now skipped on every successful swapout. The equivalent change
> >> for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink()
> >> infinite LRU walk on backup failure") tests "lret > 0", which is what
> >> was intended here as well.
> >>
> >> Before b2ed01e7ad3d the resource was taken off the bulk_move before the
> >> swapout; since then a swapped-out resource stays inside its BO's
> >> bulk_move range (and on the manager LRU) although it is unevictable.
> >> When it is later freed or the BO leaves the bulk_move
> >> (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()),
> >> ttm_resource_del_bulk_move() skips it because of its
> >> !ttm_resource_unevictable() guard, so a range endpoint in pos->first /
> >> pos->last is left pointing at freed memory. The next
> >> ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor
> >> is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(),
> >> "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL
> >> dereference in ttm_resource_manager_next() -- minutes to hours after a
> >> hibernation, or at process exit / reboot following one. Samuel
> >> Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the
> >> dangling cursor; the missing removal at swapout time is the reason it
> >> dangles.
> >>
> >> Testing the condition for success restores the removal. On an AMD
> >> Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on
> >> a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug
> >> crashed 5 of 18 hibernation cycles; a function profile of one
> >> hibernation showed 336 ttm_tt_swapout() calls and zero
> >> ttm_resource_del_bulk_move_unevictable() calls. With this change the
> >> removal happens for every swapped-out resource and 12 further cycles
> >> were clean.
> >>
> >> Fixes: b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure")
> >> Cc: stable@vger.kernel.org # v7.1+
> >> Closes: https://gitlab.freedesktop.org/drm/amd/-/issues/5387
> >> Link: https://lore.kernel.org/dri-devel/CAHYiNPa6aVacJoLOje-qZ1GyYx-9p0tN4NuP8D_eSL+UJeevXw@mail.gmail.com/
> >> Signed-off-by: Vadim Nikitushkin <bub4z0r@gmail.com>
> >> Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> >> Reviewed-by: Christian König <christian.koenig@amd.com>
> >> Signed-off-by: Christian König <christian.koenig@amd.com>
> >> Link: https://lore.kernel.org/r/20260909205028.13799-1-bub4z0r@gmail.com
> >> [ Squashed with commit fcfe64715b425262af1b36f498f9197f3537ceed
> > 
> > Do not squash, send a patch series instead please.
> 
> in this particular case that squashing is the correct approach.
> 
> AMDs mail servers mangled the initial patch so badly that I had trouble applying it and ended up pushing a broken patch upstream. The squash is basically fixing that up.

Ok, I'll take this then, thanks for the response.

greg k-h

      reply	other threads:[~2026-09-23  9:32 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <2026092253-dimness-unethical-2515@gregkh>
2026-09-22 15:33 ` [PATCH 7.2.y] drm/ttm: fix swapped-out resources never leaving their bulk_move range Vadim Nikitushkin
2026-09-23  7:41   ` Greg KH
2026-09-23  7:53     ` Christian König
2026-09-23  9:29       ` Greg KH [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=2026092332-alkaline-overthrow-494f@gregkh \
    --to=gregkh@linuxfoundation.org \
    --cc=bub4z0r@gmail.com \
    --cc=christian.koenig@amd.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=skainsworth@gmail.com \
    --cc=stable@vger.kernel.org \
    --cc=thomas.hellstrom@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox