From: "Kuehling, Felix" <felix.kuehling@amd.com>
To: Priya Hosur <Priya.Hosur@amd.com>,
amd-gfx@lists.freedesktop.org, Shaoyun.Liu@amd.com,
Lijo.Lazar@amd.com, Alexander.Deucher@amd.com,
Mario.Limonciello@amd.com, Christian.Koenig@amd.com
Cc: Pratik.Vishwakarma@amd.com, Veerabadhran.Gopalakrishnan@amd.com
Subject: Re: [PATCH 1/1] drm/amdkfd: Add TLB flush after MES queue eviction/suspension
Date: Wed, 12 Aug 2026 12:19:47 -0400 [thread overview]
Message-ID: <367307a7-c1cc-4296-93f2-de6aad690217@amd.com> (raw)
In-Reply-To: <20260812061122.435929-2-Priya.Hosur@amd.com>
On 2026-08-12 02:11, Priya Hosur wrote:
> MES (Micro Engine Scheduler) does not perform heavy-weight TLB invalidation
> after unmapping queues, unlike HWS which does this automatically. This causes
> a race condition where in-flight SDMA DMA descriptors can access memory that
> has been unmapped, leading to page faults and GPU queue hangs during SVM
> page migration.
>
> The issue manifests as KFDSVMRangeTest.MultiThreadMigrationTest/1 failures
> on gfx1151 (Ryzen AI MAX) with XNACK mode 1 enabled - the GPU compute queue
> hangs with packets submitted but never consumed.
>
> Add kfd_flush_tlb() with TLB_FLUSH_HEAVYWEIGHT in two MES code paths:
> 1. evict_process_queues_cpsch() - after MES removes queues
> 2. suspend_queues() - after MES suspends queues and mem_fence completes
>
> This ensures all in-flight memory accesses from unmapped queues are flushed
> before memory is freed or migrated.
>
> Testing on gfx1151 shows this reduces failure rate from 100% to approximately
> 7-10%. The residual failures require further investigation.
>
> Signed-off-by: Priya Hosur <Priya.Hosur@amd.com>
> ---
> .../gpu/drm/amd/amdkfd/kfd_device_queue_manager.c | 13 ++++++++++++-
> 1 file changed, 12 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> index a23384571193..5eb85290126e 100644
> --- a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> +++ b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> @@ -1450,6 +1450,14 @@ static int evict_process_queues_cpsch(struct device_queue_manager *dqm,
> dqm_evict_mqd_bo(dqm, q);
> }
>
> + /*
> + * Heavy-weight TLB flush after MES removes queues to ensure
> + * in-flight SDMA accesses complete before memory is freed/migrated.
> + * HWS does this automatically, MES does not.
I'm not sure why you call out SDMA specifically here. This affects
in-flight memory accesses from compute jobs as well. Just remove "SDMA"
from the comment. With that fixed, the patch is
Reviewed-by: Felix Kuehling <felix.kuehling@amd.com>
> + */
> + if (dqm->dev->kfd->shared_resources.enable_mes)
> + kfd_flush_tlb(pdd, TLB_FLUSH_HEAVYWEIGHT);
> +
> if (!dqm->dev->kfd->shared_resources.enable_mes) {
> pdd->last_evict_timestamp = get_jiffies_64();
> retval = execute_queues_cpsch(dqm,
> @@ -3736,8 +3744,11 @@ int suspend_queues(struct kfd_process *p,
> if (!per_device_suspended) {
> dqm_unlock(dqm);
> mutex_unlock(&p->event_mutex);
> - if (total_suspended)
> + if (total_suspended) {
> amdgpu_amdkfd_debug_mem_fence(dqm->dev->adev);
> + /* Heavy-weight TLB flush after MES suspends queues */
> + kfd_flush_tlb(pdd, TLB_FLUSH_HEAVYWEIGHT);
> + }
> continue;
> }
>
prev parent reply other threads:[~2026-08-12 16:19 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-12 6:11 [PATCH 0/1] drm/amdkfd: Fix SVM page migration hang on MES-based GPUs Priya Hosur
2026-08-12 6:11 ` [PATCH 1/1] drm/amdkfd: Add TLB flush after MES queue eviction/suspension Priya Hosur
2026-08-12 16:19 ` Kuehling, Felix [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=367307a7-c1cc-4296-93f2-de6aad690217@amd.com \
--to=felix.kuehling@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Christian.Koenig@amd.com \
--cc=Lijo.Lazar@amd.com \
--cc=Mario.Limonciello@amd.com \
--cc=Pratik.Vishwakarma@amd.com \
--cc=Priya.Hosur@amd.com \
--cc=Shaoyun.Liu@amd.com \
--cc=Veerabadhran.Gopalakrishnan@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.