AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Lang Yu <Lang.Yu@amd.com>
To: Felix Kuehling <felix.kuehling@amd.com>
Cc: Harish Kasiviswanathan <Harish.Kasiviswanathan@amd.com>,
	amd-gfx@lists.freedesktop.org
Subject: Re: [PATCH] drm/amdkfd: Insert missing TLB flush on GFX10 and later
Date: Thu, 14 Sep 2023 13:58:53 +0800	[thread overview]
Message-ID: <ZQKhHeHjIEYP0qrH@lang-desktop> (raw)
In-Reply-To: <5300374c-3f89-b80a-622c-afca07eb0e16@amd.com>

On 09/13/ , Felix Kuehling wrote:
> On 2023-09-13 6:23, Lang Yu wrote:
> > On 09/12/ , Felix Kuehling wrote:
> > > On 2023-09-11 22:52, Lang Yu wrote:
> > > > On 09/11/ , Harish Kasiviswanathan wrote:
> > > > > Heavy-weight TLB flush is required after unmap on all GPUs for
> > > > > correctness and security.
> > > > > 
> > > > > Signed-off-by: Harish Kasiviswanathan<Harish.Kasiviswanathan@amd.com>
> > > > > ---
> > > > >    drivers/gpu/drm/amd/amdkfd/kfd_priv.h | 3 +--
> > > > >    1 file changed, 1 insertion(+), 2 deletions(-)
> > > > > 
> > > > > diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_priv.h b/drivers/gpu/drm/amd/amdkfd/kfd_priv.h
> > > > > index b315311dfe2a..b9950074aee0 100644
> > > > > --- a/drivers/gpu/drm/amd/amdkfd/kfd_priv.h
> > > > > +++ b/drivers/gpu/drm/amd/amdkfd/kfd_priv.h
> > > > > @@ -1466,8 +1466,7 @@ void kfd_flush_tlb(struct kfd_process_device *pdd, enum TLB_FLUSH_TYPE type);
> > > > >    static inline bool kfd_flush_tlb_after_unmap(struct kfd_dev *dev)
> > > > >    {
> > > > > -	return KFD_GC_VERSION(dev) == IP_VERSION(9, 4, 3) ||
> > > > > -	       KFD_GC_VERSION(dev) == IP_VERSION(9, 4, 2) ||
> > > > > +	return KFD_GC_VERSION(dev) > IP_VERSION(9, 4, 2) ||
> > > > >    	       (KFD_GC_VERSION(dev) == IP_VERSION(9, 4, 1) && dev->sdma_fw_version >= 18) ||
> > > > >    	       KFD_GC_VERSION(dev) == IP_VERSION(9, 4, 0);
> > > > >    }
> > > > 1, If TLB_FLUSH_HEAVYWEIGHT is required after unmap on all GPUs
> > > > as described in commmit message, why we have this whitelist
> > > > instead of a blacklist?
> > > That was a bug that this patch is fixing. There were specific GPUs and
> > > firmware versions where the TLB flush after unmap was causing intermittent
> > > problems in specific tests. This should have always been a blacklist.
> > > 
> > > 
> > > > 2, kfd_flush_tlb(pdd, TLB_FLUSH_HEAVYWEIGHT) is also called
> > > > in svm_range_unmap_from_gpus(). Why not add this whitelist there?
> > > There was a patch that used kfd_flush_tlb_after_unmap in the SVM code. But
> > > you reverted that patch, probably because it caused more problems than it
> > > solved. SVM really must flush TLBs the way it does because it is so tightly
> > > integrated with Linux's virtual memory management and because with XNACK,
> > > memory can be unmapped while GPU work is in progress without preemting
> > > queues (implicitly flushing TLBs and caches):
> > > 
> > > commit 515d7cebc2e2d2b4f0a276d26f3b790a83cdfe06
> > > Author: Lang Yu<Lang.Yu@amd.com>
> > > Date:   Wed Apr 20 10:24:31 2022 +0800
> > > 
> > >      Revert "drm/amdkfd: only allow heavy-weight TLB flush on some ASICs for SVM too"
> > >      This reverts commit 36bf93216ecbe399c40c5e0486f0f0e3a4afa69e.
> > >      It causes SVM regressions on Vega10 with XNACK-ON. Just revert it
> > >      at the moment.
> > >      ./kfdtest --gtest_filter=KFDSVMRangeTest.MigratePolicyTest
> > >      Signed-off-by: Lang Yu<Lang.Yu@amd.com>
> > >      Reviewed-by: Philip Yang<Philip.Yang@amd.com>
> > >      Signed-off-by: Alex Deucher<alexander.deucher@amd.com>
> > > 
> > > Regards,
> > >    Felix
> > Yes, that's because kfd_flush_tlb_after_unmap() return false for Vega10(gfx901).
> > kfd_flush_tlb(pdd, TLB_FLUSH_HEAVYWEIGHT) is called unconditionally in SVM
> > for ASICs > IP_VERSION(9, 0, 0) and works well.
> > 
> > So why not relax the condition to KFD_GC_VERSION(dev) > IP_VERSION(9, 0, 0) ?
> 
> That would reintroduce the same problem that this workaround was meant to
> fix. I don't remember all the details of this, as it was years ago. I
> believe it was an intermittent hang or VM fault that was somewhat difficult
> to reproduce and investigate. Maybe Eric remembers the details as he was
> working on this bug back then. However, it was a real issue, and we got an
> SDMA firmware fix for it on GFX IP 9.4.1 as you can see from the FW version
> check in kfd_flush_tlb_after_unmap.
> 
> I would not recommend reverting this workaround at the risk of reintroducing
> a known intermittent bug that affects stability.

Got it. Thank you!

Regards,
Lang

> Regards,
>   Felix
> 
> 
> 
> > 
> > Regards,
> > Lang
> > 
> > > > Regards,
> > > > Lang
> > > > 
> > > > > -- 
> > > > > 2.34.1
> > > > > 

  reply	other threads:[~2023-09-14  5:59 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2023-09-11 19:00 [PATCH] drm/amdkfd: Insert missing TLB flush on GFX10 and later Harish Kasiviswanathan
2023-09-12  0:55 ` Felix Kuehling
2023-09-12  2:52 ` Lang Yu
2023-09-13  0:48   ` Felix Kuehling
2023-09-13 10:23     ` Lang Yu
2023-09-13 13:46       ` Felix Kuehling
2023-09-14  5:58         ` Lang Yu [this message]
2023-09-13  8:25 ` Christian König

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ZQKhHeHjIEYP0qrH@lang-desktop \
    --to=lang.yu@amd.com \
    --cc=Harish.Kasiviswanathan@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=felix.kuehling@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox