From mboxrd@z Thu Jan 1 00:00:00 1970 From: "Koenig, Christian" Subject: =?GB18030?B?UmU6ILvYuLSjuiC72Li0o7ogu9i4tKO6IEJ1ZzogYW1kZ3B1IGRybSBkcml2?= =?GB18030?B?ZXIgY2F1c2UgcHJvY2VzcyBpbnRvIERpc2sgc2xlZXAgc3RhdGU=?= Date: Fri, 6 Sep 2019 11:23:20 +0000 Message-ID: References: <2162676e-dbfa-a67d-248c-98e9eb2099c2@amd.com> <88a08dcc-2e95-9379-693f-2d3fd928aa11@amd.com> Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0792055804==" Return-path: In-Reply-To: Content-Language: en-US List-Id: Discussion list for AMD gfx List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: amd-gfx-bounces-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org Sender: "amd-gfx" To: yanhua <78666679-9uewiaClKEY@public.gmane.org>, amd-gfx Cc: "Deucher, Alexander" --===============0792055804== Content-Language: en-US Content-Type: multipart/alternative; boundary="_000_badd9ea16f78abbcbdbee11271188524amdcom_" --_000_badd9ea16f78abbcbdbee11271188524amdcom_ Content-Type: text/plain; charset="GB18030" Content-Transfer-Encoding: quoted-printable Are there anything I have missed ? Yeah, unfortunately quite a bunch of things. The fact that arm64 doesn't su= pport the PCIe NoSnoop TLP attribute is only the tip of the iceberg. You need a full "recent" driver stack, e.g. not older than a few month till= a year, for this to work. And not only the kernel, but also recent userspa= ce components. Maybe that's something you could first, e.g. install a recent version of Me= sa and/or tell Mesa to not use the SDMA at all. But since you are running i= nto an SDMA lockup with a kernel triggered page table update I see little c= hance that this work. The only other alternative I can see is the DKMS package of the pro-driver.= With that one you might be able to compile the recent driver for an older = kernel version. But I can't guarantee at all that this actually works on ARM64. Sorry that I don't have better news for you, Christian. Am 05.09.19 um 03:36 schrieb yanhua: Hi, Christian, I noticed that you said 'amdgpu is known to not work on arm64 unti= l very recently'. I found the CPU related commit with drm is "drm: disab= le uncached DMA optimization for ARM and arm64". @@ -47,6 +47,24 @@ static inline bool drm_arch_can_wc_memory(void) return false; #elif defined(CONFIG_MIPS) && defined(CONFIG_CPU_LOONGSON3) return false; +#elif defined(CONFIG_ARM) || defined(CONFIG_ARM64) + /* + * The DRM driver stack is designed to work with cache coherent dev= ices + * only, but permits an optimization to be enabled in some cases, w= here + * for some buffers, both the CPU and the GPU use uncached mappings= , + * removing the need for DMA snooping and allocation in the CPU cac= hes. + * + * The use of uncached GPU mappings relies on the correct implement= ation + * of the PCIe NoSnoop TLP attribute by the platform, otherwise the= GPU + * will use cached mappings nonetheless. On x86 platforms, this doe= s not + * seem to matter, as uncached CPU mappings will snoop the caches i= n any + * case. However, on ARM and arm64, enabling this optimization on a + * platform where NoSnoop is ignored results in loss of coherency, = which + * breaks correct operation of the device. Since we have no way of + * detecting whether NoSnoop works or not, just disable this + * optimization entirely for ARM and arm64. + */ + return false; #else return true; #endif The real effect is to in amdgpu_object.c if (!drm_arch_can_wc_memory()) bo->flags &=3D ~AMDGPU_GEM_CREATE_CPU_GTT_USWC; And we have AMDGPU_GEM_CREATE_CPU_GTT_USWC turned off in our 4.19.36 kernel= , So I think this is not the cause of my bug. Are there anything I have m= issed ? I had suggest the machine supplier to use a more newer kernel such as 5.2.2= , But they failed to do so after some try. We also backport a series patch= es from newer kernel. But still we get the bad ring timeout. We have dived into the amdgpu drm driver a long time, bu it is really diffi= cult for me, especially the hardware related ring timeout. ------------------ Yanhua ------------------ =D4=AD=CA=BC=D3=CA=BC=FE ------------------ =B7=A2=BC=FE=C8=CB: "Koenig, Christian"; =B7=A2=CB=CD=CA=B1=BC=E4: 2019=C4=EA9=D4=C23=C8=D5(=D0=C7=C6=DA=B6=FE) =CD= =ED=C9=CF9:19 =CA=D5=BC=FE=C8=CB: "yanhua"<78666679-9uewiaClKEY@public.gmane.org>;"amd-= gfx"; =B3=AD=CB=CD: "Deucher, Alexander"; =D6=F7=CC=E2: Re: =BB=D8=B8=B4=A3=BA =BB=D8=B8=B4=A3=BA Bug: amdgpu drm dri= ver cause process into Disk sleep state This is just a GPU lock, please open up a bug report on freedesktop.org and= attach the full dmesg and which version of Mesa you are using. Regards, Christian. Am 03.09.19 um 15:16 schrieb 78666679: Yes, with dmesg|grep drm , I get following. 348571.880718] [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ring sdma1 timeou= t, signaled seq=3D24423862, emitted seq=3D24423865 ------------------ =D4=AD=CA=BC=D3=CA=BC=FE ------------------ =B7=A2=BC=FE=C8=CB: "Koenig, Christian"; =B7=A2=CB=CD=CA=B1=BC=E4: 2019=C4=EA9=D4=C23=C8=D5(=D0=C7=C6=DA=B6=FE) =CD= =ED=C9=CF9:07 =CA=D5=BC=FE=C8=CB: ""<78666679-9uewiaClKEY@public.gmane.org>;"amd-gfx"; =B3=AD=CB=CD: "Deucher, Alexander"; =D6=F7=CC=E2: Re: =BB=D8=B8=B4=A3=BA Bug: amdgpu drm driver cause process i= nto Disk sleep state Well that looks like the hardware got stuck. Do you get something in the locks about a timeout on the SDMA ring? Regards, Christian. Am 03.09.19 um 14:50 schrieb 78666679: Hi Christian, Sometimes the thread blocked disk sleeping in call to amdgpu_sa_bo_= new. following is the stack trace. it seems the sa bo is used up , so th= e caller blocked waiting someone to free sa resources. D 206833 227656 [surfaceflinger] Binder:45_5 cat /proc/206833/task/227656/stack [<0>] __switch_to+0x94/0xe8 [<0>] dma_fence_wait_any_timeout+0x234/0x2d0 [<0>] amdgpu_sa_bo_new+0x468/0x540 [amdgpu] [<0>] amdgpu_ib_get+0x60/0xc8 [amdgpu] [<0>] amdgpu_job_alloc_with_ib+0x70/0xb0 [amdgpu] [<0>] amdgpu_vm_bo_update_mapping+0x2e0/0x3d8 [amdgpu] [<0>] amdgpu_vm_bo_update+0x2a0/0x710 [amdgpu] [<0>] amdgpu_gem_va_ioctl+0x46c/0x4c8 [amdgpu] [<0>] drm_ioctl_kernel+0x94/0x118 [drm] [<0>] drm_ioctl+0x1f0/0x438 [drm] [<0>] amdgpu_drm_ioctl+0x58/0x90 [amdgpu] [<0>] do_vfs_ioctl+0xc4/0x8c0 [<0>] ksys_ioctl+0x8c/0xa0 [<0>] __arm64_sys_ioctl+0x28/0x38 [<0>] el0_svc_common+0xa0/0x180 [<0>] el0_svc_handler+0x38/0x78 [<0>] el0_svc+0x8/0xc [<0>] 0xffffffffffffffff -------------------- YanHua ------------------ =D4=AD=CA=BC=D3=CA=BC=FE ------------------ =B7=A2=BC=FE=C8=CB: "Koenig, Christian"; =B7=A2=CB=CD=CA=B1=BC=E4: 2019=C4=EA9=D4=C23=C8=D5(=D0=C7=C6=DA=B6=FE) =CF= =C2=CE=E74:21 =CA=D5=BC=FE=C8=CB: ""<78666679-9uewiaClKEY@public.gmane.org>;"amd-gfx"; =B3=AD=CB=CD: "Deucher, Alexander"; =D6=F7=CC=E2: Re: Bug: amdgpu drm driver cause process into Disk sleep stat= e Hi Yanhua, please update your kernel first, cause that looks like a known issue which was recently fixed by patch "drm/scheduler: use job count instead of peek". Probably best to try the latest bleeding edge kernel and if that doesn't help please open up a bug report on https://bugs.freedesktop.org/. Regards, Christian. Am 03.09.19 um 09:35 schrieb 78666679: > Hi, Sirs: > I have a wx5100 amdgpu card, It randomly come into failure. some= times, it will cause processes into uninterruptible wait state. > > > cps-new-ondemand-0587:~ # ps aux|grep -w D > root 11268 0.0 0.0 260628 3516 ? Ssl 8=D4=C226 0:00 /us= r/sbin/gssproxy -D > root 136482 0.0 0.0 212500 572 pts/0 S+ 15:25 0:00 grep --= color=3Dauto -w D > root 370684 0.0 0.0 17972 7428 ? Ss 9=D4=C202 0:04 /us= r/sbin/sshd -D > 10066 432951 0.0 0.0 0 0 ? D 9=D4=C202 0:00 [Fa= keFinalizerDa] > root 496774 0.0 0.0 0 0 ? D 9=D4=C202 0:17 [kw= orker/8:1+eve] > cps-new-ondemand-0587:~ # cat /proc/496774/stack > [<0>] __switch_to+0x94/0xe8 > [<0>] drm_sched_entity_flush+0xf8/0x248 [gpu_sched] > [<0>] amdgpu_ctx_mgr_entity_flush+0xac/0x148 [amdgpu] > [<0>] amdgpu_flush+0x2c/0x50 [amdgpu] > [<0>] filp_close+0x40/0xa0 > [<0>] put_files_struct+0x118/0x120 > [<0>] put_files_struct+0x30/0x68 [binder_linux] > [<0>] binder_deferred_func+0x4d4/0x658 [binder_linux] > [<0>] process_one_work+0x1b4/0x3f8 > [<0>] worker_thread+0x54/0x470 > [<0>] kthread+0x134/0x138 > [<0>] ret_from_fork+0x10/0x18 > [<0>] 0xffffffffffffffff > > > > This issue troubled me a long time. looking eagerly to get help from you= ! > > > ----- > Yanhua --_000_badd9ea16f78abbcbdbee11271188524amdcom_ Content-Type: text/html; charset="GB18030" Content-ID: <785C48D37DB0584AB60F9C220A74364E-asWib9pRmPqcE4WynfumptQqCkab/8FMAL8bYrjMMd8@public.gmane.org> Content-Transfer-Encoding: quoted-printable
Are there anything I have missed ?

Yeah, unfortunately quite a bunch of things. The fact that arm64 doesn't su= pport the PCIe NoSnoop TLP attribute is only the tip of the iceberg.

You need a full "recent" driver stack, e.g. not older than a few = month till a year, for this to work. And not only the kernel, but also rece= nt userspace components.

Maybe that's something you could first, e.g. install a recent version of Me= sa and/or tell Mesa to not use the SDMA at all. But since you are running i= nto an SDMA lockup with a kernel triggered page table update I see little c= hance that this work.

The only other alternative I can see is the DKMS package of the pro-driver.= With that one you might be able to compile the recent driver for an older = kernel version.

But I can't guarantee at all that this actually works on ARM64.

Sorry that I don't have better news for you,
Christian.

Am 05.09.19 um 03:36 schrieb yanhua:
Hi, Christian,
        I noticed that you said&nbs= p; 'amdgpu is known to not work on arm64 until very recently'.   = I found the CPU related commit with drm is "drm: disable uncached DMA= optimization for ARM and arm64". 
@@ -47,6 +47,24 @@ static inline bool drm_arch_can_wc_memory(void)=
        return false;
 #elif defined(CONFIG_MIPS) && defined(CONFIG_CPU_LOONGSON3)         return false;
+#elif defined(CONFIG_ARM) || defined(CONFIG_ARM64)
+       /*
+        * The DRM driver stack is d= esigned to work with cache coherent devices
+        * only, but permits an opti= mization to be enabled in some cases, where
+        * for some buffers, both th= e CPU and the GPU use uncached mappings,
+        * removing the need for DMA= snooping and allocation in the CPU caches.
+        *
+        * The use of uncached GPU m= appings relies on the correct implementation
+        * of the PCIe NoSnoop TLP a= ttribute by the platform, otherwise the GPU
+        * will use cached mappings = nonetheless. On x86 platforms, this does not
+        * seem to matter, as uncach= ed CPU mappings will snoop the caches in any
+        * case. However, on ARM and= arm64, enabling this optimization on a
+        * platform where NoSnoop is= ignored results in loss of coherency, which
+        * breaks correct operation = of the device. Since we have no way of
+        * detecting whether NoSnoop= works or not, just disable this
+        * optimization entirely for= ARM and arm64.
+        */
+       return false;
 #else
        return true;
 #endif

The real effect is to  in amdgpu_object.c

   if (!drm_arch_can_wc_memory())
            &nb= sp;   bo->flags &=3D ~AMDGPU_GEM_CREATE_CPU_GTT_USWC;

And we have AMDGPU_GEM_= CREATE_CPU_GTT_USWC turned off in our 4.19.36 kernel, So I think this is no= t  the cause of my bug.  Are there anything I have missed ?

I had suggest the machine supplier to use a more newer kernel such as = 5.2.2, But they failed to do so after some try.  We also backport a se= ries patches from newer kernel. But still we get the bad ring timeout.

We have dived into the amdgpu drm driver a long time, bu it is really = difficult for me, especially the hardware related ring timeout.

------------------
Yanhua

------------------ =D4=AD=CA=BC=D3=CA=BC=FE ------------------
=B7=A2=BC=FE=C8=CB: "Koenig, Christian"<Chr= istian.Koenig-5C7GfCeVMHo@public.gmane.org>;
=B7=A2=CB=CD=CA=B1=BC=E4: 2019=C4=EA9=D4=C23=C8=D5(=D0=C7= =C6=DA=B6=FE) =CD=ED=C9=CF9:19
=CA=D5=BC=FE=C8=CB: "yanhua"<78666679-9uewiaClKEY@public.gmane.org>;= "amd-gfx"<amd-gfx-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org>;
=B3=AD=CB=CD: "Deucher, Alexander"<Alexande= r.Deucher-5C7GfCeVMHo@public.gmane.org>;
=D6=F7=CC=E2: Re: =BB=D8=B8=B4=A3=BA =BB=D8=B8=B4=A3=BA Bu= g: amdgpu drm driver cause process into Disk sleep state

This is just a GPU lock, please open up a bu= g report on freedesktop.org and attach the full dmesg and which version of = Mesa you are using.

Regards,
Christian.

Am 03.09.19 um 15:16 schrieb 78666679:
Yes, with dmesg|grep drm ,  I get following.

348571.880718] [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ring sdma1 t= imeout, signaled seq=3D24423862, emitted seq=3D24423865


------------------ =D4=AD=CA=BC=D3=CA=BC=FE ------------------
=B7=A2=BC=FE=C8=CB: "Koenig, Christian"<Christian.Koenig-5C7GfCeVMHo@public.gmane.org>;
=B7=A2=CB=CD=CA=B1=BC=E4: 2019=C4=EA9=D4=C23=C8=D5(=D0=C7= =C6=DA=B6=FE) =CD=ED=C9=CF9:07
=B3=AD=CB=CD: "Deucher, Alexander"<Alexander.Deucher-5C7GfCeVMHo@public.gmane.org>;
=D6=F7=CC=E2: Re: =BB=D8=B8=B4=A3=BA Bug: amdgpu drm drive= r cause process into Disk sleep state

Well that looks like the hardware got stuck.=

Do you get something in the locks about a timeout on the SDMA ring?

Regards,
Christian.

Am 03.09.19 um 14:50 schrieb 78666679:
Hi Christian,
       Sometimes the thread blocked  disk sle= eping in call to amdgpu_sa_bo_new. following is the stack trace.  it s= eems the sa bo is used up ,  so  the caller blocked waiting someo= ne to free sa resources.

D 206833 227656 [surfaceflinger] <defunct> Binder:45_5
cat /proc/206833/task/227656/stack

[<0>] __switch_to+0x94/0xe8
[<0>] dma_fence_wait_any_timeout+0x234/0x2d0
[<0>] amdgpu_sa_bo_new+0x468/0x540 [amdgpu]
[<0>] amdgpu_ib_get+0x60/0xc8 [amdgpu]
[<0>] amdgpu_job_alloc_with_ib+0x70/0xb0 [amdgpu]
[<0>] amdgpu_vm_bo_update_mapping+0x2e0/0x3d8 [amdgpu]
[<0>] amdgpu_vm_bo_update+0x2a0/0x710 [amdgpu]
[<0>] amdgpu_gem_va_ioctl+0x46c/0x4c8 [amdgpu]
[<0>] drm_ioctl_kernel+0x94/0x118 [drm]
[<0>] drm_ioctl+0x1f0/0x438 [drm]
[<0>] amdgpu_drm_ioctl+0x58/0x90 [amdgpu]
[<0>] do_vfs_ioctl+0xc4/0x8c0
[<0>] ksys_ioctl+0x8c/0xa0
[<0>] __arm64_sys_ioctl+0x28/0x38
[<0>] el0_svc_common+0xa0/0x180
[<0>] el0_svc_handler+0x38/0x78
[<0>] el0_svc+0x8/0xc
[<0>] 0xffffffffffffffff


--------------------
YanHua

------------------ =D4=AD=CA=BC=D3=CA=BC=FE ------------------
=B7=A2=BC=FE=C8=CB: "Koenig, Christian"<Christian.Koenig-5C7GfCeVMHo@public.gmane.org>;
=B7=A2=CB=CD=CA=B1=BC=E4: 2019=C4=EA9=D4=C23=C8=D5(=D0=C7= =C6=DA=B6=FE) =CF=C2=CE=E74:21
=B3=AD=CB=CD: "Deucher, Alexander"<Alexander.Deucher-5C7GfCeVMHo@public.gmane.org>;
=D6=F7=CC=E2: Re: Bug: amdgpu drm driver cause process int= o Disk sleep state

Hi Yanhua,

please update your kernel first, cause that looks like a known issue
which was recently fixed by patch "drm/scheduler: use job count instea= d
of peek".

Probably best to try the latest bleeding edge kernel and if that doesn't help please open up a bug report on https://bugs.freedesktop.org/.

Regards,
Christian.

Am 03.09.19 um 09:35 schrieb 78666679:
> Hi, Sirs:
>         I have a wx5100 amdgpu card, It = randomly come into failure.  sometimes, it will cause processes into u= ninterruptible wait state.
>
>
> cps-new-ondemand-0587:~ # ps aux|grep -w D
> root      11268  0.0  0.0 260628  3516 ?=         Ssl  8=D4=C226   0:00 /usr/sbin/= gssproxy -D
> root     136482  0.0  0.0 212500   572 pts/0&= nbsp;   S+   15:25   0:00 grep --color=3Dauto -w D
> root     370684  0.0  0.0  17972  7428 ?=         Ss   9=D4=C202   0:04 /usr/sbin/= sshd -D
> 10066    432951  0.0  0.0      0 &n= bsp;   0 ?        D    9=D4=C202 &n= bsp; 0:00 [FakeFinalizerDa]
> root     496774  0.0  0.0      0 &n= bsp;   0 ?        D    9=D4=C202 &n= bsp; 0:17 [kworker/8:1+eve]
> cps-new-ondemand-0587:~ # cat /proc/496774/stack
> [<0>] __switch_to+0x94/0xe8
> [<0>] drm_sched_entity_flush+0xf8/0x248 [gpu_sched]
> [<0>] amdgpu_ctx_mgr_entity_flush+0xac/0x148 [amdgpu]
> [<0>] amdgpu_flush+0x2c/0x50 [amdgpu]
> [<0>] filp_close+0x40/0xa0
> [<0>] put_files_struct+0x118/0x120
> [<0>] put_files_struct+0x30/0x68 [binder_linux]
> [<0>] binder_deferred_func+0x4d4/0x658 [binder_linux]
> [<0>] process_one_work+0x1b4/0x3f8
> [<0>] worker_thread+0x54/0x470
> [<0>] kthread+0x134/0x138
> [<0>] ret_from_fork+0x10/0x18
> [<0>] 0xffffffffffffffff
>
>
>
> This issue troubled me a long time.  looking eagerly to get help = from you!
>
>
> -----
> Yanhua




--_000_badd9ea16f78abbcbdbee11271188524amdcom_-- --===============0792055804== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KYW1kLWdmeCBt YWlsaW5nIGxpc3QKYW1kLWdmeEBsaXN0cy5mcmVlZGVza3RvcC5vcmcKaHR0cHM6Ly9saXN0cy5m cmVlZGVza3RvcC5vcmcvbWFpbG1hbi9saXN0aW5mby9hbWQtZ2Z4 --===============0792055804==--