From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 107152] GPU fault detected: 146 / VM_CONTEXT1_PROTECTION_FAULT / ring gfx timeout Date: Tue, 31 Jul 2018 21:41:14 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============1310216603==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [131.252.210.165]) by gabe.freedesktop.org (Postfix) with ESMTP id 1FC0A89C68 for ; Tue, 31 Jul 2018 21:41:14 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============1310216603== Content-Type: multipart/alternative; boundary="15330732740.39Be7E.21758" Content-Transfer-Encoding: 7bit --15330732740.39Be7E.21758 Date: Tue, 31 Jul 2018 21:41:14 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D107152 --- Comment #4 from dwagner --- Saw this kind of crash (still with the latest amd-staging-drm-next kernel) three times in a row today, just by playing a specific video immediately af= ter rebooting and starting X11 with mpv, before the 10 minute video ended. The video (which just shows a static cover image) can be obtained via: youtube-dl -f 248+251 'https://www.youtube.com/watch?v=3DkYKE78Pcjog' The log messages were just like reported above, I guess the additional "hw_= done or flip_done timed out" after the "GPU reset begin!" is not really relevant: Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0: GPU fault detected: 147 0x0f580402 for process Xorg pid 793 thread amdgpu_cs:0 pid 794 Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0:=20=20 VM_CONTEXT1_PROTECTION_FAULT_ADDR 0x0010C3EB Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0:=20=20 VM_CONTEXT1_PROTECTION_FAULT_STATUS 0x02004002 Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0: VM fault (0x02, vmid 1, pasid 32768) at page 1098731, read from 'TC3' (0x54433300) (4) Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0: GPU fault detected: 146 0x0c984424 for process Xorg pid 793 thread amdgpu_cs:0 pid 794 Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0:=20=20 VM_CONTEXT1_PROTECTION_FAULT_ADDR 0x00100193 Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0:=20=20 VM_CONTEXT1_PROTECTION_FAULT_STATUS 0x04044024 Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0: VM fault (0x24, vmid 2, pasid 32768) at page 1048979, read from 'TC1' (0x54433100) (68) Jul 31 22:20:25 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri= ng gfx timeout, signaled seq=3D26570, emitted seq=3D26573 Jul 31 22:20:25 ryzen kernel: amdgpu 0000:0a:00.0: GPU reset begin! Jul 31 22:20:35 ryzen kernel: [drm:amdgpu_dm_atomic_check [amdgpu]] *ERROR* [CRTC:44:crtc-0] hw_done or flip_done timed out --=20 You are receiving this mail because: You are the assignee for the bug.= --15330732740.39Be7E.21758 Date: Tue, 31 Jul 2018 21:41:14 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 4 on bug 10715= 2 from dwagner
Saw this kind of crash (still with the latest amd-staging-drm-=
next kernel)
three times in a row today, just by playing a specific video immediately af=
ter
rebooting and starting X11 with mpv, before the 10 minute video ended.
The video (which just shows a static cover image) can be obtained via:

youtube-dl -f 248+251 'https://www.youtube.com/watch?v=3DkYKE78Pcjog'

The log messages were just like reported above, I guess the additional &quo=
t;hw_done
or flip_done timed out" after the "GPU reset begin!" is not =
really relevant:

Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0: GPU fault detected: 147
0x0f580402 for process Xorg pid 793 thread amdgpu_cs:0 pid 794
Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0:=20=20
VM_CONTEXT1_PROTECTION_FAULT_ADDR   0x0010C3EB
Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0:=20=20
VM_CONTEXT1_PROTECTION_FAULT_STATUS 0x02004002
Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0: VM fault (0x02, vmid 1,
pasid 32768) at page 1098731, read from 'TC3' (0x54433300) (4)
Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0: GPU fault detected: 146
0x0c984424 for process Xorg pid 793 thread amdgpu_cs:0 pid 794
Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0:=20=20
VM_CONTEXT1_PROTECTION_FAULT_ADDR   0x00100193
Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0:=20=20
VM_CONTEXT1_PROTECTION_FAULT_STATUS 0x04044024
Jul 31 22:20:21 ryzen kernel: amdgpu 0000:0a:00.0: VM fault (0x24, vmid 2,
pasid 32768) at page 1048979, read from 'TC1' (0x54433100) (68)
Jul 31 22:20:25 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
gfx timeout, signaled seq=3D26570, emitted seq=3D26573
Jul 31 22:20:25 ryzen kernel: amdgpu 0000:0a:00.0: GPU reset begin!
Jul 31 22:20:35 ryzen kernel: [drm:amdgpu_dm_atomic_check [amdgpu]] *ERROR*
[CRTC:44:crtc-0] hw_done or flip_done timed out


You are receiving this mail because:
  • You are the assignee for the bug.
= --15330732740.39Be7E.21758-- --===============1310216603== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============1310216603==--