From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 102322] System crashes after "[drm] IP block:gmc_v8_0 is hung!" / [drm] IP block:sdma_v3_0 is hung! Date: Wed, 14 Nov 2018 00:23:15 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0095406626==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [131.252.210.165]) by gabe.freedesktop.org (Postfix) with ESMTP id 4A80F6E417 for ; Wed, 14 Nov 2018 00:23:16 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0095406626== Content-Type: multipart/alternative; boundary="15421549963.481599B.6448" Content-Transfer-Encoding: 7bit --15421549963.481599B.6448 Date: Wed, 14 Nov 2018 00:23:16 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D102322 --- Comment #68 from dwagner --- Tested today's current amd-staging-drm-next git head, to see if there has b= een any improvement over the last two months. The bad news: The 3-fps-video-replay test still crashes the driver reproduc= ably after few minutes, as long as the default automatic power management is act= ive. The mediocre news: At least it looks as if the linux kernel now survives the driver crash to some extent, I found messages in the journal like this: Nov 14 00:59:36 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri= ng sdma0 timeout, signaled seq=3D22008, emitted seq=3D22010 Nov 14 00:59:36 ryzen kernel: [drm] GPU recovery disabled. Nov 14 00:59:37 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri= ng sdma1 timeout, signaled seq=3D107, emitted seq=3D109 Nov 14 00:59:37 ryzen kernel: [drm] GPU recovery disabled. Nov 14 00:59:40 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri= ng sdma0 timeout, signaled seq=3D22008, emitted seq=3D22010 Nov 14 00:59:40 ryzen kernel: [drm] GPU recovery disabled. Nov 14 00:59:41 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri= ng sdma1 timeout, signaled seq=3D107, emitted seq=3D109 ... and so on repeating for several minutes after the screen went blank. Will test tomorrow if this means I can now collect the diagnostics outputs = that were asked for earlier. Some good news: S3 suspends/resumes are working fine right now. There are s= ome scary messages emitted upon resume, but they do not seem to have bad consequences: [ 281.465654] [drm:emulated_link_detect [amdgpu]] *ERROR* Failed to read E= DID [ 281.490719] [drm:emulated_link_detect [amdgpu]] *ERROR* Failed to read E= DID [ 282.006225] [drm] Fence fallback timer expired on ring sdma0 [ 282.512879] [drm] Fence fallback timer expired on ring sdma0 [ 282.556651] [drm] UVD and UVD ENC initialized successfully. [ 282.657771] [drm] VCE initialized successfully. --=20 You are receiving this mail because: You are the assignee for the bug.= --15421549963.481599B.6448 Date: Wed, 14 Nov 2018 00:23:16 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Comme= nt # 68 on bug 10232= 2 from dwagner
Tested today's current amd-staging-drm-next git head, to see i=
f there has been
any improvement over the last two months.

The bad news: The 3-fps-video-replay test still crashes the driver reproduc=
ably
after few minutes, as long as the default automatic power management is act=
ive.

The mediocre news: At least it looks as if the linux kernel now survives the
driver crash to some extent, I found messages in the journal like this:

Nov 14 00:59:36 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma0 timeout, signaled seq=3D22008, emitted seq=3D22010
Nov 14 00:59:36 ryzen kernel: [drm] GPU recovery disabled.
Nov 14 00:59:37 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma1 timeout, signaled seq=3D107, emitted seq=3D109
Nov 14 00:59:37 ryzen kernel: [drm] GPU recovery disabled.
Nov 14 00:59:40 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma0 timeout, signaled seq=3D22008, emitted seq=3D22010
Nov 14 00:59:40 ryzen kernel: [drm] GPU recovery disabled.
Nov 14 00:59:41 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma1 timeout, signaled seq=3D107, emitted seq=3D109

... and so on repeating for several minutes after the screen went blank.

Will test tomorrow if this means I can now collect the diagnostics outputs =
that
were asked for earlier.

Some good news: S3 suspends/resumes are working fine right now. There are s=
ome
scary messages emitted upon resume, but they do not seem to have bad
consequences:

[  281.465654] [drm:emulated_link_detect [amdgpu]] *ERROR* Failed to read E=
DID
[  281.490719] [drm:emulated_link_detect [amdgpu]] *ERROR* Failed to read E=
DID
[  282.006225] [drm] Fence fallback timer expired on ring sdma0
[  282.512879] [drm] Fence fallback timer expired on ring sdma0
[  282.556651] [drm] UVD and UVD ENC initialized successfully.
[  282.657771] [drm] VCE initialized successfully.


You are receiving this mail because:
  • You are the assignee for the bug.
= --15421549963.481599B.6448-- --===============0095406626== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============0095406626==--