From mboxrd@z Thu Jan 1 00:00:00 1970
From: bugzilla-daemon@freedesktop.org
Subject: [Bug 102322] System crashes after "[drm] IP block:gmc_v8_0 is hung!"
/ [drm] IP block:sdma_v3_0 is hung!
Date: Wed, 14 Nov 2018 00:23:15 +0000
Message-ID:
References:
Mime-Version: 1.0
Content-Type: multipart/mixed; boundary="===============0095406626=="
Return-path:
Received: from culpepper.freedesktop.org (culpepper.freedesktop.org
[131.252.210.165])
by gabe.freedesktop.org (Postfix) with ESMTP id 4A80F6E417
for ; Wed, 14 Nov 2018 00:23:16 +0000 (UTC)
In-Reply-To:
List-Unsubscribe: ,
List-Archive:
List-Post:
List-Help:
List-Subscribe: ,
Errors-To: dri-devel-bounces@lists.freedesktop.org
Sender: "dri-devel"
To: dri-devel@lists.freedesktop.org
List-Id: dri-devel@lists.freedesktop.org
--===============0095406626==
Content-Type: multipart/alternative; boundary="15421549963.481599B.6448"
Content-Transfer-Encoding: 7bit
--15421549963.481599B.6448
Date: Wed, 14 Nov 2018 00:23:16 +0000
MIME-Version: 1.0
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
X-Bugzilla-URL: http://bugs.freedesktop.org/
Auto-Submitted: auto-generated
https://bugs.freedesktop.org/show_bug.cgi?id=3D102322
--- Comment #68 from dwagner ---
Tested today's current amd-staging-drm-next git head, to see if there has b=
een
any improvement over the last two months.
The bad news: The 3-fps-video-replay test still crashes the driver reproduc=
ably
after few minutes, as long as the default automatic power management is act=
ive.
The mediocre news: At least it looks as if the linux kernel now survives the
driver crash to some extent, I found messages in the journal like this:
Nov 14 00:59:36 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma0 timeout, signaled seq=3D22008, emitted seq=3D22010
Nov 14 00:59:36 ryzen kernel: [drm] GPU recovery disabled.
Nov 14 00:59:37 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma1 timeout, signaled seq=3D107, emitted seq=3D109
Nov 14 00:59:37 ryzen kernel: [drm] GPU recovery disabled.
Nov 14 00:59:40 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma0 timeout, signaled seq=3D22008, emitted seq=3D22010
Nov 14 00:59:40 ryzen kernel: [drm] GPU recovery disabled.
Nov 14 00:59:41 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma1 timeout, signaled seq=3D107, emitted seq=3D109
... and so on repeating for several minutes after the screen went blank.
Will test tomorrow if this means I can now collect the diagnostics outputs =
that
were asked for earlier.
Some good news: S3 suspends/resumes are working fine right now. There are s=
ome
scary messages emitted upon resume, but they do not seem to have bad
consequences:
[ 281.465654] [drm:emulated_link_detect [amdgpu]] *ERROR* Failed to read E=
DID
[ 281.490719] [drm:emulated_link_detect [amdgpu]] *ERROR* Failed to read E=
DID
[ 282.006225] [drm] Fence fallback timer expired on ring sdma0
[ 282.512879] [drm] Fence fallback timer expired on ring sdma0
[ 282.556651] [drm] UVD and UVD ENC initialized successfully.
[ 282.657771] [drm] VCE initialized successfully.
--=20
You are receiving this mail because:
You are the assignee for the bug.=
--15421549963.481599B.6448
Date: Wed, 14 Nov 2018 00:23:16 +0000
MIME-Version: 1.0
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
X-Bugzilla-URL: http://bugs.freedesktop.org/
Auto-Submitted: auto-generated
Comme=
nt # 68
on bug 10232=
2
from dwagner
Tested today's current amd-staging-drm-next git head, to see i=
f there has been
any improvement over the last two months.
The bad news: The 3-fps-video-replay test still crashes the driver reproduc=
ably
after few minutes, as long as the default automatic power management is act=
ive.
The mediocre news: At least it looks as if the linux kernel now survives the
driver crash to some extent, I found messages in the journal like this:
Nov 14 00:59:36 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma0 timeout, signaled seq=3D22008, emitted seq=3D22010
Nov 14 00:59:36 ryzen kernel: [drm] GPU recovery disabled.
Nov 14 00:59:37 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma1 timeout, signaled seq=3D107, emitted seq=3D109
Nov 14 00:59:37 ryzen kernel: [drm] GPU recovery disabled.
Nov 14 00:59:40 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma0 timeout, signaled seq=3D22008, emitted seq=3D22010
Nov 14 00:59:40 ryzen kernel: [drm] GPU recovery disabled.
Nov 14 00:59:41 ryzen kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ri=
ng
sdma1 timeout, signaled seq=3D107, emitted seq=3D109
... and so on repeating for several minutes after the screen went blank.
Will test tomorrow if this means I can now collect the diagnostics outputs =
that
were asked for earlier.
Some good news: S3 suspends/resumes are working fine right now. There are s=
ome
scary messages emitted upon resume, but they do not seem to have bad
consequences:
[ 281.465654] [drm:emulated_link_detect [amdgpu]] *ERROR* Failed to read E=
DID
[ 281.490719] [drm:emulated_link_detect [amdgpu]] *ERROR* Failed to read E=
DID
[ 282.006225] [drm] Fence fallback timer expired on ring sdma0
[ 282.512879] [drm] Fence fallback timer expired on ring sdma0
[ 282.556651] [drm] UVD and UVD ENC initialized successfully.
[ 282.657771] [drm] VCE initialized successfully.
You are receiving this mail because:
- You are the assignee for the bug.
=
--15421549963.481599B.6448--
--===============0095406626==
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
Content-Disposition: inline
X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs
IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz
dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg==
--===============0095406626==--