From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 111763] ring_gfx hangs/freezes on Navi gpus Date: Tue, 05 Nov 2019 16:28:03 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0249081278==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [IPv6:2610:10:20:722:a800:ff:fe98:4b55]) by gabe.freedesktop.org (Postfix) with ESMTP id E486A6EAF3 for ; Tue, 5 Nov 2019 16:28:03 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0249081278== Content-Type: multipart/alternative; boundary="15729712838.d0577D.7169" Content-Transfer-Encoding: 7bit --15729712838.d0577D.7169 Date: Tue, 5 Nov 2019 16:28:03 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D111763 --- Comment #24 from wychuchol --- (In reply to wychuchol from comment #23) > (In reply to wychuchol from comment #19) > > After some time in Witcher 3 GOTY run with Lutris PC restarts on it's o= wn. I > > thought something is overheating (I've noticed graphic card memory in > > PSensor sometimes reaching 90 so I thought maybe that's what's happenin= g) > > but I investigated kern.log and this always happened before that autono= mous > > reset: > >=20 > > Nov 2 22:01:53 pop-os kernel: [ 979.244964] pcieport 0000:00:01.1: AE= R: > > Corrected error received: 0000:01:00.0 > > Nov 2 22:01:53 pop-os kernel: [ 979.244967] nvme 0000:01:00.0: AER: P= CIe > > Bus Error: severity=3DCorrected, type=3DData Link Layer, (Transmitter I= D) > > Nov 2 22:01:53 pop-os kernel: [ 979.244968] nvme 0000:01:00.0: AER:= =20=20 > > device [1987:5012] error status/mask=3D00001000/00006000 > > Nov 2 22:01:53 pop-os kernel: [ 979.244968] nvme 0000:01:00.0: AER:= =20=20=20 > > [12] Timeout=20=20=20=20=20=20=20=20=20=20=20=20=20=20=20 > > Nov 2 22:01:53 pop-os kernel: [ 979.262629] Emergency Sync complete >=20 > Thing with those AER errors is that they can go on and on and reset happe= ns > few minutes after the last logged error.=20 > This might be overheating, I managed to find how to output sensors readin= gs > into txt log and found that memory went up to 96 C (or rather it stayed > there for about 1m 10s) > Last reading before reset: > amdgpu-pci-2800 > Adapter: PCI adapter > vddgfx: +1.16 V=20=20 > fan1: 1551 RPM (min =3D 0 RPM, max =3D 3200 RPM) > edge: +74.0=C2=B0C (crit =3D +118.0=C2=B0C, hyst =3D -273.1=C2= =B0C) > (emerg =3D +99.0=C2=B0C) > junction: +88.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0= C) > (emerg =3D +99.0=C2=B0C) > mem: +96.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0= C) > (emerg =3D +99.0=C2=B0C) > power1: 162.00 W (cap =3D 195.00 W) >=20 > k10temp-pci-00c3 > Adapter: PCI adapter > Tdie: +70.5=C2=B0C (high =3D +70.0=C2=B0C) > Tctl: +70.5=C2=B0C=20=20 >=20 > Now the weird thing is - if this is in fact overheating why fan didn't go > beyond 1600 rpm even once.... Highest was like 1581 rpm and I don't have > silent bios switched on (sapphire pulse rx 5700 xt, lever facing away from > video ports). Okay I don't think it's overheating anymore. I found a moment in Anomaly 1.= 5.0 I can't get past without system resetting, just before a psi storm in Army Warehouses (I can provide a savefile). Last sensors reading before crash (5 second increments): amdgpu-pci-2800 Adapter: PCI adapter vddgfx: +1.01 V=20=20 fan1: 1560 RPM (min =3D 0 RPM, max =3D 3200 RPM) edge: +69.0=C2=B0C (crit =3D +118.0=C2=B0C, hyst =3D -273.1=C2=B0C) (emerg =3D +99.0=C2=B0C) junction: +84.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0C) (emerg =3D +99.0=C2=B0C) mem: +80.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0C) (emerg =3D +99.0=C2=B0C) power1: 227.00 W (cap =3D 195.00 W) k10temp-pci-00c3 Adapter: PCI adapter Tdie: +71.8=C2=B0C (high =3D +70.0=C2=B0C) Tctl: +71.8=C2=B0C --=20 You are receiving this mail because: You are the assignee for the bug.= --15729712838.d0577D.7169 Date: Tue, 5 Nov 2019 16:28:03 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Comme= nt # 24 on bug 11176= 3 from wychuchol
(In reply to wychuchol from comment #23)
> (In reply to wychuchol from comment #19)
> > After some time in Witcher 3 GOTY run with Lutris PC restarts on =
it's own. I
> > thought something is overheating (I've noticed graphic card memor=
y in
> > PSensor sometimes reaching 90 so I thought maybe that's what's ha=
ppening)
> > but I investigated kern.log and this always happened before that =
autonomous
> > reset:
> >=20
> > Nov  2 22:01:53 pop-os kernel: [  979.244964] pcieport 0000:00:01=
.1: AER:
> > Corrected error received: 0000:01:00.0
> > Nov  2 22:01:53 pop-os kernel: [  979.244967] nvme 0000:01:00.0: =
AER: PCIe
> > Bus Error: severity=3DCorrected, type=3DData Link Layer, (Transmi=
tter ID)
> > Nov  2 22:01:53 pop-os kernel: [  979.244968] nvme 0000:01:00.0: =
AER:=20=20
> > device [1987:5012] error status/mask=3D00001000/00006000
> > Nov  2 22:01:53 pop-os kernel: [  979.244968] nvme 0000:01:00.0: =
AER:=20=20=20
> > [12] Timeout=20=20=20=20=20=20=20=20=20=20=20=20=20=20=20
> > Nov  2 22:01:53 pop-os kernel: [  979.262629] Emergency Sync comp=
lete
>=20
> Thing with those AER errors is that they can go on and on and reset ha=
ppens
> few minutes after the last logged error.=20
> This might be overheating, I managed to find how to output sensors rea=
dings
> into txt log and found that memory went up to 96 C (or rather it stayed
> there for about 1m 10s)
> Last reading before reset:
> amdgpu-pci-2800
> Adapter: PCI adapter
> vddgfx:       +1.16 V=20=20
> fan1:        1551 RPM  (min =3D    0 RPM, max =3D 3200 RPM)
> edge:         +74.0=C2=B0C  (crit =3D +118.0=C2=B0C, hyst =3D -273.1=
=C2=B0C)
>                        (emerg =3D +99.0=C2=B0C)
> junction:     +88.0=C2=B0C  (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=
=B0C)
>                        (emerg =3D +99.0=C2=B0C)
> mem:          +96.0=C2=B0C  (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=
=B0C)
>                        (emerg =3D +99.0=C2=B0C)
> power1:      162.00 W  (cap =3D 195.00 W)
>=20
> k10temp-pci-00c3
> Adapter: PCI adapter
> Tdie:         +70.5=C2=B0C  (high =3D +70.0=C2=B0C)
> Tctl:         +70.5=C2=B0C=20=20
>=20
> Now the weird thing is - if this is in fact overheating why fan didn't=
 go
> beyond 1600 rpm even once.... Highest was like 1581 rpm and I don't ha=
ve
> silent bios switched on (sapphire pulse rx 5700 xt, lever facing away =
from
> video ports).

Okay I don't think it's overheating anymore. I found a moment in Anomaly 1.=
5.0
I can't get past without system resetting, just before a psi storm in Army
Warehouses (I can provide a savefile).

Last sensors reading before crash (5 second increments):
amdgpu-pci-2800
Adapter: PCI adapter
vddgfx:       +1.01 V=20=20
fan1:        1560 RPM  (min =3D    0 RPM, max =3D 3200 RPM)
edge:         +69.0=C2=B0C  (crit =3D +118.0=C2=B0C, hyst =3D -273.1=C2=B0C)
                       (emerg =3D +99.0=C2=B0C)
junction:     +84.0=C2=B0C  (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0C)
                       (emerg =3D +99.0=C2=B0C)
mem:          +80.0=C2=B0C  (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0C)
                       (emerg =3D +99.0=C2=B0C)
power1:      227.00 W  (cap =3D 195.00 W)

k10temp-pci-00c3
Adapter: PCI adapter
Tdie:         +71.8=C2=B0C  (high =3D +70.0=C2=B0C)
Tctl:         +71.8=C2=B0C


You are receiving this mail because:
  • You are the assignee for the bug.
= --15729712838.d0577D.7169-- --===============0249081278== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVs --===============0249081278==--