From mboxrd@z Thu Jan 1 00:00:00 1970
From: bugzilla-daemon@freedesktop.org
Subject: [Bug 111763] ring_gfx hangs/freezes on Navi gpus
Date: Tue, 05 Nov 2019 16:28:03 +0000
Message-ID:
References:
Mime-Version: 1.0
Content-Type: multipart/mixed; boundary="===============0249081278=="
Return-path:
Received: from culpepper.freedesktop.org (culpepper.freedesktop.org
[IPv6:2610:10:20:722:a800:ff:fe98:4b55])
by gabe.freedesktop.org (Postfix) with ESMTP id E486A6EAF3
for ; Tue, 5 Nov 2019 16:28:03 +0000 (UTC)
In-Reply-To:
List-Unsubscribe: ,
List-Archive:
List-Post:
List-Help:
List-Subscribe: ,
Errors-To: dri-devel-bounces@lists.freedesktop.org
Sender: "dri-devel"
To: dri-devel@lists.freedesktop.org
List-Id: dri-devel@lists.freedesktop.org
--===============0249081278==
Content-Type: multipart/alternative; boundary="15729712838.d0577D.7169"
Content-Transfer-Encoding: 7bit
--15729712838.d0577D.7169
Date: Tue, 5 Nov 2019 16:28:03 +0000
MIME-Version: 1.0
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
X-Bugzilla-URL: http://bugs.freedesktop.org/
Auto-Submitted: auto-generated
https://bugs.freedesktop.org/show_bug.cgi?id=3D111763
--- Comment #24 from wychuchol ---
(In reply to wychuchol from comment #23)
> (In reply to wychuchol from comment #19)
> > After some time in Witcher 3 GOTY run with Lutris PC restarts on it's o=
wn. I
> > thought something is overheating (I've noticed graphic card memory in
> > PSensor sometimes reaching 90 so I thought maybe that's what's happenin=
g)
> > but I investigated kern.log and this always happened before that autono=
mous
> > reset:
> >=20
> > Nov 2 22:01:53 pop-os kernel: [ 979.244964] pcieport 0000:00:01.1: AE=
R:
> > Corrected error received: 0000:01:00.0
> > Nov 2 22:01:53 pop-os kernel: [ 979.244967] nvme 0000:01:00.0: AER: P=
CIe
> > Bus Error: severity=3DCorrected, type=3DData Link Layer, (Transmitter I=
D)
> > Nov 2 22:01:53 pop-os kernel: [ 979.244968] nvme 0000:01:00.0: AER:=
=20=20
> > device [1987:5012] error status/mask=3D00001000/00006000
> > Nov 2 22:01:53 pop-os kernel: [ 979.244968] nvme 0000:01:00.0: AER:=
=20=20=20
> > [12] Timeout=20=20=20=20=20=20=20=20=20=20=20=20=20=20=20
> > Nov 2 22:01:53 pop-os kernel: [ 979.262629] Emergency Sync complete
>=20
> Thing with those AER errors is that they can go on and on and reset happe=
ns
> few minutes after the last logged error.=20
> This might be overheating, I managed to find how to output sensors readin=
gs
> into txt log and found that memory went up to 96 C (or rather it stayed
> there for about 1m 10s)
> Last reading before reset:
> amdgpu-pci-2800
> Adapter: PCI adapter
> vddgfx: +1.16 V=20=20
> fan1: 1551 RPM (min =3D 0 RPM, max =3D 3200 RPM)
> edge: +74.0=C2=B0C (crit =3D +118.0=C2=B0C, hyst =3D -273.1=C2=
=B0C)
> (emerg =3D +99.0=C2=B0C)
> junction: +88.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0=
C)
> (emerg =3D +99.0=C2=B0C)
> mem: +96.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0=
C)
> (emerg =3D +99.0=C2=B0C)
> power1: 162.00 W (cap =3D 195.00 W)
>=20
> k10temp-pci-00c3
> Adapter: PCI adapter
> Tdie: +70.5=C2=B0C (high =3D +70.0=C2=B0C)
> Tctl: +70.5=C2=B0C=20=20
>=20
> Now the weird thing is - if this is in fact overheating why fan didn't go
> beyond 1600 rpm even once.... Highest was like 1581 rpm and I don't have
> silent bios switched on (sapphire pulse rx 5700 xt, lever facing away from
> video ports).
Okay I don't think it's overheating anymore. I found a moment in Anomaly 1.=
5.0
I can't get past without system resetting, just before a psi storm in Army
Warehouses (I can provide a savefile).
Last sensors reading before crash (5 second increments):
amdgpu-pci-2800
Adapter: PCI adapter
vddgfx: +1.01 V=20=20
fan1: 1560 RPM (min =3D 0 RPM, max =3D 3200 RPM)
edge: +69.0=C2=B0C (crit =3D +118.0=C2=B0C, hyst =3D -273.1=C2=B0C)
(emerg =3D +99.0=C2=B0C)
junction: +84.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0C)
(emerg =3D +99.0=C2=B0C)
mem: +80.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0C)
(emerg =3D +99.0=C2=B0C)
power1: 227.00 W (cap =3D 195.00 W)
k10temp-pci-00c3
Adapter: PCI adapter
Tdie: +71.8=C2=B0C (high =3D +70.0=C2=B0C)
Tctl: +71.8=C2=B0C
--=20
You are receiving this mail because:
You are the assignee for the bug.=
--15729712838.d0577D.7169
Date: Tue, 5 Nov 2019 16:28:03 +0000
MIME-Version: 1.0
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
X-Bugzilla-URL: http://bugs.freedesktop.org/
Auto-Submitted: auto-generated
Comme=
nt # 24
on bug 11176=
3
from wychuchol
(In reply to wychuchol from comment #23)
> (In reply to wychuchol from comment #19)
> > After some time in Witcher 3 GOTY run with Lutris PC restarts on =
it's own. I
> > thought something is overheating (I've noticed graphic card memor=
y in
> > PSensor sometimes reaching 90 so I thought maybe that's what's ha=
ppening)
> > but I investigated kern.log and this always happened before that =
autonomous
> > reset:
> >=20
> > Nov 2 22:01:53 pop-os kernel: [ 979.244964] pcieport 0000:00:01=
.1: AER:
> > Corrected error received: 0000:01:00.0
> > Nov 2 22:01:53 pop-os kernel: [ 979.244967] nvme 0000:01:00.0: =
AER: PCIe
> > Bus Error: severity=3DCorrected, type=3DData Link Layer, (Transmi=
tter ID)
> > Nov 2 22:01:53 pop-os kernel: [ 979.244968] nvme 0000:01:00.0: =
AER:=20=20
> > device [1987:5012] error status/mask=3D00001000/00006000
> > Nov 2 22:01:53 pop-os kernel: [ 979.244968] nvme 0000:01:00.0: =
AER:=20=20=20
> > [12] Timeout=20=20=20=20=20=20=20=20=20=20=20=20=20=20=20
> > Nov 2 22:01:53 pop-os kernel: [ 979.262629] Emergency Sync comp=
lete
>=20
> Thing with those AER errors is that they can go on and on and reset ha=
ppens
> few minutes after the last logged error.=20
> This might be overheating, I managed to find how to output sensors rea=
dings
> into txt log and found that memory went up to 96 C (or rather it stayed
> there for about 1m 10s)
> Last reading before reset:
> amdgpu-pci-2800
> Adapter: PCI adapter
> vddgfx: +1.16 V=20=20
> fan1: 1551 RPM (min =3D 0 RPM, max =3D 3200 RPM)
> edge: +74.0=C2=B0C (crit =3D +118.0=C2=B0C, hyst =3D -273.1=
=C2=B0C)
> (emerg =3D +99.0=C2=B0C)
> junction: +88.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=
=B0C)
> (emerg =3D +99.0=C2=B0C)
> mem: +96.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=
=B0C)
> (emerg =3D +99.0=C2=B0C)
> power1: 162.00 W (cap =3D 195.00 W)
>=20
> k10temp-pci-00c3
> Adapter: PCI adapter
> Tdie: +70.5=C2=B0C (high =3D +70.0=C2=B0C)
> Tctl: +70.5=C2=B0C=20=20
>=20
> Now the weird thing is - if this is in fact overheating why fan didn't=
go
> beyond 1600 rpm even once.... Highest was like 1581 rpm and I don't ha=
ve
> silent bios switched on (sapphire pulse rx 5700 xt, lever facing away =
from
> video ports).
Okay I don't think it's overheating anymore. I found a moment in Anomaly 1.=
5.0
I can't get past without system resetting, just before a psi storm in Army
Warehouses (I can provide a savefile).
Last sensors reading before crash (5 second increments):
amdgpu-pci-2800
Adapter: PCI adapter
vddgfx: +1.01 V=20=20
fan1: 1560 RPM (min =3D 0 RPM, max =3D 3200 RPM)
edge: +69.0=C2=B0C (crit =3D +118.0=C2=B0C, hyst =3D -273.1=C2=B0C)
(emerg =3D +99.0=C2=B0C)
junction: +84.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0C)
(emerg =3D +99.0=C2=B0C)
mem: +80.0=C2=B0C (crit =3D +99.0=C2=B0C, hyst =3D -273.1=C2=B0C)
(emerg =3D +99.0=C2=B0C)
power1: 227.00 W (cap =3D 195.00 W)
k10temp-pci-00c3
Adapter: PCI adapter
Tdie: +71.8=C2=B0C (high =3D +70.0=C2=B0C)
Tctl: +71.8=C2=B0C
You are receiving this mail because:
- You are the assignee for the bug.
=
--15729712838.d0577D.7169--
--===============0249081278==
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
Content-Disposition: inline
X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs
IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz
dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVs
--===============0249081278==--