From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 100666] amdgpu coolers never stoping linux Date: Wed, 28 Feb 2018 03:07:11 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0528416188==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [IPv6:2610:10:20:722:a800:ff:fe98:4b55]) by gabe.freedesktop.org (Postfix) with ESMTP id 7FA4D6E075 for ; Wed, 28 Feb 2018 03:07:11 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0528416188== Content-Type: multipart/alternative; boundary="15197872310.FC3Cf0ed.11345" Content-Transfer-Encoding: 7bit --15197872310.FC3Cf0ed.11345 Date: Wed, 28 Feb 2018 03:07:11 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D100666 --- Comment #13 from Alex Deucher --- (In reply to Luke McKee from comment #11) > In this case it was on topic. The link explains how to use fancontrol scr= ipt > from lm_sensors to work around fan control issues. I saw on another ticket > when I first posted here that dc=3D1 fixed the fancontrol issues. Finally= I > got dc=3D1 working and still it doesn't resolve the dpm fancontrol issues= on > my platform. dc and powerplay are largely independent. It's generally not likely that o= ne will affect the other.=20=20 >=20 > https://github.com/kobalicek/amdtweak > as root > # ./amdtweak --card 0 --verbose --extract-bios /tmp/amdbios.bin > fails. The sysfs shows that the powerplay tables are not proper too. >=20 I'm not familiar with that tool or how it goes about attempting to fetch the vbios. The driver uses several mechanism to fetch it depending on the platform. It's possible that tool does something weird to fetch the vbios = and it's possible that tool incorrectly interprets some of the vbios tables. > [ 4969.713277] resource sanity check: requesting [mem > 0x000c0000-0x000dffff], which spans more than PCI Bus 0000:00 [mem > 0x000c0000-0x000c3fff window] > [ 4969.713283] caller pci_map_rom+0x66/0xf0 mapping multiple BARs > [ 4969.713289] amdgpu 0000:01:00.0: Invalid PCI ROM header signature: > expecting 0xaa55, got 0xffff This last message is from the pci subsystem and is harmless. If the driver were not able to load the vbios, it would fail to load. >=20 > If it can't read it's powerplay table because it can't read the bios maybe > that's why there is all these problems. The driver is able to load the vbios image just fine. If it wasn't able to= , or if there was a major problem with one of the tables, the driver would fail = to load. >=20 >=20 > (In reply to Alex Deucher from comment #9) > >=20 > > Please stop posting this on every bug report. >=20 > https://bugs.freedesktop.org/show_bug.cgi?id=3D100666#c0 > Also the users above on this ticket above here when they grepped their dm= esg > wouldn't have output any powerplay mes.sages because they grepped radeon > instead of amdgpu >=20 > [ 10.124232] amdgpu: [powerplay]=20 > failed to send message 309 ret is 254=20 > [ 10.124248] amdgpu: [powerplay]=20 > failed to send pre message 14e ret is 254=20 >=20 There are lots of reasons an smu message might fail. Just because you see = an smu message failure does not mean you are seeing the same issue as someone else. It's like a GPU hang. There are lots of potential root causes. --=20 You are receiving this mail because: You are the assignee for the bug.= --15197872310.FC3Cf0ed.11345 Date: Wed, 28 Feb 2018 03:07:11 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Comme= nt # 13 on bug 10066= 6 from Alex Deucher
(In reply to Luke McKee from comment #11)
> In this case it was on topic. The link explains =
how to use fancontrol script
> from lm_sensors to work around fan control issues. I saw on another ti=
cket
> when I first posted here that dc=3D1 fixed the fancontrol issues. Fina=
lly I
> got dc=3D1 working and still it doesn't resolve the dpm fancontrol iss=
ues on
> my platform.

dc and powerplay are largely independent.  It's generally not likely that o=
ne
will affect the other.=20=20

>=20
> https://github.com/k=
obalicek/amdtweak
> as root
> # ./amdtweak  --card 0 --verbose --extract-bios /tmp/amdbios.bin
> fails. The sysfs shows that the powerplay tables are not proper too.
> 

I'm not familiar with that tool or how it goes about attempting to fetch the
vbios.  The driver uses several mechanism to fetch it depending on the
platform.  It's possible that tool does something weird to fetch the vbios =
and
it's possible that tool incorrectly interprets some of the vbios tables.

> [ 4969.713277] resource sanity check: requesting=
 [mem
> 0x000c0000-0x000dffff], which spans more than PCI Bus 0000:00 [mem
> 0x000c0000-0x000c3fff window]
> [ 4969.713283] caller pci_map_rom+0x66/0xf0 mapping multiple BARs
> [ 4969.713289] amdgpu 0000:01:00.0: Invalid PCI ROM header signature:
> expecting 0xaa55, got 0xffff

This last message is from the pci subsystem and is harmless.  If the driver
were not able to load the vbios, it would fail to load.

>=20
> If it can't read it's powerplay table because it can't read the bios m=
aybe
> that's why there is all these problems.

The driver is able to load the vbios image just fine.  If it wasn't able to=
, or
if there was a major problem with one of the tables, the driver would fail =
to
load.

>=20
>=20
>  (In reply to Alex Deucher from comment #9)
> >=20
> > Please stop posting this on every bug report.
>=20
> https://bugs.freedesktop.org/show_b=
ug.cgi?id=3D100666#c0
> Also the users above on this ticket above here when they grepped their=
 dmesg
> wouldn't have output any powerplay mes.sages because they grepped rade=
on
> instead of amdgpu
>=20
> [   10.124232] amdgpu: [powerplay]=20
>                 failed to send message 309 ret is 254=20
> [   10.124248] amdgpu: [powerplay]=20
>                 failed to send pre message 14e ret is 254=20
> 

There are lots of reasons an smu message might fail.  Just because you see =
an
smu message failure does not mean you are seeing the same issue as someone
else.  It's like a GPU hang.  There are lots of potential root causes.


You are receiving this mail because:
  • You are the assignee for the bug.
= --15197872310.FC3Cf0ed.11345-- --===============0528416188== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============0528416188==--