From mboxrd@z Thu Jan 1 00:00:00 1970
From: bugzilla-daemon@freedesktop.org
Subject: [Bug 93341] Semi-random GPU lockups on radeonsi with a RadeonHD 7770
(when playing videos, running OpenGL games, WebGL apps,
or after extended periods of time)
Date: Tue, 11 Apr 2017 22:20:05 +0000
Message-ID:
References:
Mime-Version: 1.0
Content-Type: multipart/mixed; boundary="===============0021322733=="
Return-path:
Received: from culpepper.freedesktop.org (culpepper.freedesktop.org
[IPv6:2610:10:20:722:a800:ff:fe98:4b55])
by gabe.freedesktop.org (Postfix) with ESMTP id A5DC06E158
for ; Tue, 11 Apr 2017 22:20:05 +0000 (UTC)
In-Reply-To:
List-Unsubscribe: ,
List-Archive:
List-Post:
List-Help:
List-Subscribe: ,
Errors-To: dri-devel-bounces@lists.freedesktop.org
Sender: "dri-devel"
To: dri-devel@lists.freedesktop.org
List-Id: dri-devel@lists.freedesktop.org
--===============0021322733==
Content-Type: multipart/alternative; boundary="14919492050.DeD33Aa8.854";
charset="UTF-8"
--14919492050.DeD33Aa8.854
Date: Tue, 11 Apr 2017 22:20:05 +0000
MIME-Version: 1.0
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
X-Bugzilla-URL: http://bugs.freedesktop.org/
Auto-Submitted: auto-generated
https://bugs.freedesktop.org/show_bug.cgi?id=3D93341
--- Comment #26 from Jean-Fran=C3=A7ois Fortin Tam ---
OK, I've got good news... Julien, thanks to the crazy furry donut "torture
test" you suggested, I was able to finally pinpoint the real trigger for th=
is
bug.
My understanding is that on Radeons (well, at least the Radeon HD 7770), th=
ere
is an emergency mechanism in the hardware (or firmware/microcode maybe) that
activates self-throttling of performances when the GPU reaches a critical
temperature. Normally, the video driver is supposed to handle this state ch=
ange
gracefully, however the radeonsi/radeon/amdgpu driver on Linux does not, so=
the
kernel panics because the driver went belly up.
During additional testing today, where I forced my GPU to overheat, I was a=
ble
to determine that the critical point is the same as on Windows: 113 degrees
Celsius. As soon as you go over 112... boom, dead radeonsi driver + kernel =
oops
(with the same error messages as my previous logs above). Additionally,
lm_sensors thinks the temperature has instantly jumped to 511 degrees Celsi=
us
(!), and the readings stay stuck at 511 Celsius.
"Duh! Just get better cooling!" might sound like a workaround (just like
keeping the case open), but nope, technically, it's still a software/driver
issue: the Linux driver should handle such scenarios gracefully just as wel=
l as
the Windows driver. In Windows, breaching the 110-113 degrees Celsius limit
results in the video driver simply dropping frames massively, continuing to
function at reduced performance (ie: going from 40-60 fps to 10-15 fps on o=
ne
of my benchmarks). The system never crashes.
So the bug here, as I understand it, is that the radeonsi driver on Linux d=
oes
not handle the event where the hardware force-throttles itself.
---------
Contextual notes:
The reason why I only started experiencing this issue in December 2015 (as =
I've
had the GPU since 2012) was that I changed my PC case then, which means a
different airflow and cooling behavior... And the reason why it was so hard=
to
get consistent crashes here was that when I was trying to troubleshoot it, I
was sometimes doing it with the case closed, sometimes with the case open (=
when
trying with a different power supply unit using a "siamese transplant" acro=
ss
another computer, for example). If I keep my case open, the card will never
reach the critical temperature and so the issue will not happen. I might ge=
t a
system "freeze" (possibly saying "*ERROR*
si_restrict_performance_levels_before_switch failed") after many hours of
torture testing, but the symptoms are different (the screen does not turn o=
ff,
image stays on with everything frozen, and nothing else in the logs) and so=
I
presume that to be a different issue.
--=20
You are receiving this mail because:
You are the assignee for the bug.=
--14919492050.DeD33Aa8.854
Date: Tue, 11 Apr 2017 22:20:05 +0000
MIME-Version: 1.0
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
X-Bugzilla-URL: http://bugs.freedesktop.org/
Auto-Submitted: auto-generated
Commen=
t # 26
on bug 93341<=
/a>
from Jean-Fran=C3=A7ois Fortin Tam
OK, I've got good news... Julien, thanks to the crazy furry do=
nut "torture
test" you suggested, I was able to finally pinpoint the real trigger f=
or this
bug.
My understanding is that on Radeons (well, at least the Radeon HD 7770), th=
ere
is an emergency mechanism in the hardware (or firmware/microcode maybe) that
activates self-throttling of performances when the GPU reaches a critical
temperature. Normally, the video driver is supposed to handle this state ch=
ange
gracefully, however the radeonsi/radeon/amdgpu driver on Linux does not, so=
the
kernel panics because the driver went belly up.
During additional testing today, where I forced my GPU to overheat, I was a=
ble
to determine that the critical point is the same as on Windows: 113 degrees
Celsius. As soon as you go over 112... boom, dead radeonsi driver + kernel =
oops
(with the same error messages as my previous logs above). Additionally,
lm_sensors thinks the temperature has instantly jumped to 511 degrees Celsi=
us
(!), and the readings stay stuck at 511 Celsius.
"Duh! Just get better cooling!" might sound like a workaround (ju=
st like
keeping the case open), but nope, technically, it's still a software/driver
issue: the Linux driver should handle such scenarios gracefully just as wel=
l as
the Windows driver. In Windows, breaching the 110-113 degrees Celsius limit
results in the video driver simply dropping frames massively, continuing to
function at reduced performance (ie: going from 40-60 fps to 10-15 fps on o=
ne
of my benchmarks). The system never crashes.
So the bug here, as I understand it, is that the radeonsi driver on Linux d=
oes
not handle the event where the hardware force-throttles itself.
---------
Contextual notes:
The reason why I only started experiencing this issue in December 2015 (as =
I've
had the GPU since 2012) was that I changed my PC case then, which means a
different airflow and cooling behavior... And the reason why it was so hard=
to
get consistent crashes here was that when I was trying to troubleshoot it, I
was sometimes doing it with the case closed, sometimes with the case open (=
when
trying with a different power supply unit using a "siamese transplant&=
quot; across
another computer, for example). If I keep my case open, the card will never
reach the critical temperature and so the issue will not happen. I might ge=
t a
system "freeze" (possibly saying "*ERROR*
si_restrict_performance_levels_before_switch failed") after many hours=
of
torture testing, but the symptoms are different (the screen does not turn o=
ff,
image stays on with everything frozen, and nothing else in the logs) and so=
I
presume that to be a different issue.
You are receiving this mail because:
- You are the assignee for the bug.
=
--14919492050.DeD33Aa8.854--
--===============0021322733==
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
Content-Disposition: inline
X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs
IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz
dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg==
--===============0021322733==--