From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 93341] Semi-random GPU lockups on radeonsi with a RadeonHD 7770 (when playing videos, running OpenGL games, WebGL apps, or after extended periods of time) Date: Tue, 11 Apr 2017 22:20:05 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0021322733==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [IPv6:2610:10:20:722:a800:ff:fe98:4b55]) by gabe.freedesktop.org (Postfix) with ESMTP id A5DC06E158 for ; Tue, 11 Apr 2017 22:20:05 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0021322733== Content-Type: multipart/alternative; boundary="14919492050.DeD33Aa8.854"; charset="UTF-8" --14919492050.DeD33Aa8.854 Date: Tue, 11 Apr 2017 22:20:05 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D93341 --- Comment #26 from Jean-Fran=C3=A7ois Fortin Tam --- OK, I've got good news... Julien, thanks to the crazy furry donut "torture test" you suggested, I was able to finally pinpoint the real trigger for th= is bug. My understanding is that on Radeons (well, at least the Radeon HD 7770), th= ere is an emergency mechanism in the hardware (or firmware/microcode maybe) that activates self-throttling of performances when the GPU reaches a critical temperature. Normally, the video driver is supposed to handle this state ch= ange gracefully, however the radeonsi/radeon/amdgpu driver on Linux does not, so= the kernel panics because the driver went belly up. During additional testing today, where I forced my GPU to overheat, I was a= ble to determine that the critical point is the same as on Windows: 113 degrees Celsius. As soon as you go over 112... boom, dead radeonsi driver + kernel = oops (with the same error messages as my previous logs above). Additionally, lm_sensors thinks the temperature has instantly jumped to 511 degrees Celsi= us (!), and the readings stay stuck at 511 Celsius. "Duh! Just get better cooling!" might sound like a workaround (just like keeping the case open), but nope, technically, it's still a software/driver issue: the Linux driver should handle such scenarios gracefully just as wel= l as the Windows driver. In Windows, breaching the 110-113 degrees Celsius limit results in the video driver simply dropping frames massively, continuing to function at reduced performance (ie: going from 40-60 fps to 10-15 fps on o= ne of my benchmarks). The system never crashes. So the bug here, as I understand it, is that the radeonsi driver on Linux d= oes not handle the event where the hardware force-throttles itself. --------- Contextual notes: The reason why I only started experiencing this issue in December 2015 (as = I've had the GPU since 2012) was that I changed my PC case then, which means a different airflow and cooling behavior... And the reason why it was so hard= to get consistent crashes here was that when I was trying to troubleshoot it, I was sometimes doing it with the case closed, sometimes with the case open (= when trying with a different power supply unit using a "siamese transplant" acro= ss another computer, for example). If I keep my case open, the card will never reach the critical temperature and so the issue will not happen. I might ge= t a system "freeze" (possibly saying "*ERROR* si_restrict_performance_levels_before_switch failed") after many hours of torture testing, but the symptoms are different (the screen does not turn o= ff, image stays on with everything frozen, and nothing else in the logs) and so= I presume that to be a different issue. --=20 You are receiving this mail because: You are the assignee for the bug.= --14919492050.DeD33Aa8.854 Date: Tue, 11 Apr 2017 22:20:05 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 26 on bug 93341<= /a> from Jean-Fran=C3=A7ois Fortin Tam
OK, I've got good news... Julien, thanks to the crazy furry do=
nut "torture
test" you suggested, I was able to finally pinpoint the real trigger f=
or this
bug.

My understanding is that on Radeons (well, at least the Radeon HD 7770), th=
ere
is an emergency mechanism in the hardware (or firmware/microcode maybe) that
activates self-throttling of performances when the GPU reaches a critical
temperature. Normally, the video driver is supposed to handle this state ch=
ange
gracefully, however the radeonsi/radeon/amdgpu driver on Linux does not, so=
 the
kernel panics because the driver went belly up.

During additional testing today, where I forced my GPU to overheat, I was a=
ble
to determine that the critical point is the same as on Windows: 113 degrees
Celsius. As soon as you go over 112... boom, dead radeonsi driver + kernel =
oops
(with the same error messages as my previous logs above). Additionally,
lm_sensors thinks the temperature has instantly jumped to 511 degrees Celsi=
us
(!), and the readings stay stuck at 511 Celsius.

"Duh! Just get better cooling!" might sound like a workaround (ju=
st like
keeping the case open), but nope, technically, it's still a software/driver
issue: the Linux driver should handle such scenarios gracefully just as wel=
l as
the Windows driver. In Windows, breaching the 110-113 degrees Celsius limit
results in the video driver simply dropping frames massively, continuing to
function at reduced performance (ie: going from 40-60 fps to 10-15 fps on o=
ne
of my benchmarks). The system never crashes.

So the bug here, as I understand it, is that the radeonsi driver on Linux d=
oes
not handle the event where the hardware force-throttles itself.

---------
Contextual notes:
The reason why I only started experiencing this issue in December 2015 (as =
I've
had the GPU since 2012) was that I changed my PC case then, which means a
different airflow and cooling behavior... And the reason why it was so hard=
 to
get consistent crashes here was that when I was trying to troubleshoot it, I
was sometimes doing it with the case closed, sometimes with the case open (=
when
trying with a different power supply unit using a "siamese transplant&=
quot; across
another computer, for example). If I keep my case open, the card will never
reach the critical temperature and so the issue will not happen. I might ge=
t a
system "freeze" (possibly saying "*ERROR*
si_restrict_performance_levels_before_switch failed") after many hours=
 of
torture testing, but the symptoms are different (the screen does not turn o=
ff,
image stays on with everything frozen, and nothing else in the logs) and so=
 I
presume that to be a different issue.


You are receiving this mail because:
  • You are the assignee for the bug.
= --14919492050.DeD33Aa8.854-- --===============0021322733== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============0021322733==--