From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 88301] Dota causes GPU fault and kernel hang Date: Thu, 15 Jan 2015 20:06:38 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============1114542409==" Return-path: Received: from culpepper.freedesktop.org (unknown [131.252.210.165]) by gabe.freedesktop.org (Postfix) with ESMTP id 67CCD6E021 for ; Thu, 15 Jan 2015 12:06:38 -0800 (PST) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============1114542409== Content-Type: multipart/alternative; boundary="1421352398.aaf20.804"; charset="UTF-8" --1421352398.aaf20.804 Date: Thu, 15 Jan 2015 20:06:38 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" https://bugs.freedesktop.org/show_bug.cgi?id=88301 --- Comment #2 from Tilman Sauerbeck --- Sorry about that misinformation. It's not a hang at all since I'm still able to use sysrq to reboot. What's happening after the GPU faults is that apparently the driver attempts to get the card back into working shape, but fails to do so. X doesn't become usable again after the GPU faults anyway. Here's the kernel log following the GPU faults: radeon 0000:01:00.0: ring 0 stalled for more than 10428msec radeon 0000:01:00.0: GPU lockup (current fence id 0x0000000000107f5c last fence id 0x00000000001080f7 on ring 0) radeon 0000:01:00.0: failed to get a new IB (-35) [drm:radeon_cs_ib_fill] *ERROR* Failed to get ib ! radeon 0000:01:00.0: Saved 7977 dwords of commands on ring 0. radeon 0000:01:00.0: GPU softreset: 0x00000009 [snipped list of registers that were reset (I think)] [drm] probing gen 2 caps for device 1002:5a16 = 31cd02/0 [drm] PCIE gen 2 link speeds already enabled [drm] PCIE GART of 1024M enabled (table at 0x000000000078C000). radeon 0000:01:00.0: WB enabled radeon 0000:01:00.0: fence driver on ring 0 use gpu addr 0x0000000080000c00 and cpu addr 0xffff8800bac4cc00 radeon 0000:01:00.0: fence driver on ring 1 use gpu addr 0x0000000080000c04 and cpu addr 0xffff8800bac4cc04 radeon 0000:01:00.0: fence driver on ring 2 use gpu addr 0x0000000080000c08 and cpu addr 0xffff8800bac4cc08 radeon 0000:01:00.0: fence driver on ring 3 use gpu addr 0x0000000080000c0c and cpu addr 0xffff8800bac4cc0c radeon 0000:01:00.0: fence driver on ring 4 use gpu addr 0x0000000080000c10 and cpu addr 0xffff8800bac4cc10 radeon 0000:01:00.0: fence driver on ring 5 use gpu addr 0x0000000000076c98 and cpu addr 0xffffc90010c36c98 radeon 0000:01:00.0: fence driver on ring 6 use gpu addr 0x0000000080000c18 and cpu addr 0xffff8800bac4cc18 radeon 0000:01:00.0: fence driver on ring 7 use gpu addr 0x0000000080000c1c and cpu addr 0xffff8800bac4cc1c [drm] ring test on 0 succeeded in 3 usecs [drm:cik_ring_test] *ERROR* radeon: ring 1 test failed (scratch(0x3010C)=0xCAFEDEAD) [drm:cik_ring_test] *ERROR* radeon: ring 2 test failed (scratch(0x3010C)=0xCAFEDEAD) [drm:cik_sdma_ring_test] *ERROR* radeon: ring 3 test failed (0xCAFEDEAD) [drm:cik_resume] *ERROR* cik startup failed on resume [drm:radeon_pm_resume_dpm] *ERROR* radeon: dpm resume failed -- You are receiving this mail because: You are the assignee for the bug. --1421352398.aaf20.804 Date: Thu, 15 Jan 2015 20:06:38 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8"

Comment # 2 on bug 88301 from
Sorry about that misinformation.
It's not a hang at all since I'm still able to use sysrq to reboot.

What's happening after the GPU faults is that apparently the driver attempts to
get the card back into working shape, but fails to do so. X doesn't become
usable again after the GPU faults anyway.

Here's the kernel log following the GPU faults:

radeon 0000:01:00.0: ring 0 stalled for more than 10428msec
radeon 0000:01:00.0: GPU lockup (current fence id 0x0000000000107f5c last fence
id 0x00000000001080f7 on ring 0)
radeon 0000:01:00.0: failed to get a new IB (-35)
[drm:radeon_cs_ib_fill] *ERROR* Failed to get ib !
radeon 0000:01:00.0: Saved 7977 dwords of commands on ring 0.
radeon 0000:01:00.0: GPU softreset: 0x00000009
[snipped list of registers that were reset (I think)]

[drm] probing gen 2 caps for device 1002:5a16 = 31cd02/0
[drm] PCIE gen 2 link speeds already enabled
[drm] PCIE GART of 1024M enabled (table at 0x000000000078C000).
radeon 0000:01:00.0: WB enabled
radeon 0000:01:00.0: fence driver on ring 0 use gpu addr 0x0000000080000c00 and
cpu addr 0xffff8800bac4cc00
radeon 0000:01:00.0: fence driver on ring 1 use gpu addr 0x0000000080000c04 and
cpu addr 0xffff8800bac4cc04
radeon 0000:01:00.0: fence driver on ring 2 use gpu addr 0x0000000080000c08 and
cpu addr 0xffff8800bac4cc08
radeon 0000:01:00.0: fence driver on ring 3 use gpu addr 0x0000000080000c0c and
cpu addr 0xffff8800bac4cc0c
radeon 0000:01:00.0: fence driver on ring 4 use gpu addr 0x0000000080000c10 and
cpu addr 0xffff8800bac4cc10
radeon 0000:01:00.0: fence driver on ring 5 use gpu addr 0x0000000000076c98 and
cpu addr 0xffffc90010c36c98
radeon 0000:01:00.0: fence driver on ring 6 use gpu addr 0x0000000080000c18 and
cpu addr 0xffff8800bac4cc18
radeon 0000:01:00.0: fence driver on ring 7 use gpu addr 0x0000000080000c1c and
cpu addr 0xffff8800bac4cc1c
[drm] ring test on 0 succeeded in 3 usecs
[drm:cik_ring_test] *ERROR* radeon: ring 1 test failed
(scratch(0x3010C)=0xCAFEDEAD)
[drm:cik_ring_test] *ERROR* radeon: ring 2 test failed
(scratch(0x3010C)=0xCAFEDEAD)
[drm:cik_sdma_ring_test] *ERROR* radeon: ring 3 test failed (0xCAFEDEAD)
[drm:cik_resume] *ERROR* cik startup failed on resume
[drm:radeon_pm_resume_dpm] *ERROR* radeon: dpm resume failed


You are receiving this mail because:
  • You are the assignee for the bug.
--1421352398.aaf20.804-- --===============1114542409== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHA6Ly9saXN0 cy5mcmVlZGVza3RvcC5vcmcvbWFpbG1hbi9saXN0aW5mby9kcmktZGV2ZWwK --===============1114542409==--