From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 109692] deadlock occurs during GPU reset Date: Tue, 26 Feb 2019 10:42:22 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============1772859566==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [131.252.210.165]) by gabe.freedesktop.org (Postfix) with ESMTP id 238BB89CF5 for ; Tue, 26 Feb 2019 10:42:22 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============1772859566== Content-Type: multipart/alternative; boundary="15511777420.C59e6E.25471" Content-Transfer-Encoding: 7bit --15511777420.C59e6E.25471 Date: Tue, 26 Feb 2019 10:42:22 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D109692 --- Comment #10 from mikhail.v.gavrilov@gmail.com --- Even without reproducing GPU hang in kernel log I found "suspicious RCU usa= ge" and some errors. [drm:amdgpu_ctx_mgr_entity_fini [amdgpu]] *ERROR* ctx 000000002caf7aed is s= till alive [drm:amdgpu_ctx_mgr_fini [amdgpu]] *ERROR* ctx 000000002caf7aed is still al= ive =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D WARNING: suspicious RCU usage 5.0.0-rc1-drm-next-kernel+ #1 Tainted: G C=20=20=20=20=20=20=20 ----------------------------- include/linux/rcupdate.h:280 Illegal context switch in RCU read-side critic= al section! other info that might help us debug this: rcu_scheduler_active =3D 2, debug_locks =3D 1 3 locks held by CrashBandicootN/26312: #0: 00000000eb680bad (&f->f_pos_lock){+.+.}, at: __fdget_pos+0x4d/0x60 #1: 00000000b3a3c406 (&p->lock){+.+.}, at: seq_read+0x38/0x410 #2: 000000007c893f05 (rcu_read_lock){....}, at: dev_seq_start+0x5/0x100 stack backtrace: CPU: 8 PID: 26312 Comm: CrashBandicootN Tainted: G C=20=20=20=20=20= =20=20 5.0.0-rc1-drm-next-kernel+ #1 Hardware name: System manufacturer System Product Name/ROG STRIX X470-I GAM= ING, BIOS 1103 11/16/2018 Call Trace: dump_stack+0x85/0xc0 ___might_sleep+0x100/0x180 __mutex_lock+0x61/0x930 ? igb_get_stats64+0x29/0x80 [igb] ? seq_vprintf+0x33/0x50 ? igb_get_stats64+0x29/0x80 [igb] igb_get_stats64+0x29/0x80 [igb] dev_get_stats+0x5c/0xc0 dev_seq_printf_stats+0x33/0xe0 dev_seq_show+0x10/0x30 seq_read+0x2fa/0x410 proc_reg_read+0x3c/0x60 __vfs_read+0x37/0x1b0 vfs_read+0xb2/0x170 ksys_read+0x52/0xc0 do_syscall_64+0x5c/0xa0 entry_SYSCALL_64_after_hwframe+0x49/0xbe RIP: 0033:0x7f2188d8934c Code: ec 28 48 89 54 24 18 48 89 74 24 10 89 7c 24 08 e8 79 c9 01 00 48 8b = 54 24 18 48 8b 74 24 10 41 89 c0 8b 7c 24 08 31 c0 0f 05 <48> 3d 00 f0 ff ff 7= 7 30 44 89 c7 48 89 44 24 08 e8 af c9 01 00 48 RSP: 002b:000000000023f010 EFLAGS: 00000246 ORIG_RAX: 0000000000000000 RAX: ffffffffffffffda RBX: 000000007d11b6d0 RCX: 00007f2188d8934c RDX: 0000000000000400 RSI: 000000007d0dd4f0 RDI: 000000000000007b RBP: 0000000000000d68 R08: 0000000000000000 R09: 0000000000000000 R10: 00007f2188621c40 R11: 0000000000000246 R12: 00007f2188e59740 R13: 00007f2188e5a340 R14: 00000000000001ff R15: 000000007d11b6d0 This only occures when I use "amd-staging-drm-next". --=20 You are receiving this mail because: You are the assignee for the bug.= --15511777420.C59e6E.25471 Date: Tue, 26 Feb 2019 10:42:22 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Comme= nt # 10 on bug 10969= 2 from mikhail.v.gavrilov@gmail.com
Even without reproducing GPU hang in kernel log I found "=
suspicious RCU usage"
and some errors.

[drm:amdgpu_ctx_mgr_entity_fini [amdgpu]] *ERROR* ctx 000000002caf7aed is s=
till
alive
[drm:amdgpu_ctx_mgr_fini [amdgpu]] *ERROR* ctx 000000002caf7aed is still al=
ive

=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D
WARNING: suspicious RCU usage
5.0.0-rc1-drm-next-kernel+ #1 Tainted: G         C=20=20=20=20=20=20=20
-----------------------------
include/linux/rcupdate.h:280 Illegal context switch in RCU read-side critic=
al
section!

other info that might help us debug this:

rcu_scheduler_active =3D 2, debug_locks =3D 1
3 locks held by CrashBandicootN/26312:
 #0: 00000000eb680bad (&f->f_pos_lock){+.+.}, at: __fdget_pos+0x4d/0=
x60
 #1: 00000000b3a3c406 (&p->lock){+.+.}, at: seq_read+0x38/0x410
 #2: 000000007c893f05 (rcu_read_lock){....}, at: dev_seq_start+0x5/0x100

stack backtrace:
CPU: 8 PID: 26312 Comm: CrashBandicootN Tainted: G         C=20=20=20=20=20=
=20=20
5.0.0-rc1-drm-next-kernel+ #1
Hardware name: System manufacturer System Product Name/ROG STRIX X470-I GAM=
ING,
BIOS 1103 11/16/2018
Call Trace:
 dump_stack+0x85/0xc0
 ___might_sleep+0x100/0x180
 __mutex_lock+0x61/0x930
 ? igb_get_stats64+0x29/0x80 [igb]
 ? seq_vprintf+0x33/0x50
 ? igb_get_stats64+0x29/0x80 [igb]
 igb_get_stats64+0x29/0x80 [igb]
 dev_get_stats+0x5c/0xc0
 dev_seq_printf_stats+0x33/0xe0
 dev_seq_show+0x10/0x30
 seq_read+0x2fa/0x410
 proc_reg_read+0x3c/0x60
 __vfs_read+0x37/0x1b0
 vfs_read+0xb2/0x170
 ksys_read+0x52/0xc0
 do_syscall_64+0x5c/0xa0
 entry_SYSCALL_64_after_hwframe+0x49/0xbe
RIP: 0033:0x7f2188d8934c
Code: ec 28 48 89 54 24 18 48 89 74 24 10 89 7c 24 08 e8 79 c9 01 00 48 8b =
54
24 18 48 8b 74 24 10 41 89 c0 8b 7c 24 08 31 c0 0f 05 <48> 3d 00 f0 f=
f ff 77 30
44 89 c7 48 89 44 24 08 e8 af c9 01 00 48
RSP: 002b:000000000023f010 EFLAGS: 00000246 ORIG_RAX: 0000000000000000
RAX: ffffffffffffffda RBX: 000000007d11b6d0 RCX: 00007f2188d8934c
RDX: 0000000000000400 RSI: 000000007d0dd4f0 RDI: 000000000000007b
RBP: 0000000000000d68 R08: 0000000000000000 R09: 0000000000000000
R10: 00007f2188621c40 R11: 0000000000000246 R12: 00007f2188e59740
R13: 00007f2188e5a340 R14: 00000000000001ff R15: 000000007d11b6d0

This only occures when I use "amd-staging-drm-next".


You are receiving this mail because:
  • You are the assignee for the bug.
= --15511777420.C59e6E.25471-- --===============1772859566== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVs --===============1772859566==--