From mboxrd@z Thu Jan 1 00:00:00 1970
From: bugzilla-daemon-CC+yJ3UmIYqDUpFQwHEjaQ@public.gmane.org
Subject: [Bug 100567] Nouveau system freeze fifo: SCHED_ERROR 0a
[CTXSW_TIMEOUT]
Date: Sat, 05 Jan 2019 21:29:37 +0000
Message-ID:
References:
Mime-Version: 1.0
Content-Type: multipart/mixed; boundary="===============0186099139=="
Return-path:
In-Reply-To:
List-Unsubscribe: ,
List-Archive:
List-Post:
List-Help:
List-Subscribe: ,
Errors-To: nouveau-bounces-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org
Sender: "Nouveau"
To: nouveau-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org
List-Id: nouveau.vger.kernel.org
--===============0186099139==
Content-Type: multipart/alternative; boundary="15467237774.ec1D6f.438"
Content-Transfer-Encoding: 7bit
--15467237774.ec1D6f.438
Date: Sat, 5 Jan 2019 21:29:37 +0000
MIME-Version: 1.0
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
X-Bugzilla-URL: http://bugs.freedesktop.org/
Auto-Submitted: auto-generated
https://bugs.freedesktop.org/show_bug.cgi?id=3D100567
--- Comment #18 from kenorb-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org ---
The same problem on Ubuntu 18.10, kernel 4.18.0-13.
I've got 4x GPU: GTX 1080 Ti (3-Way SLI Connector), NVIDIA GeForce GTX 1080=
Ti
graphics card with 3584 cores.
$ uname -a
Linux Ubuntu-PC 4.18.0-13-generic #14-Ubuntu SMP Wed Dec 5 09:04:24 UTC 2018
x86_64 x86_64 x86_64 GNU/Linux
Errors in kern.log file:
nouveau 0000:65:00.0: fifo: SCHED_ERROR 0a [CTXSW_TIMEOUT]
nouveau 0000:65:00.0: fifo: runlist 0: scheduled for recovery
nouveau 0000:65:00.0: fifo: channel 2: killed
nouveau 0000:65:00.0: fifo: engine 0: scheduled for recovery
nouveau 0000:65:00.0: Xorg[5447]: channel 2 killed!
nouveau 0000:65:00.0: systemd-logind[3394]: nv50cal_space: -16
nouveau 0000:65:00.0: systemd-logind[3394]: nv50cal_space: -16
(the same message repeated 800x over and over again)
The system got freeze (no mouse or keyboard reaction), however kernel react=
ed
on few Magic SysRq keys, so here are some stack traces:
INFO: task kworker/u72:8:492 blocked for more than 120 seconds.
Tainted: G O 4.18.0-13-generic #14-Ubuntu
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
kworker/u72:8 D 0 492 2 0x80000000
Workqueue: events_unbound nv50_disp_atomic_commit_work [nouveau]
Call Trace at 20:25:50:
__schedule+0x29e/0x840
schedule+0x2c/0x80
schedule_timeout+0x258/0x360
? nv50_wndw_atomic_destroy_state+0x1d/0x20 [nouveau]
dma_fence_default_wait+0x1fc/0x260
? dma_fence_release+0xa0/0xa0
dma_fence_wait_timeout+0x3e/0xf0
drm_atomic_helper_wait_for_fences+0x3f/0xc0 [drm_kms_helper]
nv50_disp_atomic_commit_tail+0x78/0x860 [nouveau]
? __switch_to_asm+0x40/0x70
? __switch_to_asm+0x34/0x70
nv50_disp_atomic_commit_work+0x12/0x20 [nouveau]
process_one_work+0x20f/0x3c0
worker_thread+0x34/0x400
kthread+0x120/0x140
? pwq_unbound_release_workfn+0xd0/0xd0
? kthread_bind+0x40/0x40
ret_from_fork+0x35/0x40
Same call trace at 20:29:51 (few minutes later while Xorg was frozen):
Workqueue: events_unbound nv50_disp_atomic_commit_work [nouveau]
Call Trace:
__schedule+0x29e/0x840
? apic_timer_interrupt+0xa/0x20
? __drm_crtc_commit_free+0x12/0x20 [drm]
schedule+0x2c/0x80
schedule_timeout+0x258/0x360
? nv50_wndw_atomic_destroy_state+0x1d/0x20 [nouveau]
dma_fence_default_wait+0x1fc/0x260
? dma_fence_release+0xa0/0xa0
dma_fence_wait_timeout+0x3e/0xf0
drm_atomic_helper_wait_for_fences+0x3f/0xc0 [drm_kms_helper]
nv50_disp_atomic_commit_tail+0x78/0x860 [nouveau]
? __switch_to_asm+0x40/0x70
? __switch_to_asm+0x34/0x70
nv50_disp_atomic_commit_work+0x12/0x20 [nouveau]
process_one_work+0x20f/0x3c0
worker_thread+0x34/0x400
kthread+0x120/0x140
? pwq_unbound_release_workfn+0xd0/0xd0
? kthread_bind+0x40/0x40
ret_from_fork+0x35/0x40
Another one:
INFO: task Xorg:5447 blocked for more than 120 seconds.
Tainted: G O 4.18.0-13-generic #14-Ubuntu
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
Xorg D 0 5447 5445 0x00000004
Call Trace:
__schedule+0x29e/0x840
schedule+0x2c/0x80
schedule_preempt_disabled+0xe/0x10
__ww_mutex_lock.isra.6+0x3c1/0x660
__ww_mutex_lock_slowpath+0x16/0x20
ww_mutex_lock+0x34/0x50
drm_modeset_lock+0x6e/0xb0 [drm]
drm_crtc_get_sequence_ioctl+0xbc/0x190 [drm]
? drm_wait_vblank_ioctl+0x610/0x610 [drm]
drm_ioctl_kernel+0xa4/0xf0 [drm]
drm_ioctl+0x227/0x400 [drm]
? drm_wait_vblank_ioctl+0x610/0x610 [drm]
? do_iter_write+0xe1/0x1a0
? do_iter_write+0xe1/0x1a0
nouveau_drm_ioctl+0x73/0xc0 [nouveau]
do_vfs_ioctl+0xa8/0x620
? __sys_recvmsg+0x88/0xa0
ksys_ioctl+0x67/0x90
__x64_sys_ioctl+0x1a/0x20
do_syscall_64+0x5a/0x110
entry_SYSCALL_64_after_hwframe+0x44/0xa9
RIP: 0033:0x7f3f654b93c7
Code: Bad RIP value.
RSP: 002b:00007ffd57bbf168 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
RAX: ffffffffffffffda RBX: 00007ffd57bbf200 RCX: 00007f3f654b93c7
RDX: 00007ffd57bbf1a0 RSI: 00000000c018643b RDI: 000000000000000e
RBP: 00007ffd57bbf1a0 R08: 0000000000000000 R09: 00005646eb8ff7c0
R10: 00005646eb54ad30 R11: 0000000000000246 R12: 00000000c018643b
R13: 000000000000000e R14: 00005646eb54b800 R15: 00005646eb466880
Full log: https://gist.github.com/kenorb/5b95caa1694dbf7f030ccc808a110856
--=20
You are receiving this mail because:
You are the assignee for the bug.=
--15467237774.ec1D6f.438
Date: Sat, 5 Jan 2019 21:29:37 +0000
MIME-Version: 1.0
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
X-Bugzilla-URL: http://bugs.freedesktop.org/
Auto-Submitted: auto-generated
Comme=
nt # 18
on bug 10056=
7
from kenorb@gmail.com=
a>
The same problem on Ubuntu 18.10, kernel 4.18.0-13.
I've got 4x GPU: GTX 1080 Ti (3-Way SLI Connector), NVIDIA GeForce GTX 1080=
Ti
graphics card with 3584 cores.
$ uname -a
Linux Ubuntu-PC 4.18.0-13-generic #14-Ubuntu SMP Wed Dec 5 09:04:24 UTC 2018
x86_64 x86_64 x86_64 GNU/Linux
Errors in kern.log file:
nouveau 0000:65:00.0: fifo: SCHED_ERROR 0a [CTXSW_TIMEOUT]
nouveau 0000:65:00.0: fifo: runlist 0: scheduled for recovery
nouveau 0000:65:00.0: fifo: channel 2: killed
nouveau 0000:65:00.0: fifo: engine 0: scheduled for recovery
nouveau 0000:65:00.0: Xorg[5447]: channel 2 killed!
nouveau 0000:65:00.0: systemd-logind[3394]: nv50cal_space: -16
nouveau 0000:65:00.0: systemd-logind[3394]: nv50cal_space: -16
(the same message repeated 800x over and over again)
The system got freeze (no mouse or keyboard reaction), however kernel react=
ed
on few Magic SysRq keys, so here are some stack traces:
INFO: task kworker/u72:8:492 blocked for more than 120 seconds.
Tainted: G O 4.18.0-13-generic #14-Ubuntu
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables th=
is message.
kworker/u72:8 D 0 492 2 0x80000000
Workqueue: events_unbound nv50_disp_atomic_commit_work [nouveau]
Call Trace at 20:25:50:
__schedule+0x29e/0x840
schedule+0x2c/0x80
schedule_timeout+0x258/0x360
? nv50_wndw_atomic_destroy_state+0x1d/0x20 [nouveau]
dma_fence_default_wait+0x1fc/0x260
? dma_fence_release+0xa0/0xa0
dma_fence_wait_timeout+0x3e/0xf0
drm_atomic_helper_wait_for_fences+0x3f/0xc0 [drm_kms_helper]
nv50_disp_atomic_commit_tail+0x78/0x860 [nouveau]
? __switch_to_asm+0x40/0x70
? __switch_to_asm+0x34/0x70
nv50_disp_atomic_commit_work+0x12/0x20 [nouveau]
process_one_work+0x20f/0x3c0
worker_thread+0x34/0x400
kthread+0x120/0x140
? pwq_unbound_release_workfn+0xd0/0xd0
? kthread_bind+0x40/0x40
ret_from_fork+0x35/0x40
Same call trace at 20:29:51 (few minutes later while Xorg was frozen):
Workqueue: events_unbound nv50_disp_atomic_commit_work [nouveau]
Call Trace:
__schedule+0x29e/0x840
? apic_timer_interrupt+0xa/0x20
? __drm_crtc_commit_free+0x12/0x20 [drm]
schedule+0x2c/0x80
schedule_timeout+0x258/0x360
? nv50_wndw_atomic_destroy_state+0x1d/0x20 [nouveau]
dma_fence_default_wait+0x1fc/0x260
? dma_fence_release+0xa0/0xa0
dma_fence_wait_timeout+0x3e/0xf0
drm_atomic_helper_wait_for_fences+0x3f/0xc0 [drm_kms_helper]
nv50_disp_atomic_commit_tail+0x78/0x860 [nouveau]
? __switch_to_asm+0x40/0x70
? __switch_to_asm+0x34/0x70
nv50_disp_atomic_commit_work+0x12/0x20 [nouveau]
process_one_work+0x20f/0x3c0
worker_thread+0x34/0x400
kthread+0x120/0x140
? pwq_unbound_release_workfn+0xd0/0xd0
? kthread_bind+0x40/0x40
ret_from_fork+0x35/0x40
Another one:
INFO: task Xorg:5447 blocked for more than 120 seconds.
Tainted: G O 4.18.0-13-generic #14-Ubuntu
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables th=
is message.
Xorg D 0 5447 5445 0x00000004
Call Trace:
__schedule+0x29e/0x840
schedule+0x2c/0x80
schedule_preempt_disabled+0xe/0x10
__ww_mutex_lock.isra.6+0x3c1/0x660
__ww_mutex_lock_slowpath+0x16/0x20
ww_mutex_lock+0x34/0x50
drm_modeset_lock+0x6e/0xb0 [drm]
drm_crtc_get_sequence_ioctl+0xbc/0x190 [drm]
? drm_wait_vblank_ioctl+0x610/0x610 [drm]
drm_ioctl_kernel+0xa4/0xf0 [drm]
drm_ioctl+0x227/0x400 [drm]
? drm_wait_vblank_ioctl+0x610/0x610 [drm]
? do_iter_write+0xe1/0x1a0
? do_iter_write+0xe1/0x1a0
nouveau_drm_ioctl+0x73/0xc0 [nouveau]
do_vfs_ioctl+0xa8/0x620
? __sys_recvmsg+0x88/0xa0
ksys_ioctl+0x67/0x90
__x64_sys_ioctl+0x1a/0x20
do_syscall_64+0x5a/0x110
entry_SYSCALL_64_after_hwframe+0x44/0xa9
RIP: 0033:0x7f3f654b93c7
Code: Bad RIP value.
RSP: 002b:00007ffd57bbf168 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
RAX: ffffffffffffffda RBX: 00007ffd57bbf200 RCX: 00007f3f654b93c7
RDX: 00007ffd57bbf1a0 RSI: 00000000c018643b RDI: 000000000000000e
RBP: 00007ffd57bbf1a0 R08: 0000000000000000 R09: 00005646eb8ff7c0
R10: 00005646eb54ad30 R11: 0000000000000246 R12: 00000000c018643b
R13: 000000000000000e R14: 00005646eb54b800 R15: 00005646eb466880
Full log: https://gist.github.com/kenorb/5b95caa1694dbf7f030ccc808a110856<=
/a>
You are receiving this mail because:
- You are the assignee for the bug.
=
--15467237774.ec1D6f.438--
--===============0186099139==
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
Content-Disposition: inline
X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KTm91dmVhdSBt
YWlsaW5nIGxpc3QKTm91dmVhdUBsaXN0cy5mcmVlZGVza3RvcC5vcmcKaHR0cHM6Ly9saXN0cy5m
cmVlZGVza3RvcC5vcmcvbWFpbG1hbi9saXN0aW5mby9ub3V2ZWF1Cg==
--===============0186099139==--