* [syzbot] [nfs?] INFO: task hung in nfsd_umount
@ 2024-07-07 4:37 syzbot
2024-07-07 10:49 ` Jeff Layton
2026-08-16 10:14 ` syzbot
0 siblings, 2 replies; 16+ messages in thread
From: syzbot @ 2024-07-07 4:37 UTC (permalink / raw)
To: Dai.Ngo, chuck.lever, jlayton, kolga, linux-kernel, linux-nfs,
neilb, syzkaller-bugs, tom
Hello,
syzbot found the following issue on:
HEAD commit: 1dd28064d416 Merge tag 'integrity-v6.10-fix' of ssh://ra.k..
git tree: upstream
console output: https://syzkaller.appspot.com/x/log.txt?x=16814d76980000
kernel config: https://syzkaller.appspot.com/x/.config?x=1ace69f521989b1f
dashboard link: https://syzkaller.appspot.com/bug?extid=b568ba42c85a332a88ee
compiler: Debian clang version 15.0.6, GNU ld (GNU Binutils for Debian) 2.40
Unfortunately, I don't have any reproducer for this issue yet.
Downloadable assets:
disk image: https://storage.googleapis.com/syzbot-assets/20c723869b92/disk-1dd28064.raw.xz
vmlinux: https://storage.googleapis.com/syzbot-assets/c2a9a7382516/vmlinux-1dd28064.xz
kernel image: https://storage.googleapis.com/syzbot-assets/9872fe6f2853/bzImage-1dd28064.xz
IMPORTANT: if you fix the issue, please add the following tag to the commit:
Reported-by: syzbot+b568ba42c85a332a88ee@syzkaller.appspotmail.com
INFO: task syz.0.2871:13167 blocked for more than 143 seconds.
Not tainted 6.10.0-rc6-syzkaller-00212-g1dd28064d416 #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz.0.2871 state:D stack:26704 pid:13167 tgid:13165 ppid:10304 flags:0x00000004
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5408 [inline]
__schedule+0x1796/0x49d0 kernel/sched/core.c:6745
__schedule_loop kernel/sched/core.c:6822 [inline]
schedule+0x14b/0x320 kernel/sched/core.c:6837
schedule_preempt_disabled+0x13/0x30 kernel/sched/core.c:6894
__mutex_lock_common kernel/locking/mutex.c:684 [inline]
__mutex_lock+0x6a4/0xd70 kernel/locking/mutex.c:752
nfsd_shutdown_threads+0x4e/0xd0 fs/nfsd/nfssvc.c:632
nfsd_umount+0x43/0xd0 fs/nfsd/nfsctl.c:1412
deactivate_locked_super+0xc4/0x130 fs/super.c:473
put_fs_context+0x94/0x780 fs/fs_context.c:516
fscontext_release+0x65/0x80 fs/fsopen.c:73
__fput+0x24a/0x8a0 fs/file_table.c:422
task_work_run+0x24f/0x310 kernel/task_work.c:180
resume_user_mode_work include/linux/resume_user_mode.h:50 [inline]
exit_to_user_mode_loop kernel/entry/common.c:114 [inline]
exit_to_user_mode_prepare include/linux/entry-common.h:328 [inline]
__syscall_exit_to_user_mode_work kernel/entry/common.c:207 [inline]
syscall_exit_to_user_mode+0x168/0x360 kernel/entry/common.c:218
do_syscall_64+0x100/0x230 arch/x86/entry/common.c:89
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f11fe775bd9
RSP: 002b:00007f11ff494048 EFLAGS: 00000246 ORIG_RAX: 00000000000001b4
RAX: 0000000000000000 RBX: 00007f11fe903f60 RCX: 00007f11fe775bd9
RDX: 0000000000000000 RSI: ffffffffffffffff RDI: 0000000000000003
RBP: 00007f11fe7e4aa1 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 000000000000000b R14: 00007f11fe903f60 R15: 00007ffc04bba328
</TASK>
Showing all locks held in the system:
1 lock held by khungtaskd/30:
#0: ffffffff8e333f20 (rcu_read_lock){....}-{1:2}, at: rcu_lock_acquire include/linux/rcupdate.h:329 [inline]
#0: ffffffff8e333f20 (rcu_read_lock){....}-{1:2}, at: rcu_read_lock include/linux/rcupdate.h:781 [inline]
#0: ffffffff8e333f20 (rcu_read_lock){....}-{1:2}, at: debug_show_all_locks+0x55/0x2a0 kernel/locking/lockdep.c:6614
3 locks held by kworker/u8:4/66:
#0: ffff888015089148 ((wq_completion)events_unbound){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3223 [inline]
#0: ffff888015089148 ((wq_completion)events_unbound){+.+.}-{0:0}, at: process_scheduled_works+0x90a/0x1830 kernel/workqueue.c:3329
#1: ffffc900020afd00 ((work_completion)(&map->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3224 [inline]
#1: ffffc900020afd00 ((work_completion)(&map->work)){+.+.}-{0:0}, at: process_scheduled_works+0x945/0x1830 kernel/workqueue.c:3329
#2: ffffffff8e3391c0 (rcu_state.barrier_mutex){+.+.}-{3:3}, at: rcu_barrier+0x4c/0x530 kernel/rcu/tree.c:4448
4 locks held by udevd/4534:
#0: ffff88805a0fdd58 (&p->lock){+.+.}-{3:3}, at: seq_read_iter+0xb7/0xd60 fs/seq_file.c:182
#1: ffff88806542ec88 (&of->mutex#2){+.+.}-{3:3}, at: kernfs_seq_start+0x53/0x3b0 fs/kernfs/file.c:154
#2: ffff88805af532d8 (kn->active#5){++++}-{0:0}, at: kernfs_seq_start+0x72/0x3b0 fs/kernfs/file.c:155
#3: ffff8880686c90e8 (&dev->mutex){....}-{3:3}, at: device_lock include/linux/device.h:1009 [inline]
#3: ffff8880686c90e8 (&dev->mutex){....}-{3:3}, at: uevent_show+0x17d/0x340 drivers/base/core.c:2743
2 locks held by getty/4842:
#0: ffff88802f9d10a0 (&tty->ldisc_sem){++++}-{0:0}, at: tty_ldisc_ref_wait+0x25/0x70 drivers/tty/tty_ldisc.c:243
#1: ffffc9000312b2f0 (&ldata->atomic_read_lock){+.+.}-{3:3}, at: n_tty_read+0x6b5/0x1e10 drivers/tty/n_tty.c:2211
3 locks held by kworker/u9:10/5101:
#0: ffff88802e1fa148 ((wq_completion)hci2){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3223 [inline]
#0: ffff88802e1fa148 ((wq_completion)hci2){+.+.}-{0:0}, at: process_scheduled_works+0x90a/0x1830 kernel/workqueue.c:3329
#1: ffffc90003827d00 ((work_completion)(&hdev->cmd_sync_work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3224 [inline]
#1: ffffc90003827d00 ((work_completion)(&hdev->cmd_sync_work)){+.+.}-{0:0}, at: process_scheduled_works+0x945/0x1830 kernel/workqueue.c:3329
#2: ffff888020eb4d88 (&hdev->req_lock){+.+.}-{3:3}, at: hci_cmd_sync_work+0x1ec/0x400 net/bluetooth/hci_sync.c:322
2 locks held by syz.4.2691/12623:
#0: ffffffff8f63ad70 (cb_lock){++++}-{3:3}, at: genl_rcv+0x19/0x40 net/netlink/genetlink.c:1218
#1: ffffffff8e5fff48 (nfsd_mutex){+.+.}-{3:3}, at: nfsd_nl_listener_set_doit+0x12d/0x1a90 fs/nfsd/nfsctl.c:1940
2 locks held by syz.0.2871/13167:
#0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: __super_lock fs/super.c:56 [inline]
#0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: __super_lock_excl fs/super.c:71 [inline]
#0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: deactivate_super+0xb5/0xf0 fs/super.c:505
#1: ffffffff8e5fff48 (nfsd_mutex){+.+.}-{3:3}, at: nfsd_shutdown_threads+0x4e/0xd0 fs/nfsd/nfssvc.c:632
1 lock held by udevd/14823:
2 locks held by syz-executor/15993:
#0: ffffffff8f5d4908 (rtnl_mutex){+.+.}-{3:3}, at: rtnl_lock net/core/rtnetlink.c:79 [inline]
#0: ffffffff8f5d4908 (rtnl_mutex){+.+.}-{3:3}, at: rtnetlink_rcv_msg+0x842/0x1180 net/core/rtnetlink.c:6632
#1: ffffffff8e3392f8 (rcu_state.exp_mutex){+.+.}-{3:3}, at: exp_funnel_lock kernel/rcu/tree_exp.h:323 [inline]
#1: ffffffff8e3392f8 (rcu_state.exp_mutex){+.+.}-{3:3}, at: synchronize_rcu_expedited+0x451/0x830 kernel/rcu/tree_exp.h:939
1 lock held by syz.4.3895/16063:
#0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: __super_lock fs/super.c:58 [inline]
#0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: super_lock+0x27c/0x400 fs/super.c:120
4 locks held by kvm-nx-lpage-re/16077:
#0: ffffffff8e361f28 (cgroup_mutex){+.+.}-{3:3}, at: cgroup_lock include/linux/cgroup.h:368 [inline]
#0: ffffffff8e361f28 (cgroup_mutex){+.+.}-{3:3}, at: cgroup_attach_task_all+0x27/0xe0 kernel/cgroup/cgroup-v1.c:61
#1: ffffffff8e1ce5b0 (cpu_hotplug_lock){++++}-{0:0}, at: cgroup_attach_lock+0x11/0x40 kernel/cgroup/cgroup.c:2413
#2: ffffffff8e362110 (cgroup_threadgroup_rwsem){++++}-{0:0}, at: cgroup_attach_task_all+0x31/0xe0 kernel/cgroup/cgroup-v1.c:62
#3: ffffffff8e3392f8 (rcu_state.exp_mutex){+.+.}-{3:3}, at: exp_funnel_lock kernel/rcu/tree_exp.h:291 [inline]
#3: ffffffff8e3392f8 (rcu_state.exp_mutex){+.+.}-{3:3}, at: synchronize_rcu_expedited+0x381/0x830 kernel/rcu/tree_exp.h:939
=============================================
NMI backtrace for cpu 1
CPU: 1 PID: 30 Comm: khungtaskd Not tainted 6.10.0-rc6-syzkaller-00212-g1dd28064d416 #0
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 06/07/2024
Call Trace:
<TASK>
__dump_stack lib/dump_stack.c:88 [inline]
dump_stack_lvl+0x241/0x360 lib/dump_stack.c:114
nmi_cpu_backtrace+0x49c/0x4d0 lib/nmi_backtrace.c:113
nmi_trigger_cpumask_backtrace+0x198/0x320 lib/nmi_backtrace.c:62
trigger_all_cpu_backtrace include/linux/nmi.h:162 [inline]
check_hung_uninterruptible_tasks kernel/hung_task.c:223 [inline]
watchdog+0xfde/0x1020 kernel/hung_task.c:379
kthread+0x2f0/0x390 kernel/kthread.c:389
ret_from_fork+0x4b/0x80 arch/x86/kernel/process.c:147
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:244
</TASK>
Sending NMI from CPU 1 to CPUs 0:
NMI backtrace for cpu 0
CPU: 0 PID: 1109 Comm: kworker/u8:6 Not tainted 6.10.0-rc6-syzkaller-00212-g1dd28064d416 #0
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 06/07/2024
Workqueue: writeback wb_workfn (flush-8:0)
RIP: 0010:bio_add_page+0xb7/0x840 block/bio.c:1118
Code: 05 00 00 8b 5d 00 45 89 f4 41 f7 d4 89 df 44 89 e6 e8 dd 6f 11 fd 44 39 e3 76 0a e8 13 6e 11 fd e9 20 04 00 00 48 89 6c 24 28 <49> 8d 5f 70 48 89 d8 48 c1 e8 03 48 89 44 24 30 42 0f b6 04 28 84
RSP: 0018:ffffc900045ce610 EFLAGS: 00000213
RAX: 0000000000000000 RBX: 000000000000a000 RCX: ffff888022311e00
RDX: ffff888022311e00 RSI: 00000000ffffefff RDI: 000000000000a000
RBP: ffff88802eb37a28 R08: ffffffff8484b8a3 R09: 1ffffd4000090b30
R10: dffffc0000000000 R11: fffff94000090b31 R12: 00000000ffffefff
R13: dffffc0000000000 R14: 0000000000001000 R15: ffff88802eb37a00
FS: 0000000000000000(0000) GS:ffff8880b9400000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 0000001b30e11ff8 CR3: 000000000e132000 CR4: 00000000003526f0
DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 0000000000000400
Call Trace:
<NMI>
</NMI>
<TASK>
bio_add_folio+0x52/0x80 block/bio.c:1160
io_submit_add_bh fs/ext4/page-io.c:422 [inline]
ext4_bio_write_folio+0x1691/0x1da0 fs/ext4/page-io.c:560
mpage_submit_folio+0x1af/0x230 fs/ext4/inode.c:1869
mpage_process_page_bufs+0x6c9/0x8d0 fs/ext4/inode.c:1982
mpage_prepare_extent_to_map+0xec7/0x1c80 fs/ext4/inode.c:2490
ext4_do_writepages+0xc52/0x3d40 fs/ext4/inode.c:2632
ext4_writepages+0x213/0x3c0 fs/ext4/inode.c:2768
do_writepages+0x359/0x870 mm/page-writeback.c:2656
__writeback_single_inode+0x165/0x10b0 fs/fs-writeback.c:1651
writeback_sb_inodes+0x99c/0x1380 fs/fs-writeback.c:1947
__writeback_inodes_wb+0x11b/0x260 fs/fs-writeback.c:2018
wb_writeback+0x495/0xd40 fs/fs-writeback.c:2129
wb_check_old_data_flush fs/fs-writeback.c:2233 [inline]
wb_do_writeback fs/fs-writeback.c:2286 [inline]
wb_workfn+0xba1/0x1090 fs/fs-writeback.c:2314
process_one_work kernel/workqueue.c:3248 [inline]
process_scheduled_works+0xa2c/0x1830 kernel/workqueue.c:3329
worker_thread+0x86d/0xd50 kernel/workqueue.c:3409
kthread+0x2f0/0x390 kernel/kthread.c:389
ret_from_fork+0x4b/0x80 arch/x86/kernel/process.c:147
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:244
</TASK>
---
This report is generated by a bot. It may contain errors.
See https://goo.gl/tpsmEJ for more information about syzbot.
syzbot engineers can be reached at syzkaller@googlegroups.com.
syzbot will keep track of this issue. See:
https://goo.gl/tpsmEJ#status for how to communicate with syzbot.
If the report is already addressed, let syzbot know by replying with:
#syz fix: exact-commit-title
If you want to overwrite report's subsystems, reply with:
#syz set subsystems: new-subsystem
(See the list of subsystem names on the web dashboard)
If the report is a duplicate of another one, reply with:
#syz dup: exact-subject-of-another-report
If you want to undo deduplication, reply with:
#syz undup
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-07-07 4:37 [syzbot] [nfs?] INFO: task hung in nfsd_umount syzbot
@ 2024-07-07 10:49 ` Jeff Layton
2024-07-08 0:07 ` NeilBrown
2026-08-16 10:14 ` syzbot
1 sibling, 1 reply; 16+ messages in thread
From: Jeff Layton @ 2024-07-07 10:49 UTC (permalink / raw)
To: syzbot, Dai.Ngo, chuck.lever, kolga, linux-kernel, linux-nfs,
neilb, syzkaller-bugs, tom
On Sat, 2024-07-06 at 21:37 -0700, syzbot wrote:
> Hello,
>
> syzbot found the following issue on:
>
> HEAD commit: 1dd28064d416 Merge tag 'integrity-v6.10-fix' of ssh://ra.k..
> git tree: upstream
> console output: https://syzkaller.appspot.com/x/log.txt?x=16814d76980000
> kernel config: https://syzkaller.appspot.com/x/.config?x=1ace69f521989b1f
> dashboard link: https://syzkaller.appspot.com/bug?extid=b568ba42c85a332a88ee
> compiler: Debian clang version 15.0.6, GNU ld (GNU Binutils for Debian) 2.40
>
> Unfortunately, I don't have any reproducer for this issue yet.
>
> Downloadable assets:
> disk image: https://storage.googleapis.com/syzbot-assets/20c723869b92/disk-1dd28064.raw.xz
> vmlinux: https://storage.googleapis.com/syzbot-assets/c2a9a7382516/vmlinux-1dd28064.xz
> kernel image: https://storage.googleapis.com/syzbot-assets/9872fe6f2853/bzImage-1dd28064.xz
>
> IMPORTANT: if you fix the issue, please add the following tag to the commit:
> Reported-by: syzbot+b568ba42c85a332a88ee@syzkaller.appspotmail.com
>
> INFO: task syz.0.2871:13167 blocked for more than 143 seconds.
> Not tainted 6.10.0-rc6-syzkaller-00212-g1dd28064d416 #0
> "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
> task:syz.0.2871 state:D stack:26704 pid:13167 tgid:13165 ppid:10304 flags:0x00000004
> Call Trace:
> <TASK>
> context_switch kernel/sched/core.c:5408 [inline]
> __schedule+0x1796/0x49d0 kernel/sched/core.c:6745
> __schedule_loop kernel/sched/core.c:6822 [inline]
> schedule+0x14b/0x320 kernel/sched/core.c:6837
> schedule_preempt_disabled+0x13/0x30 kernel/sched/core.c:6894
> __mutex_lock_common kernel/locking/mutex.c:684 [inline]
> __mutex_lock+0x6a4/0xd70 kernel/locking/mutex.c:752
> nfsd_shutdown_threads+0x4e/0xd0 fs/nfsd/nfssvc.c:632
> nfsd_umount+0x43/0xd0 fs/nfsd/nfsctl.c:1412
> deactivate_locked_super+0xc4/0x130 fs/super.c:473
> put_fs_context+0x94/0x780 fs/fs_context.c:516
> fscontext_release+0x65/0x80 fs/fsopen.c:73
> __fput+0x24a/0x8a0 fs/file_table.c:422
> task_work_run+0x24f/0x310 kernel/task_work.c:180
> resume_user_mode_work include/linux/resume_user_mode.h:50 [inline]
> exit_to_user_mode_loop kernel/entry/common.c:114 [inline]
> exit_to_user_mode_prepare include/linux/entry-common.h:328 [inline]
> __syscall_exit_to_user_mode_work kernel/entry/common.c:207 [inline]
> syscall_exit_to_user_mode+0x168/0x360 kernel/entry/common.c:218
> do_syscall_64+0x100/0x230 arch/x86/entry/common.c:89
> entry_SYSCALL_64_after_hwframe+0x77/0x7f
> RIP: 0033:0x7f11fe775bd9
> RSP: 002b:00007f11ff494048 EFLAGS: 00000246 ORIG_RAX: 00000000000001b4
> RAX: 0000000000000000 RBX: 00007f11fe903f60 RCX: 00007f11fe775bd9
> RDX: 0000000000000000 RSI: ffffffffffffffff RDI: 0000000000000003
> RBP: 00007f11fe7e4aa1 R08: 0000000000000000 R09: 0000000000000000
> R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
> R13: 000000000000000b R14: 00007f11fe903f60 R15: 00007ffc04bba328
> </TASK>
>
> Showing all locks held in the system:
> 1 lock held by khungtaskd/30:
> #0: ffffffff8e333f20 (rcu_read_lock){....}-{1:2}, at: rcu_lock_acquire include/linux/rcupdate.h:329 [inline]
> #0: ffffffff8e333f20 (rcu_read_lock){....}-{1:2}, at: rcu_read_lock include/linux/rcupdate.h:781 [inline]
> #0: ffffffff8e333f20 (rcu_read_lock){....}-{1:2}, at: debug_show_all_locks+0x55/0x2a0 kernel/locking/lockdep.c:6614
> 3 locks held by kworker/u8:4/66:
> #0: ffff888015089148 ((wq_completion)events_unbound){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3223 [inline]
> #0: ffff888015089148 ((wq_completion)events_unbound){+.+.}-{0:0}, at: process_scheduled_works+0x90a/0x1830 kernel/workqueue.c:3329
> #1: ffffc900020afd00 ((work_completion)(&map->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3224 [inline]
> #1: ffffc900020afd00 ((work_completion)(&map->work)){+.+.}-{0:0}, at: process_scheduled_works+0x945/0x1830 kernel/workqueue.c:3329
> #2: ffffffff8e3391c0 (rcu_state.barrier_mutex){+.+.}-{3:3}, at: rcu_barrier+0x4c/0x530 kernel/rcu/tree.c:4448
> 4 locks held by udevd/4534:
> #0: ffff88805a0fdd58 (&p->lock){+.+.}-{3:3}, at: seq_read_iter+0xb7/0xd60 fs/seq_file.c:182
> #1: ffff88806542ec88 (&of->mutex#2){+.+.}-{3:3}, at: kernfs_seq_start+0x53/0x3b0 fs/kernfs/file.c:154
> #2: ffff88805af532d8 (kn->active#5){++++}-{0:0}, at: kernfs_seq_start+0x72/0x3b0 fs/kernfs/file.c:155
> #3: ffff8880686c90e8 (&dev->mutex){....}-{3:3}, at: device_lock include/linux/device.h:1009 [inline]
> #3: ffff8880686c90e8 (&dev->mutex){....}-{3:3}, at: uevent_show+0x17d/0x340 drivers/base/core.c:2743
> 2 locks held by getty/4842:
> #0: ffff88802f9d10a0 (&tty->ldisc_sem){++++}-{0:0}, at: tty_ldisc_ref_wait+0x25/0x70 drivers/tty/tty_ldisc.c:243
> #1: ffffc9000312b2f0 (&ldata->atomic_read_lock){+.+.}-{3:3}, at: n_tty_read+0x6b5/0x1e10 drivers/tty/n_tty.c:2211
> 3 locks held by kworker/u9:10/5101:
> #0: ffff88802e1fa148 ((wq_completion)hci2){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3223 [inline]
> #0: ffff88802e1fa148 ((wq_completion)hci2){+.+.}-{0:0}, at: process_scheduled_works+0x90a/0x1830 kernel/workqueue.c:3329
> #1: ffffc90003827d00 ((work_completion)(&hdev->cmd_sync_work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3224 [inline]
> #1: ffffc90003827d00 ((work_completion)(&hdev->cmd_sync_work)){+.+.}-{0:0}, at: process_scheduled_works+0x945/0x1830 kernel/workqueue.c:3329
> #2: ffff888020eb4d88 (&hdev->req_lock){+.+.}-{3:3}, at: hci_cmd_sync_work+0x1ec/0x400 net/bluetooth/hci_sync.c:322
> 2 locks held by syz.4.2691/12623:
> #0: ffffffff8f63ad70 (cb_lock){++++}-{3:3}, at: genl_rcv+0x19/0x40 net/netlink/genetlink.c:1218
> #1: ffffffff8e5fff48 (nfsd_mutex){+.+.}-{3:3}, at: nfsd_nl_listener_set_doit+0x12d/0x1a90 fs/nfsd/nfsctl.c:1940
> 2 locks held by syz.0.2871/13167:
> #0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: __super_lock fs/super.c:56 [inline]
> #0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: __super_lock_excl fs/super.c:71 [inline]
> #0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: deactivate_super+0xb5/0xf0 fs/super.c:505
> #1: ffffffff8e5fff48 (nfsd_mutex){+.+.}-{3:3}, at: nfsd_shutdown_threads+0x4e/0xd0 fs/nfsd/nfssvc.c:632
Above there are two tasks that are holding/blocked on the nfsd_mutex.
The first is the stuck thread. The second is a task that took it in
nfsd_nl_listener_set_doit. I don't see a way to exit that function
without releasing that mutex, so it seems likely that thread is stuck
too for some reason. Unfortunately, we don't have a stack trace from
that task, so I can't tell what it's doing.
> 1 lock held by udevd/14823:
> 2 locks held by syz-executor/15993:
> #0: ffffffff8f5d4908 (rtnl_mutex){+.+.}-{3:3}, at: rtnl_lock net/core/rtnetlink.c:79 [inline]
> #0: ffffffff8f5d4908 (rtnl_mutex){+.+.}-{3:3}, at: rtnetlink_rcv_msg+0x842/0x1180 net/core/rtnetlink.c:6632
> #1: ffffffff8e3392f8 (rcu_state.exp_mutex){+.+.}-{3:3}, at: exp_funnel_lock kernel/rcu/tree_exp.h:323 [inline]
> #1: ffffffff8e3392f8 (rcu_state.exp_mutex){+.+.}-{3:3}, at: synchronize_rcu_expedited+0x451/0x830 kernel/rcu/tree_exp.h:939
> 1 lock held by syz.4.3895/16063:
> #0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: __super_lock fs/super.c:58 [inline]
> #0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: super_lock+0x27c/0x400 fs/super.c:120
> 4 locks held by kvm-nx-lpage-re/16077:
> #0: ffffffff8e361f28 (cgroup_mutex){+.+.}-{3:3}, at: cgroup_lock include/linux/cgroup.h:368 [inline]
> #0: ffffffff8e361f28 (cgroup_mutex){+.+.}-{3:3}, at: cgroup_attach_task_all+0x27/0xe0 kernel/cgroup/cgroup-v1.c:61
> #1: ffffffff8e1ce5b0 (cpu_hotplug_lock){++++}-{0:0}, at: cgroup_attach_lock+0x11/0x40 kernel/cgroup/cgroup.c:2413
> #2: ffffffff8e362110 (cgroup_threadgroup_rwsem){++++}-{0:0}, at: cgroup_attach_task_all+0x31/0xe0 kernel/cgroup/cgroup-v1.c:62
> #3: ffffffff8e3392f8 (rcu_state.exp_mutex){+.+.}-{3:3}, at: exp_funnel_lock kernel/rcu/tree_exp.h:291 [inline]
> #3: ffffffff8e3392f8 (rcu_state.exp_mutex){+.+.}-{3:3}, at: synchronize_rcu_expedited+0x381/0x830 kernel/rcu/tree_exp.h:939
>
> =============================================
>
> NMI backtrace for cpu 1
> CPU: 1 PID: 30 Comm: khungtaskd Not tainted 6.10.0-rc6-syzkaller-00212-g1dd28064d416 #0
> Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 06/07/2024
> Call Trace:
> <TASK>
> __dump_stack lib/dump_stack.c:88 [inline]
> dump_stack_lvl+0x241/0x360 lib/dump_stack.c:114
> nmi_cpu_backtrace+0x49c/0x4d0 lib/nmi_backtrace.c:113
> nmi_trigger_cpumask_backtrace+0x198/0x320 lib/nmi_backtrace.c:62
> trigger_all_cpu_backtrace include/linux/nmi.h:162 [inline]
> check_hung_uninterruptible_tasks kernel/hung_task.c:223 [inline]
> watchdog+0xfde/0x1020 kernel/hung_task.c:379
> kthread+0x2f0/0x390 kernel/kthread.c:389
> ret_from_fork+0x4b/0x80 arch/x86/kernel/process.c:147
> ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:244
> </TASK>
> Sending NMI from CPU 1 to CPUs 0:
> NMI backtrace for cpu 0
> CPU: 0 PID: 1109 Comm: kworker/u8:6 Not tainted 6.10.0-rc6-syzkaller-00212-g1dd28064d416 #0
> Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 06/07/2024
> Workqueue: writeback wb_workfn (flush-8:0)
> RIP: 0010:bio_add_page+0xb7/0x840 block/bio.c:1118
> Code: 05 00 00 8b 5d 00 45 89 f4 41 f7 d4 89 df 44 89 e6 e8 dd 6f 11 fd 44 39 e3 76 0a e8 13 6e 11 fd e9 20 04 00 00 48 89 6c 24 28 <49> 8d 5f 70 48 89 d8 48 c1 e8 03 48 89 44 24 30 42 0f b6 04 28 84
> RSP: 0018:ffffc900045ce610 EFLAGS: 00000213
> RAX: 0000000000000000 RBX: 000000000000a000 RCX: ffff888022311e00
> RDX: ffff888022311e00 RSI: 00000000ffffefff RDI: 000000000000a000
> RBP: ffff88802eb37a28 R08: ffffffff8484b8a3 R09: 1ffffd4000090b30
> R10: dffffc0000000000 R11: fffff94000090b31 R12: 00000000ffffefff
> R13: dffffc0000000000 R14: 0000000000001000 R15: ffff88802eb37a00
> FS: 0000000000000000(0000) GS:ffff8880b9400000(0000) knlGS:0000000000000000
> CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> CR2: 0000001b30e11ff8 CR3: 000000000e132000 CR4: 00000000003526f0
> DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
> DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 0000000000000400
> Call Trace:
> <NMI>
> </NMI>
> <TASK>
> bio_add_folio+0x52/0x80 block/bio.c:1160
> io_submit_add_bh fs/ext4/page-io.c:422 [inline]
> ext4_bio_write_folio+0x1691/0x1da0 fs/ext4/page-io.c:560
> mpage_submit_folio+0x1af/0x230 fs/ext4/inode.c:1869
> mpage_process_page_bufs+0x6c9/0x8d0 fs/ext4/inode.c:1982
> mpage_prepare_extent_to_map+0xec7/0x1c80 fs/ext4/inode.c:2490
> ext4_do_writepages+0xc52/0x3d40 fs/ext4/inode.c:2632
> ext4_writepages+0x213/0x3c0 fs/ext4/inode.c:2768
> do_writepages+0x359/0x870 mm/page-writeback.c:2656
> __writeback_single_inode+0x165/0x10b0 fs/fs-writeback.c:1651
> writeback_sb_inodes+0x99c/0x1380 fs/fs-writeback.c:1947
> __writeback_inodes_wb+0x11b/0x260 fs/fs-writeback.c:2018
> wb_writeback+0x495/0xd40 fs/fs-writeback.c:2129
> wb_check_old_data_flush fs/fs-writeback.c:2233 [inline]
> wb_do_writeback fs/fs-writeback.c:2286 [inline]
> wb_workfn+0xba1/0x1090 fs/fs-writeback.c:2314
> process_one_work kernel/workqueue.c:3248 [inline]
> process_scheduled_works+0xa2c/0x1830 kernel/workqueue.c:3329
> worker_thread+0x86d/0xd50 kernel/workqueue.c:3409
> kthread+0x2f0/0x390 kernel/kthread.c:389
> ret_from_fork+0x4b/0x80 arch/x86/kernel/process.c:147
> ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:244
> </TASK>
>
>
> ---
> This report is generated by a bot. It may contain errors.
> See https://goo.gl/tpsmEJ for more information about syzbot.
> syzbot engineers can be reached at syzkaller@googlegroups.com.
>
> syzbot will keep track of this issue. See:
> https://goo.gl/tpsmEJ#status for how to communicate with syzbot.
>
> If the report is already addressed, let syzbot know by replying with:
> #syz fix: exact-commit-title
>
> If you want to overwrite report's subsystems, reply with:
> #syz set subsystems: new-subsystem
> (See the list of subsystem names on the web dashboard)
>
> If the report is a duplicate of another one, reply with:
> #syz dup: exact-subject-of-another-report
>
> If you want to undo deduplication, reply with:
> #syz undup
--
Jeff Layton <jlayton@kernel.org>
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-07-07 10:49 ` Jeff Layton
@ 2024-07-08 0:07 ` NeilBrown
2024-09-21 7:58 ` Harald Dunkel
0 siblings, 1 reply; 16+ messages in thread
From: NeilBrown @ 2024-07-08 0:07 UTC (permalink / raw)
To: syzbot
Cc: Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-kernel, linux-nfs,
syzkaller-bugs, tom
On Sun, 07 Jul 2024, Jeff Layton wrote:
> On Sat, 2024-07-06 at 21:37 -0700, syzbot wrote:
> > Hello,
> >
> > syzbot found the following issue on:
> >
> > HEAD commit: 1dd28064d416 Merge tag 'integrity-v6.10-fix' of ssh://ra.k..
...
> > 2 locks held by syz.4.2691/12623:
> > #0: ffffffff8f63ad70 (cb_lock){++++}-{3:3}, at: genl_rcv+0x19/0x40 net/netlink/genetlink.c:1218
> > #1: ffffffff8e5fff48 (nfsd_mutex){+.+.}-{3:3}, at: nfsd_nl_listener_set_doit+0x12d/0x1a90 fs/nfsd/nfsctl.c:1940
> > 2 locks held by syz.0.2871/13167:
> > #0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: __super_lock fs/super.c:56 [inline]
> > #0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: __super_lock_excl fs/super.c:71 [inline]
> > #0: ffff88804e1920e0 (&type->s_umount_key#77){++++}-{3:3}, at: deactivate_super+0xb5/0xf0 fs/super.c:505
> > #1: ffffffff8e5fff48 (nfsd_mutex){+.+.}-{3:3}, at: nfsd_shutdown_threads+0x4e/0xd0 fs/nfsd/nfssvc.c:632
>
> Above there are two tasks that are holding/blocked on the nfsd_mutex.
> The first is the stuck thread. The second is a task that took it in
> nfsd_nl_listener_set_doit. I don't see a way to exit that function
> without releasing that mutex, so it seems likely that thread is stuck
> too for some reason. Unfortunately, we don't have a stack trace from
> that task, so I can't tell what it's doing.
>
We can guess though. It isn't waiting for a lock - that would show in
the above list - so it might be waiting for a wakeup, or might be
spinning.
The only wake-up I can imagine is in one of the memory-allocation calls,
but if the system were running out of memory we would probably see
messages about that.
I wonder if it could be looping in svc_xprt_destroy_all(), and sitting
in the msleep() when the hang is detected so there are no locks to
report. I can't see while it would block there.
It would really help to get a full task list.
There is a sysctl for that: /proc/sys/kernel/hung_task_all_cpu_backtrace
Could that be enabled?
NeilBrown
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-07-08 0:07 ` NeilBrown
@ 2024-09-21 7:58 ` Harald Dunkel
2024-09-28 7:41 ` Harald Dunkel
0 siblings, 1 reply; 16+ messages in thread
From: Harald Dunkel @ 2024-09-21 7:58 UTC (permalink / raw)
To: NeilBrown, syzbot
Cc: Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-kernel, linux-nfs,
syzkaller-bugs, tom
NeilBrown wrote:
>
> We can guess though. It isn't waiting for a lock - that would show in
> the above list - so it might be waiting for a wakeup, or might be
> spinning.
> The only wake-up I can imagine is in one of the memory-allocation calls,
> but if the system were running out of memory we would probably see
> messages about that.
>
I have seen something like this. I am running NFS inside a container,
using legacy cgroup. When it got stuck it claimed I cannot login
into the container due to out of memory. When it happens again, I
can send you the exact error message. The next hung nfsd is overdue,
anyway.
> I wonder if it could be looping in svc_xprt_destroy_all(), and sitting
> in the msleep() when the hang is detected so there are no locks to
> report. I can't see while it would block there.
>
> It would really help to get a full task list.
> There is a sysctl for that: /proc/sys/kernel/hung_task_all_cpu_backtrace
>
> Could that be enabled?
>
I have enabled it on my NFS server (echo 1 >/proc/.../hung_task_all_cpu_backtrace).
Regards
Harri
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-09-21 7:58 ` Harald Dunkel
@ 2024-09-28 7:41 ` Harald Dunkel
2024-09-28 22:23 ` NeilBrown
0 siblings, 1 reply; 16+ messages in thread
From: Harald Dunkel @ 2024-09-28 7:41 UTC (permalink / raw)
To: NeilBrown, syzbot
Cc: Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
[-- Attachment #1: Type: text/plain, Size: 1125 bytes --]
Hi folks,
On 2024-09-21 09:58:55, Harald Dunkel wrote:
> NeilBrown wrote:
>>
>> We can guess though. It isn't waiting for a lock - that would show in
>> the above list - so it might be waiting for a wakeup, or might be
>> spinning.
>> The only wake-up I can imagine is in one of the memory-allocation calls,
>> but if the system were running out of memory we would probably see
>> messages about that.
>>
>
> I have seen something like this. I am running NFS inside a container,
> using legacy cgroup. When it got stuck it claimed I cannot login
> into the container due to out of memory. When it happens again, I
> can send you the exact error message. The next hung nfsd is overdue,
> anyway.
>
my NFS server got stuck again last night. Unfortunately the service was
recovered by a colleague, so I had no chance to check the memory. Attached
you can find the log files of both nfs container and LXC server, with
/proc/sys/kernel/hung_task_all_cpu_backtrace set to 1.
I dropped the kernel mailing list from this reply, due to large attachments.
Hopefully this was OK?
Hope this helps. Please mail if I can help
Harri
[-- Attachment #2: log.nfs01.txt.gz --]
[-- Type: application/gzip, Size: 3806 bytes --]
[-- Attachment #3: log.nasl006.txt.gz --]
[-- Type: application/gzip, Size: 4870 bytes --]
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-09-28 7:41 ` Harald Dunkel
@ 2024-09-28 22:23 ` NeilBrown
2024-09-29 8:23 ` Harald Dunkel
0 siblings, 1 reply; 16+ messages in thread
From: NeilBrown @ 2024-09-28 22:23 UTC (permalink / raw)
To: Harald Dunkel
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
On Sat, 28 Sep 2024, Harald Dunkel wrote:
> Hi folks,
>
> On 2024-09-21 09:58:55, Harald Dunkel wrote:
> > NeilBrown wrote:
> >>
> >> We can guess though. It isn't waiting for a lock - that would show in
> >> the above list - so it might be waiting for a wakeup, or might be
> >> spinning.
> >> The only wake-up I can imagine is in one of the memory-allocation calls,
> >> but if the system were running out of memory we would probably see
> >> messages about that.
> >>
> >
> > I have seen something like this. I am running NFS inside a container,
> > using legacy cgroup. When it got stuck it claimed I cannot login
> > into the container due to out of memory. When it happens again, I
> > can send you the exact error message. The next hung nfsd is overdue,
> > anyway.
> >
>
> my NFS server got stuck again last night. Unfortunately the service was
> recovered by a colleague, so I had no chance to check the memory. Attached
> you can find the log files of both nfs container and LXC server, with
> /proc/sys/kernel/hung_task_all_cpu_backtrace set to 1.
>
> I dropped the kernel mailing list from this reply, due to large attachments.
> Hopefully this was OK?
>
> Hope this helps. Please mail if I can help
> Harri
>
Thanks for the logs. The point to flush_workqueue() being a problem,
presumably from nfsd4_probe_callback_sync(), though I'm not 100% sure of
that. Maybe some deadlock in the callback code. I'm not very familiar
with that code and nothing immediately jumps out.
I had thought that hung_task_all_cpu_backtrace would show a backtrace of
*all* tasks - I missed the "cpu" in there.
If if it happens again and if you can
echo t > /proc/sysrq-trigger
to get stack traces of everything, that might help. Maybe it won't be
necessary if I or someone else can spot a deadlock with
flush_workqueue().
Thanks,
NeilBrown
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-09-28 22:23 ` NeilBrown
@ 2024-09-29 8:23 ` Harald Dunkel
2024-09-29 9:59 ` NeilBrown
0 siblings, 1 reply; 16+ messages in thread
From: Harald Dunkel @ 2024-09-29 8:23 UTC (permalink / raw)
To: NeilBrown
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
Hi Neil,
On 2024-09-29 00:23:18, NeilBrown wrote:
>
> Thanks for the logs. The point to flush_workqueue() being a problem,
> presumably from nfsd4_probe_callback_sync(), though I'm not 100% sure of
> that. Maybe some deadlock in the callback code. I'm not very familiar
> with that code and nothing immediately jumps out.
>
> I had thought that hung_task_all_cpu_backtrace would show a backtrace of
> *all* tasks - I missed the "cpu" in there.
> If if it happens again and if you can
> echo t > /proc/sysrq-trigger
> to get stack traces of everything, that might help. Maybe it won't be
> necessary if I or someone else can spot a deadlock with
> flush_workqueue().
>
I just learned that kernel.hung_task_panic = 1 should have been set, too.
Sorry for that. Currently I have
kernel.hung_task_panic = 1
kernel.hung_task_all_cpu_backtrace = 1
Please confirm.
I have set a watchdog to run the sysrq trigger on the NFS server if df
on an NFS client doesn't respond.
Regards
Harri
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-09-29 8:23 ` Harald Dunkel
@ 2024-09-29 9:59 ` NeilBrown
2024-10-01 10:21 ` Harald Dunkel
0 siblings, 1 reply; 16+ messages in thread
From: NeilBrown @ 2024-09-29 9:59 UTC (permalink / raw)
To: Harald Dunkel
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
On Sun, 29 Sep 2024, Harald Dunkel wrote:
> Hi Neil,
>
> On 2024-09-29 00:23:18, NeilBrown wrote:
> >
> > Thanks for the logs. The point to flush_workqueue() being a problem,
> > presumably from nfsd4_probe_callback_sync(), though I'm not 100% sure of
> > that. Maybe some deadlock in the callback code. I'm not very familiar
> > with that code and nothing immediately jumps out.
> >
> > I had thought that hung_task_all_cpu_backtrace would show a backtrace of
> > *all* tasks - I missed the "cpu" in there.
> > If if it happens again and if you can
> > echo t > /proc/sysrq-trigger
> > to get stack traces of everything, that might help. Maybe it won't be
> > necessary if I or someone else can spot a deadlock with
> > flush_workqueue().
> >
>
> I just learned that kernel.hung_task_panic = 1 should have been set, too.
> Sorry for that. Currently I have
>
> kernel.hung_task_panic = 1
> kernel.hung_task_all_cpu_backtrace = 1
>
> Please confirm.
You DON'T want hung_task_panic in this case. If the system panics, it
might do so before the watchdog you mention below fires. We really need
the sysrq-trigger output and the panic might interfere with that.
Thanks,
NeilBrown
>
> I have set a watchdog to run the sysrq trigger on the NFS server if df
> on an NFS client doesn't respond.
>
>
> Regards
> Harri
>
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-09-29 9:59 ` NeilBrown
@ 2024-10-01 10:21 ` Harald Dunkel
2024-10-02 13:55 ` Harald Dunkel
0 siblings, 1 reply; 16+ messages in thread
From: Harald Dunkel @ 2024-10-01 10:21 UTC (permalink / raw)
To: NeilBrown
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
NeilBrown wrote:
> On Sun, 29 Sep 2024, Harald Dunkel wrote:
>>
>> I just learned that kernel.hung_task_panic = 1 should have been set, too.
>> Sorry for that. Currently I have
>>
>> kernel.hung_task_panic = 1
>> kernel.hung_task_all_cpu_backtrace = 1
>>
>> Please confirm.
>
> You DON'T want hung_task_panic in this case. If the system panics, it
> might do so before the watchdog you mention below fires. We really need
> the sysrq-trigger output and the panic might interfere with that.
>
Fixed. hung_task_panic has been reset to 0. Thank you for your reply.
Regards
Harri
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-10-01 10:21 ` Harald Dunkel
@ 2024-10-02 13:55 ` Harald Dunkel
2024-10-02 14:06 ` Harald Dunkel
0 siblings, 1 reply; 16+ messages in thread
From: Harald Dunkel @ 2024-10-02 13:55 UTC (permalink / raw)
To: NeilBrown
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
[-- Attachment #1: Type: text/plain, Size: 301 bytes --]
Hi folks,
Seems that 1 out of 8 nfsd got stuck again, but it wasn't a permanent
problem. Currently NFS seems to work fine. Nevertheless, I created the
sysrq dump (attached) and kept the host running.
My watchdog did not trigger, ie it was not affected by the NFS problem
yet.
Hope this helps
Harri
[-- Attachment #2: syslog.lxc_server.txt.gz --]
[-- Type: application/gzip, Size: 41577 bytes --]
[-- Attachment #3: syslog.nfs_container.txt.gz --]
[-- Type: application/gzip, Size: 44451 bytes --]
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-10-02 13:55 ` Harald Dunkel
@ 2024-10-02 14:06 ` Harald Dunkel
2024-10-04 12:21 ` Harald Dunkel
2024-10-04 23:37 ` NeilBrown
0 siblings, 2 replies; 16+ messages in thread
From: Harald Dunkel @ 2024-10-02 14:06 UTC (permalink / raw)
To: NeilBrown
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
[-- Attachment #1: Type: text/plain, Size: 30 bytes --]
PS: dmesg -T attached as well.
[-- Attachment #2: dmesg-T.txt.gz --]
[-- Type: application/gzip, Size: 17133 bytes --]
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-10-02 14:06 ` Harald Dunkel
@ 2024-10-04 12:21 ` Harald Dunkel
2024-10-04 23:37 ` NeilBrown
1 sibling, 0 replies; 16+ messages in thread
From: Harald Dunkel @ 2024-10-04 12:21 UTC (permalink / raw)
To: NeilBrown
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
On 2024-10-02 16:06:14, Harald Dunkel wrote:
> PS: dmesg -T attached as well.
PPS: I learned another thing today: I have to disable kernel
logging in the containers. Sorry about that. The nfsd died
again today, but I doubt that the new log files are very useful.
Please mail if you are interested.
The NFS server has been rebooted (without imklog syslog module
in the containers) and the watchdog has been started again.
Regards
Harri
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-10-02 14:06 ` Harald Dunkel
2024-10-04 12:21 ` Harald Dunkel
@ 2024-10-04 23:37 ` NeilBrown
2024-10-07 10:51 ` Harald Dunkel
1 sibling, 1 reply; 16+ messages in thread
From: NeilBrown @ 2024-10-04 23:37 UTC (permalink / raw)
To: Harald Dunkel
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
On Thu, 03 Oct 2024, Harald Dunkel wrote:
> PS: dmesg -T attached as well.
This doesn't contain stacks for all processes - there are too many and
the buffer wrapped.
This is addressed with the log_buf_len kernel parameter.
1M is probably enough, so maybe set log_buf_len=4M to be sure you have
more than enough.
Unfortunately it requires a reboot to change this.
NeilBrown
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-10-04 23:37 ` NeilBrown
@ 2024-10-07 10:51 ` Harald Dunkel
2024-10-07 18:54 ` Harald Dunkel
0 siblings, 1 reply; 16+ messages in thread
From: Harald Dunkel @ 2024-10-07 10:51 UTC (permalink / raw)
To: NeilBrown
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
Hi Neil,
ACK. Debian provided a new kernel version 6.1.112 for Bookworm, anyway.
May I send you a task dump of the running system (before nfsd getting
stuck) after the reboot, just for verification?
Regards
Harri
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-10-07 10:51 ` Harald Dunkel
@ 2024-10-07 18:54 ` Harald Dunkel
0 siblings, 0 replies; 16+ messages in thread
From: Harald Dunkel @ 2024-10-07 18:54 UTC (permalink / raw)
To: NeilBrown
Cc: syzbot, Jeff Layton, Dai.Ngo, chuck.lever, kolga, linux-nfs,
syzkaller-bugs, tom
On 2024-10-07 12:51:08, Harald Dunkel wrote:
> Hi Neil,
>
> ACK. Debian provided a new kernel version 6.1.112 for Bookworm, anyway.
> May I send you a task dump of the running system (before nfsd getting
> stuck) after the reboot, just for verification?
>
Rebooted. The changelog for the new kernel mentions some nfsd fixes:
- nfsd: Simplify code around svc_exit_thread() call in nfsd()
- nfsd: separate nfsd_last_thread() from nfsd_put()
- NFSD: simplify error paths in nfsd_svc()
- nfsd: call nfsd_last_thread() before final nfsd_put()
- nfsd: drop the nfsd_put helper
- nfsd: don't call locks_release_private() twice concurrently
- nfsd: Fix a regression in nfsd_setattr()
Do you think these fixes could be relevant here?
Regards
Harri
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [syzbot] [nfs?] INFO: task hung in nfsd_umount
2024-07-07 4:37 [syzbot] [nfs?] INFO: task hung in nfsd_umount syzbot
2024-07-07 10:49 ` Jeff Layton
@ 2026-08-16 10:14 ` syzbot
1 sibling, 0 replies; 16+ messages in thread
From: syzbot @ 2026-08-16 10:14 UTC (permalink / raw)
To: Dai.Ngo, cel, chuck.lever, dai.ngo, harri, jlayton, kolga,
linux-kernel, linux-nfs, neil, neilb, okorniev, syzkaller-bugs,
tom
syzbot has found a reproducer for the following issue on:
HEAD commit: 3eb40771c00a Merge tag 'soc-fixes-7.2-3' of git://git.kern..
git tree: upstream
console+strace: https://syzkaller.appspot.com/x/log.txt?x=1309ba79580000
kernel config: https://syzkaller.appspot.com/x/.config?x=1d67342c314f228d
dashboard link: https://syzkaller.appspot.com/bug?extid=b568ba42c85a332a88ee
compiler: gcc (Debian 14.2.0-19) 14.2.0, GNU ld (GNU Binutils for Debian) 2.44
syz repro: https://syzkaller.appspot.com/x/repro.syz?x=156daa25580000
C reproducer: https://syzkaller.appspot.com/x/repro.c?x=13ad9a79580000
Downloadable assets:
disk image: https://storage.googleapis.com/syzbot-assets/21569197c3cb/disk-3eb40771.raw.xz
vmlinux: https://storage.googleapis.com/syzbot-assets/09af5299a851/vmlinux-3eb40771.xz
kernel image: https://storage.googleapis.com/syzbot-assets/c657b46695b7/bzImage-3eb40771.xz
IMPORTANT: if you fix the issue, please add the following tag to the commit:
Reported-by: syzbot+b568ba42c85a332a88ee@syzkaller.appspotmail.com
INFO: task syz-executor347:5620 blocked for more than 143 seconds.
Not tainted syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz-executor347 state:D stack:22184 pid:5620 tgid:5620 ppid:5617 task_flags:0x400140 flags:0x00080001
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5510 [inline]
__schedule+0x125c/0x6730 kernel/sched/core.c:7234
__schedule_loop kernel/sched/core.c:7311 [inline]
schedule+0xdd/0x2c0 kernel/sched/core.c:7326
schedule_preempt_disabled+0x13/0x30 kernel/sched/core.c:7383
__mutex_lock_common kernel/locking/mutex.c:726 [inline]
__mutex_lock+0xccc/0x1bd0 kernel/locking/mutex.c:821
nfsd_shutdown_threads+0x5b/0xf0 fs/nfsd/nfssvc.c:576
nfsd_umount+0x3b/0x60 fs/nfsd/nfsctl.c:1365
deactivate_locked_super+0xc1/0x1b0 fs/super.c:477
deactivate_super fs/super.c:510 [inline]
deactivate_super+0xe7/0x110 fs/super.c:506
cleanup_mnt+0x21f/0x450 fs/namespace.c:1317
task_work_run+0x150/0x240 kernel/task_work.c:233
ptrace_notify+0xf3/0x120 kernel/signal.c:2531
ptrace_report_syscall include/linux/ptrace.h:416 [inline]
ptrace_report_syscall_exit include/linux/ptrace.h:478 [inline]
arch_ptrace_report_syscall_exit include/linux/entry-common.h:213 [inline]
syscall_exit_work include/linux/entry-common.h:248 [inline]
syscall_exit_to_user_mode_work include/linux/entry-common.h:279 [inline]
syscall_exit_to_user_mode include/linux/entry-common.h:316 [inline]
do_syscall_64+0x6c0/0x870 arch/x86/entry/syscall_64.c:100
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f056f6eec17
RSP: 002b:00007ffff209c618 EFLAGS: 00000206 ORIG_RAX: 00000000000000a6
RAX: 0000000000000000 RBX: 000000000001399c RCX: 00007f056f6eec17
RDX: 0000000000000000 RSI: 0000000000000009 RDI: 00007ffff209c6d0
RBP: 00007ffff209c6d0 R08: 00007ffff209d6d0 R09: 00000000ffffffff
R10: 0000000000000000 R11: 0000000000000206 R12: 00007ffff209d740
R13: 0000555587301780 R14: 00007ffff209d740 R15: 0000000000000001
</TASK>
INFO: task syz-executor347:5775 blocked for more than 143 seconds.
Not tainted syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz-executor347 state:D stack:27448 pid:5775 tgid:5775 ppid:5618 task_flags:0x400140 flags:0x00080002
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5510 [inline]
__schedule+0x125c/0x6730 kernel/sched/core.c:7234
__schedule_loop kernel/sched/core.c:7311 [inline]
schedule+0xdd/0x2c0 kernel/sched/core.c:7326
schedule_preempt_disabled+0x13/0x30 kernel/sched/core.c:7383
__mutex_lock_common kernel/locking/mutex.c:726 [inline]
__mutex_lock+0xccc/0x1bd0 kernel/locking/mutex.c:821
write_ports+0xa5/0xcc0 fs/nfsd/nfsctl.c:860
nfsctl_transaction_write+0x106/0x1a0 fs/nfsd/nfsctl.c:112
vfs_write+0x2aa/0x1050 fs/read_write.c:685
ksys_write+0x12a/0x250 fs/read_write.c:739
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x115/0x870 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f056f6eea19
RSP: 002b:00007ffff209d708 EFLAGS: 00000246 ORIG_RAX: 0000000000000001
RAX: ffffffffffffffda RBX: 645f6473666e2f2e RCX: 00007f056f6eea19
RDX: 0000000000000005 RSI: 0000200000000700 RDI: 0000000000000003
RBP: 0000000000000000 R08: 00007f056f75a222 R09: 00007f056f75a222
R10: 00007f056f75a222 R11: 0000000000000246 R12: 00007f056f75a47b
R13: 0000000000000001 R14: 00007ffff209d740 R15: 0000000000000000
</TASK>
INFO: task syz-executor347:5776 blocked for more than 143 seconds.
Not tainted syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz-executor347 state:D stack:27088 pid:5776 tgid:5776 ppid:5615 task_flags:0x400140 flags:0x00080002
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5510 [inline]
__schedule+0x125c/0x6730 kernel/sched/core.c:7234
__schedule_loop kernel/sched/core.c:7311 [inline]
schedule+0xdd/0x2c0 kernel/sched/core.c:7326
schedule_preempt_disabled+0x13/0x30 kernel/sched/core.c:7383
__mutex_lock_common kernel/locking/mutex.c:726 [inline]
__mutex_lock+0xccc/0x1bd0 kernel/locking/mutex.c:821
write_ports+0xa5/0xcc0 fs/nfsd/nfsctl.c:860
nfsctl_transaction_write+0x106/0x1a0 fs/nfsd/nfsctl.c:112
vfs_write+0x2aa/0x1050 fs/read_write.c:685
ksys_write+0x12a/0x250 fs/read_write.c:739
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x115/0x870 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f056f6eea19
RSP: 002b:00007ffff209d708 EFLAGS: 00000246 ORIG_RAX: 0000000000000001
RAX: ffffffffffffffda RBX: 645f6473666e2f2e RCX: 00007f056f6eea19
RDX: 0000000000000005 RSI: 0000200000000700 RDI: 0000000000000003
RBP: 0000000000000000 R08: 00007f056f75a222 R09: 00007f056f75a222
R10: 00007f056f75a222 R11: 0000000000000246 R12: 00007f056f75a47b
R13: 0000000000000001 R14: 00007ffff209d740 R15: 0000000000000000
</TASK>
INFO: task syz-executor347:5777 blocked for more than 144 seconds.
Not tainted syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz-executor347 state:D stack:27848 pid:5777 tgid:5777 ppid:5614 task_flags:0x400140 flags:0x00080002
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5510 [inline]
__schedule+0x125c/0x6730 kernel/sched/core.c:7234
__schedule_loop kernel/sched/core.c:7311 [inline]
schedule+0xdd/0x2c0 kernel/sched/core.c:7326
schedule_preempt_disabled+0x13/0x30 kernel/sched/core.c:7383
__mutex_lock_common kernel/locking/mutex.c:726 [inline]
__mutex_lock+0xccc/0x1bd0 kernel/locking/mutex.c:821
write_ports+0xa5/0xcc0 fs/nfsd/nfsctl.c:860
nfsctl_transaction_write+0x106/0x1a0 fs/nfsd/nfsctl.c:112
vfs_write+0x2aa/0x1050 fs/read_write.c:685
ksys_write+0x12a/0x250 fs/read_write.c:739
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x115/0x870 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f056f6eea19
RSP: 002b:00007ffff209d708 EFLAGS: 00000246 ORIG_RAX: 0000000000000001
RAX: ffffffffffffffda RBX: 645f6473666e2f2e RCX: 00007f056f6eea19
RDX: 0000000000000005 RSI: 0000200000000700 RDI: 0000000000000003
RBP: 0000000000000000 R08: 00007f056f75a222 R09: 00007f056f75a222
R10: 00007f056f75a222 R11: 0000000000000246 R12: 00007f056f75a47b
R13: 0000000000000001 R14: 00007ffff209d740 R15: 0000000000000000
</TASK>
Showing all locks held in the system:
1 lock held by khungtaskd/31:
#0: ffffffff8ede8200 (rcu_read_lock){....}-{1:3}, at: rcu_lock_acquire include/linux/rcupdate.h:300 [inline]
#0: ffffffff8ede8200 (rcu_read_lock){....}-{1:3}, at: rcu_read_lock include/linux/rcupdate.h:840 [inline]
#0: ffffffff8ede8200 (rcu_read_lock){....}-{1:3}, at: debug_show_all_locks+0x3d/0x184 kernel/locking/lockdep.c:6775
5 locks held by kworker/u8:6/152:
2 locks held by getty/5348:
#0: ffff888037c7d0a0 (&tty->ldisc_sem){++++}-{0:0}, at: tty_ldisc_ref_wait+0x24/0x80 drivers/tty/tty_ldisc.c:243
#1: ffffc900032332e8 (&ldata->atomic_read_lock){+.+.}-{4:4}, at: n_tty_read+0x419/0x14e0 drivers/tty/n_tty.c:2211
2 locks held by syz-executor347/5619:
#0: ffff88806d1b80d8 (&type->s_umount_key#67){+.+.}-{4:4}, at: __super_lock fs/super.c:58 [inline]
#0: ffff88806d1b80d8 (&type->s_umount_key#67){+.+.}-{4:4}, at: __super_lock_excl fs/super.c:73 [inline]
#0: ffff88806d1b80d8 (&type->s_umount_key#67){+.+.}-{4:4}, at: deactivate_super fs/super.c:509 [inline]
#0: ffff88806d1b80d8 (&type->s_umount_key#67){+.+.}-{4:4}, at: deactivate_super+0xdf/0x110 fs/super.c:506
#1: ffffffff8f272480 (nfsd_mutex){+.+.}-{4:4}, at: nfsd_shutdown_threads+0x5b/0xf0 fs/nfsd/nfssvc.c:576
2 locks held by syz-executor347/5620:
#0: ffff88806d2220d8 (&type->s_umount_key#67){+.+.}-{4:4}, at: __super_lock fs/super.c:58 [inline]
#0: ffff88806d2220d8 (&type->s_umount_key#67){+.+.}-{4:4}, at: __super_lock_excl fs/super.c:73 [inline]
#0: ffff88806d2220d8 (&type->s_umount_key#67){+.+.}-{4:4}, at: deactivate_super fs/super.c:509 [inline]
#0: ffff88806d2220d8 (&type->s_umount_key#67){+.+.}-{4:4}, at: deactivate_super+0xdf/0x110 fs/super.c:506
#1: ffffffff8f272480 (nfsd_mutex){+.+.}-{4:4}, at: nfsd_shutdown_threads+0x5b/0xf0 fs/nfsd/nfssvc.c:576
2 locks held by syz-executor347/5775:
#0: ffff888079e4c450 (sb_writers#9){.+.+}-{0:0}, at: ksys_write+0x12a/0x250 fs/read_write.c:739
#1: ffffffff8f272480 (nfsd_mutex){+.+.}-{4:4}, at: write_ports+0xa5/0xcc0 fs/nfsd/nfsctl.c:860
2 locks held by syz-executor347/5776:
#0: ffff88807af5e450 (sb_writers#9){.+.+}-{0:0}, at: ksys_write+0x12a/0x250 fs/read_write.c:739
#1: ffffffff8f272480 (nfsd_mutex){+.+.}-{4:4}, at: write_ports+0xa5/0xcc0 fs/nfsd/nfsctl.c:860
2 locks held by syz-executor347/5777:
#0: ffff8880246f6450 (sb_writers#9){.+.+}-{0:0}, at: ksys_write+0x12a/0x250 fs/read_write.c:739
#1: ffffffff8f272480 (nfsd_mutex){+.+.}-{4:4}, at: write_ports+0xa5/0xcc0 fs/nfsd/nfsctl.c:860
=============================================
NMI backtrace for cpu 0
CPU: 0 UID: 0 PID: 31 Comm: khungtaskd Not tainted syzkaller #0 PREEMPT(full)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 07/16/2026
Call Trace:
<TASK>
__dump_stack lib/dump_stack.c:94 [inline]
dump_stack_lvl+0x100/0x190 lib/dump_stack.c:120
nmi_cpu_backtrace.cold+0x12d/0x151 lib/nmi_backtrace.c:122
nmi_trigger_cpumask_backtrace+0x21c/0x2a0 lib/nmi_backtrace.c:65
trigger_all_cpu_backtrace include/linux/nmi.h:162 [inline]
__sys_info lib/sys_info.c:157 [inline]
sys_info+0x141/0x190 lib/sys_info.c:165
check_hung_uninterruptible_tasks kernel/hung_task.c:353 [inline]
watchdog+0xcb1/0x1030 kernel/hung_task.c:561
kthread+0x370/0x450 kernel/kthread.c:436
ret_from_fork+0x72b/0xd50 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
Sending NMI from CPU 0 to CPUs 1:
NMI backtrace for cpu 1
CPU: 1 UID: 0 PID: 13 Comm: kworker/u8:1 Not tainted syzkaller #0 PREEMPT(full)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 07/16/2026
Workqueue: bat_events batadv_mcast_mla_update
RIP: 0010:_static_cpu_has arch/x86/include/asm/cpufeature.h:101 [inline]
RIP: 0010:addr_has_metadata mm/kasan/kasan.h:334 [inline]
RIP: 0010:check_region_inline mm/kasan/generic.c:188 [inline]
RIP: 0010:kasan_check_range+0x31/0x1e0 mm/kasan/generic.c:200
Code: 0f 84 7a 01 00 00 48 89 f8 41 54 41 89 d0 48 01 f0 55 53 0f 82 e6 00 00 00 eb 0f cc cc cc 48 b8 00 00 00 00 00 00 00 ff eb 0a <48> b8 00 00 00 00 00 80 ff ff 48 39 c7 0f 82 c2 00 00 00 4c 8d 54
RSP: 0018:ffffc90000127a60 EFLAGS: 00000282
RAX: ffffc90000127b34 RBX: ffffc90000127b30 RCX: ffffffff8b8815dd
RDX: 0000000000000001 RSI: 0000000000000004 RDI: ffffc90000127b30
RBP: 0000000000000004 R08: 0000000000000001 R09: 0000000000000000
R10: 0000000000000001 R11: 0000000000000000 R12: 0000000000000000
R13: 0000000000000000 R14: ffff88807832d8a8 R15: 0000000000000000
FS: 0000000000000000(0000) GS:ffff888123cde000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00005620a0b86660 CR3: 000000000eb94000 CR4: 00000000003526f0
Call Trace:
<TASK>
__asan_memset+0x23/0x50 mm/kasan/shadow.c:84
batadv_mcast_mla_flags_get net/batman-adv/multicast.c:283 [inline]
__batadv_mcast_mla_update net/batman-adv/multicast.c:907 [inline]
batadv_mcast_mla_update+0x10d/0x3200 net/batman-adv/multicast.c:946
process_one_work+0xa23/0x1940 kernel/workqueue.c:3322
process_scheduled_works kernel/workqueue.c:3405 [inline]
worker_thread+0x5ef/0xe50 kernel/workqueue.c:3486
kthread+0x370/0x450 kernel/kthread.c:436
ret_from_fork+0x72b/0xd50 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
---
If you want syzbot to run the reproducer, reply with:
#syz test: git://repo/address.git branch-or-commit-hash
If you attach or paste a git patch, syzbot will apply it before testing.
^ permalink raw reply [flat|nested] 16+ messages in thread
end of thread, other threads:[~2026-08-16 10:14 UTC | newest]
Thread overview: 16+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2024-07-07 4:37 [syzbot] [nfs?] INFO: task hung in nfsd_umount syzbot
2024-07-07 10:49 ` Jeff Layton
2024-07-08 0:07 ` NeilBrown
2024-09-21 7:58 ` Harald Dunkel
2024-09-28 7:41 ` Harald Dunkel
2024-09-28 22:23 ` NeilBrown
2024-09-29 8:23 ` Harald Dunkel
2024-09-29 9:59 ` NeilBrown
2024-10-01 10:21 ` Harald Dunkel
2024-10-02 13:55 ` Harald Dunkel
2024-10-02 14:06 ` Harald Dunkel
2024-10-04 12:21 ` Harald Dunkel
2024-10-04 23:37 ` NeilBrown
2024-10-07 10:51 ` Harald Dunkel
2024-10-07 18:54 ` Harald Dunkel
2026-08-16 10:14 ` syzbot
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.