Netdev List
 help / color / mirror / Atom feed
* [syzbot] [mm?] INFO: rcu detected stall in unmap_region
@ 2026-07-03  7:03 syzbot
  2026-07-22  7:14 ` Nikolay Ivchenko
  0 siblings, 1 reply; 2+ messages in thread
From: syzbot @ 2026-07-03  7:03 UTC (permalink / raw)
  To: akpm, jannh, liam, linux-kernel, linux-mm, ljs, netdev, pfalcato,
	syzkaller-bugs, vbabka

Hello,

syzbot found the following issue on:

HEAD commit:    32f1c2bbb26a net: airoha: dma map xmit frags with skb_frag..
git tree:       net
console output: https://syzkaller.appspot.com/x/log.txt?x=116c2c0a580000
kernel config:  https://syzkaller.appspot.com/x/.config?x=86ba763b42fa66a
dashboard link: https://syzkaller.appspot.com/bug?extid=2ad5ec205a38c46522b3
compiler:       Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
syz repro:      https://syzkaller.appspot.com/x/repro.syz?x=132f5861580000

Downloadable assets:
disk image: https://storage.googleapis.com/syzbot-assets/7b7c3a22a8ed/disk-32f1c2bb.raw.xz
vmlinux: https://storage.googleapis.com/syzbot-assets/168b43c87305/vmlinux-32f1c2bb.xz
kernel image: https://storage.googleapis.com/syzbot-assets/70704720d284/bzImage-32f1c2bb.xz

IMPORTANT: if you fix the issue, please add the following tag to the commit:
Reported-by: syzbot+2ad5ec205a38c46522b3@syzkaller.appspotmail.com

rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
rcu: 	0-...!: (1 GPs behind) idle=6664/1/0x4000000000000000 softirq=17730/17732 fqs=2
rcu: 	(detected by 1, t=10502 jiffies, g=17101, q=1895 ncpus=2)
Sending NMI from CPU 1 to CPUs 0:
NMI backtrace for cpu 0
CPU: 0 UID: 0 PID: 6010 Comm: modprobe Not tainted syzkaller #0 PREEMPT(full) 
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 05/09/2026
RIP: 0010:lock_release+0x5/0x3c0 kernel/locking/lockdep.c:5876
Code: e9 4f fe ff ff 41 bf 2f 00 00 00 e9 06 ff ff ff 0f 1f 44 00 00 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 f3 0f 1e fa 55 <41> 57 41 56 41 55 41 54 53 48 83 ec 30 49 89 f5 49 89 fe 65 48 8b
RSP: 0018:ffffc90000007db0 EFLAGS: 00000082
RAX: 932ea09fa7cced00 RBX: 0000000000000092 RCX: 0000000000000002
RDX: 0000000000000001 RSI: ffffffff81b30e19 RDI: ffffffff9a73d4c8
RBP: ffff88807e29b300 R08: 0000000000000003 R09: 0000000000000004
R10: dffffc0000000000 R11: fffff52000000f98 R12: ffff8880b86281c0
R13: dffffc0000000000 R14: ffffffff9a73d4b0 R15: ffff8880b8628350
FS:  0000000000000000(0000) GS:ffff888125226000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007ffae81e84e8 CR3: 000000007ee52000 CR4: 00000000003526f0
Call Trace:
 <IRQ>
 __raw_spin_unlock_irqrestore include/linux/spinlock_api_smp.h:176 [inline]
 _raw_spin_unlock_irqrestore+0x1b/0x80 kernel/locking/spinlock.c:198
 debug_hrtimer_deactivate kernel/time/hrtimer.c:490 [inline]
 __run_hrtimer kernel/time/hrtimer.c:2000 [inline]
 __hrtimer_run_queues+0x239/0xa10 kernel/time/hrtimer.c:2096
 hrtimer_interrupt+0x448/0x910 kernel/time/hrtimer.c:2215
 local_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1051 [inline]
 __sysvec_apic_timer_interrupt+0x102/0x430 arch/x86/kernel/apic/apic.c:1068
 instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1062 [inline]
 sysvec_apic_timer_interrupt+0xa1/0xc0 arch/x86/kernel/apic/apic.c:1062
 </IRQ>
 <TASK>
 asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:674
RIP: 0010:do_zap_pte_range mm/memory.c:1819 [inline]
RIP: 0010:zap_pte_range mm/memory.c:1934 [inline]
RIP: 0010:zap_pmd_range mm/memory.c:2020 [inline]
RIP: 0010:zap_pud_range mm/memory.c:2048 [inline]
RIP: 0010:zap_p4d_range mm/memory.c:2069 [inline]
RIP: 0010:__zap_vma_range+0xfbd/0x4f10 mm/memory.c:2109
Code: 4c 89 64 24 30 4c 89 e6 48 83 e6 9f 31 ff e8 3a 4a ad ff 4d 89 e7 49 83 e4 9f 75 2f 48 8b 44 24 10 4c 01 e8 48 83 f8 01 74 28 <e8> 3e 45 ad ff 48 83 c3 08 49 ff c5 4d 89 f4 eb a3 4d 89 ef e8 2a
RSP: 0018:ffffc90003287000 EFLAGS: 00000286
RAX: ffffffffffffff73 RBX: ffff888034e9d190 RCX: ffff88802dcf1f00
RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000000
RBP: ffffc900032872f0 R08: ffff888078979803 R09: 1ffff1100f12f300
R10: dffffc0000000000 R11: ffffed100f12f301 R12: 0000000000000000
R13: 0000000000000032 R14: dffffc0000000000 R15: 0000000000000000
 unmap_vmas+0x390/0x550 mm/memory.c:2178
 unmap_region+0x208/0x330 mm/vma.c:488
 vms_clear_ptes mm/vma.c:1303 [inline]
 vms_clean_up_area mm/vma.c:1315 [inline]
 __mmap_setup mm/vma.c:2476 [inline]
 __mmap_region mm/vma.c:2756 [inline]
 mmap_region+0xc62/0x2310 mm/vma.c:2860
 do_mmap+0xc3b/0x10c0 mm/mmap.c:560
 vm_mmap_pgoff+0x272/0x4e0 mm/util.c:581
 ksys_mmap_pgoff+0x4dc/0x760 mm/mmap.c:606
 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
 do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
 entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7ffae8219242
Code: 08 00 04 00 00 eb e2 90 41 f7 c1 ff 0f 00 00 75 27 55 89 cd 53 48 89 fb 48 85 ff 74 33 41 89 ea 48 89 df b8 09 00 00 00 0f 05 <48> 3d 00 f0 ff ff 77 5e 5b 5d c3 0f 1f 00 c7 05 46 40 01 00 16 00
RSP: 002b:00007ffec64cd428 EFLAGS: 00000206 ORIG_RAX: 0000000000000009
RAX: ffffffffffffffda RBX: 00007ffae7f73000 RCX: 00007ffae8219242
RDX: 0000000000000005 RSI: 000000000014e000 RDI: 00007ffae7f73000
RBP: 0000000000000812 R08: 0000000000000000 R09: 0000000000028000
R10: 0000000000000812 R11: 0000000000000206 R12: 00007ffec64cd478
R13: 00007ffae81ed5f0 R14: 00007ffec64cdc60 R15: 00000fffd8c99a88
 </TASK>
rcu: rcu_preempt kthread starved for 10477 jiffies! g17101 f0x0 RCU_GP_WAIT_FQS(5) ->state=0x0 ->cpu=1
rcu: 	Unless rcu_preempt kthread gets sufficient CPU time, OOM is now expected behavior.
rcu: RCU grace-period kthread stack dump:
task:rcu_preempt     state:R  running task     stack:28192 pid:16    tgid:16    ppid:2      task_flags:0x208040 flags:0x00080000
Call Trace:
 <TASK>
 context_switch kernel/sched/core.c:5510 [inline]
 __schedule+0x17d9/0x56c0 kernel/sched/core.c:7234
 __schedule_loop kernel/sched/core.c:7311 [inline]
 schedule+0x164/0x2b0 kernel/sched/core.c:7326
 schedule_timeout+0x152/0x2c0 kernel/time/sleep_timeout.c:99
 rcu_gp_fqs_loop+0x30c/0x11f0 kernel/rcu/tree.c:2123
 rcu_gp_kthread+0x9e/0x2b0 kernel/rcu/tree.c:2325
 kthread+0x388/0x470 kernel/kthread.c:436
 ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
 </TASK>
rcu: Stack dump where RCU GP kthread last ran:
CPU: 1 UID: 0 PID: 4992 Comm: udevd Not tainted syzkaller #0 PREEMPT(full) 
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 05/09/2026
RIP: 0010:csd_lock_wait kernel/smp.c:342 [inline]
RIP: 0010:smp_call_function_many_cond+0x10b0/0x14b0 kernel/smp.c:892
Code: c0 75 73 41 8b 1e 89 de 83 e6 01 31 ff e8 d8 1e 0c 00 83 e3 01 48 bb 00 00 00 00 00 fc ff df 75 07 e8 84 1a 0c 00 eb 37 f3 90 <41> 0f b6 04 1c 84 c0 75 10 41 f7 06 01 00 00 00 74 1e e8 69 1a 0c
RSP: 0018:ffffc90003837580 EFLAGS: 00000293
RAX: ffffffff81ba5717 RBX: dffffc0000000000 RCX: ffff88807e5a0000
RDX: 0000000000000000 RSI: 0000000000000001 RDI: 0000000000000000
RBP: ffffc900038376a8 R08: ffffffff9032d2f7 R09: 1ffffffff2065a5e
R10: dffffc0000000000 R11: fffffbfff2065a5f R12: 1ffff110170c85fd
R13: ffff8880b873c448 R14: ffff8880b8642fe8 R15: 0000000000000000
FS:  00007fdd3255d880(0000) GS:ffff888125326000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00002c919703e000 CR3: 000000007e7f2000 CR4: 00000000003526f0
Call Trace:
 <TASK>
 on_each_cpu_cond_mask+0x3f/0x80 kernel/smp.c:1057
 __flush_tlb_multi arch/x86/include/asm/paravirt.h:46 [inline]
 flush_tlb_multi arch/x86/mm/tlb.c:1361 [inline]
 flush_tlb_mm_range+0x5c4/0x1090 arch/x86/mm/tlb.c:1451
 dup_mmap+0x1786/0x1d90 mm/mmap.c:1905
 dup_mm kernel/fork.c:1538 [inline]
 copy_mm+0x11a/0x480 kernel/fork.c:1590
 copy_process+0x1e99/0x43f0 kernel/fork.c:2288
 kernel_clone+0x2d7/0x940 kernel/fork.c:2747
 __do_sys_clone kernel/fork.c:2888 [inline]
 __se_sys_clone kernel/fork.c:2872 [inline]
 __x64_sys_clone+0x1b6/0x230 kernel/fork.c:2872
 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
 do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
 entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7fdd31ef1636
Code: 89 df e8 6d e8 f6 ff 45 31 c0 31 d2 31 f6 64 48 8b 04 25 10 00 00 00 bf 11 00 20 01 4c 8d 90 d0 02 00 00 b8 38 00 00 00 0f 05 <48> 3d 00 f0 ff ff 77 52 89 c5 85 c0 75 31 64 48 8b 04 25 10 00 00
RSP: 002b:00007ffcd95f6670 EFLAGS: 00000246 ORIG_RAX: 0000000000000038
RAX: ffffffffffffffda RBX: 00007ffcd95f6678 RCX: 00007fdd31ef1636
RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000001200011
RBP: 0000564712dd5910 R08: 0000000000000000 R09: 0000564712ddd8e0
R10: 00007fdd3255db50 R11: 0000000000000246 R12: 00007ffcd95f6a30
R13: 0000000000000000 R14: 0000000000000000 R15: 0000000000000000
 </TASK>


---
This report is generated by a bot. It may contain errors.
See https://goo.gl/tpsmEJ for more information about syzbot.
syzbot engineers can be reached at syzkaller@googlegroups.com.

syzbot will keep track of this issue. See:
https://goo.gl/tpsmEJ#status for how to communicate with syzbot.

If the report is already addressed, let syzbot know by replying with:
#syz fix: exact-commit-title

If you want syzbot to run the reproducer, reply with:
#syz test: git://repo/address.git branch-or-commit-hash
If you attach or paste a git patch, syzbot will apply it before testing.

If you want to overwrite report's subsystems, reply with:
#syz set subsystems: new-subsystem
(See the list of subsystem names on the web dashboard)

If the report is a duplicate of another one, reply with:
#syz dup: exact-subject-of-another-report

If you want to undo deduplication, reply with:
#syz undup

^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [syzbot] [mm?] INFO: rcu detected stall in unmap_region
  2026-07-03  7:03 [syzbot] [mm?] INFO: rcu detected stall in unmap_region syzbot
@ 2026-07-22  7:14 ` Nikolay Ivchenko
  0 siblings, 0 replies; 2+ messages in thread
From: Nikolay Ivchenko @ 2026-07-22  7:14 UTC (permalink / raw)
  To: syzbot+2ad5ec205a38c46522b3
  Cc: akpm, jannh, liam, linux-kernel, linux-mm, ljs, netdev, pfalcato,
	syzkaller-bugs, vbabka, vinicius.gomes, jhs, jiri, davem,
	edumazet, kuba, pabeni, horms

Hi,

On syzbot report, syzbot wrote:
> Hello,
>
> syzbot found the following issue on:
>
> HEAD commit:    32f1c2bbb26a net: airoha: dma map xmit frags with skb_frag..
> git tree:       net
> console output: https://syzkaller.appspot.com/x/log.txt?x=116c2c0a580000
> kernel config:  https://syzkaller.appspot.com/x/.config?x=86ba763b42fa66a
> dashboard link: https://syzkaller.appspot.com/bug?extid=2ad5ec205a38c46522b3
> compiler:       Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
> syz repro:      https://syzkaller.appspot.com/x/repro.syz?x=132f5861580000
>
> Downloadable assets:
> disk image: https://storage.googleapis.com/syzbot-assets/7b7c3a22a8ed/disk-32f1c2bb.raw.xz
> vmlinux: https://storage.googleapis.com/syzbot-assets/168b43c87305/vmlinux-32f1c2bb.xz
> kernel image: https://storage.googleapis.com/syzbot-assets/70704720d284/bzImage-32f1c2bb.xz
>
> IMPORTANT: if you fix the issue, please add the following tag to the commit:
> Reported-by: syzbot+2ad5ec205a38c46522b3@syzkaller.appspotmail.com
>
> rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
> rcu: 	0-...!: (1 GPs behind) idle=6664/1/0x4000000000000000 softirq=17730/17732 fqs=2
> rcu: 	(detected by 1, t=10502 jiffies, g=17101, q=1895 ncpus=2)
> Sending NMI from CPU 1 to CPUs 0:
> NMI backtrace for cpu 0
> CPU: 0 UID: 0 PID: 6010 Comm: modprobe Not tainted syzkaller #0 PREEMPT(full)
> ...
> Call Trace:
> <IRQ>
> __raw_spin_unlock_irqrestore include/linux/spinlock_api_smp.h:176 [inline]
> _raw_spin_unlock_irqrestore+0x1b/0x80 kernel/locking/spinlock.c:198
> debug_hrtimer_deactivate kernel/time/hrtimer.c:490 [inline]
> __run_hrtimer kernel/time/hrtimer.c:2000 [inline]
> __hrtimer_run_queues+0x239/0xa10 kernel/time/hrtimer.c:2096
> hrtimer_interrupt+0x448/0x910 kernel/time/hrtimer.c:2215
> local_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1051 [inline]
> __sysvec_apic_timer_interrupt+0x102/0x430 arch/x86/kernel/apic/apic.c:1068
> instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1062 [inline]
> sysvec_apic_timer_interrupt+0xa1/0xc0 arch/x86/kernel/apic/apic.c:1062
> </IRQ>

I have analyzed this issue in sch_taprio, and configuring extremely short
schedule entry intervals (e.g., 700 ns) in software mode
(!FULL_OFFLOAD_IS_ENABLED) seems to be the root cause, leading to an hrtimer
interrupt storm that locks up the CPU and results in RCU stalls and softlockups.
When testing with larger intervals, the lockup completely disappeared.
Furthermore, no matter how many times I captured this stall, NMI backtraces
consistently showed the CPU trapped inside the timer handler
(advance_sched / hrtimer_interrupt).

Please note that this bug can show up in different execution contexts and with
various crash titles depending on what the CPU was doing when the interrupt
storm hit. As such, despite what the subject line of this report suggests, this
is not a memory management issue — the root cause is entirely in sch_taprio
(networking).

=== Cause Analysis ===

Currently, fill_sched_entry() validates schedule intervals against a minimum
duration using length_to_duration(q, ETH_ZLEN):

    int min_duration = length_to_duration(q, ETH_ZLEN);

    [...]

    if (interval < min_duration) {
        NL_SET_ERR_MSG(extack, "Invalid interval for schedule entry");
        return -EINVAL;
    }

On high-speed interfaces (e.g., veth, which defaults to 10 Gbps),
transmitting 60 bytes (ETH_ZLEN) takes only ~48 ns. Consequently, an interval
such as 700 ns passes validation because 700 ns > 48 ns.

However, in software scheduling mode (!FULL_OFFLOAD_IS_ENABLED), the hrtimer
handling overhead easily exceeds 700 ns, particularly in virtualized
environments or on slower CPUs, leading to CPU lockups and RCU stalls.

=== Minimal Reproducer ===

Based on the reproducer provided by syzbot, I have created a minimal shell
script reproducer:

#!/bin/bash
ip link del dev veth0 2>/dev/null
ip link add dev veth0 numtxqueues 4 type veth peer name veth1
ip link set dev veth0 up

# 700 ns interval passes validation on 10Gbps veth, causing softlockup:
tc qdisc add dev veth0 parent root handle 1: taprio \
    num_tc 2 \
    map 0 1 \
    queues 1@0 1@1 \
    sched-entry S 01 700 \
    clockid CLOCK_TAI

Note that after running this script, you may need to wait about 20-30 seconds
before the RCU stall or softlockup warning appears in dmesg.

=== Discussion ===

I would like to ask for opinions on how this problem should be addressed.
One approach is to enforce a minimum software interval threshold at
configuration time in fill_sched_entry() when
!FULL_OFFLOAD_IS_ENABLED(q->flags):

    if (!FULL_OFFLOAD_IS_ENABLED(q->flags) && interval < NSEC_PER_USEC) {
        NL_SET_ERR_MSG_MOD(extack, "Interval too small for software mode");
        return -EINVAL;
    }

However, I am doubtful whether this is the correct way to fix the issue.
Hardcoding a fixed lower bound (such as 1 us or NSEC_PER_USEC) is a heuristic.
An interval that works safely on high-performance hardware might still cause
softlockups on slower hardware or inside heavily loaded virtual machines, while
a conservative threshold might unnecessarily reject valid configurations.

Where should this problem ideally be solved? If it belongs at the input
validation level, how can we properly validate input data when we cannot know
in advance whether a given CPU will handle the processing load?
Conversely, if we should try to detect this issue at runtime, how exactly
should that be implemented?

Best regards,
Nikolay Ivchenko

#syz set subsystems: net

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-07-22  7:14 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-03  7:03 [syzbot] [mm?] INFO: rcu detected stall in unmap_region syzbot
2026-07-22  7:14 ` Nikolay Ivchenko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox