* [syzbot] [kernel?] general protection fault in timers_dead_cpu @ 2026-08-14 7:28 syzbot 2026-08-17 21:14 ` Thomas Gleixner 0 siblings, 1 reply; 9+ messages in thread From: syzbot @ 2026-08-14 7:28 UTC (permalink / raw) To: linux-kernel, peterz, syzkaller-bugs, tglx Hello, syzbot found the following issue on: HEAD commit: db2ddb871435 Linux 7.2-rc7 git tree: upstream console output: https://syzkaller.appspot.com/x/log.txt?x=14ec2149580000 kernel config: https://syzkaller.appspot.com/x/.config?x=c44651ea7dd2f307 dashboard link: https://syzkaller.appspot.com/bug?extid=74de56995244fe32ffe2 compiler: gcc (Debian 14.2.0-19) 14.2.0, GNU ld (GNU Binutils for Debian) 2.44 syz repro: https://syzkaller.appspot.com/x/repro.syz?x=12ec2149580000 C reproducer: https://syzkaller.appspot.com/x/repro.c?x=16d962c6580000 Downloadable assets: disk image (non-bootable): https://storage.googleapis.com/syzbot-assets/d900f083ada3/non_bootable_disk-db2ddb87.raw.xz vmlinux: https://storage.googleapis.com/syzbot-assets/698def9fcf7a/vmlinux-db2ddb87.xz kernel image: https://storage.googleapis.com/syzbot-assets/fd8b6091a563/bzImage-db2ddb87.xz IMPORTANT: if you fix the issue, please add the following tag to the commit: Reported-by: syzbot+74de56995244fe32ffe2@syzkaller.appspotmail.com smpboot: CPU 1 is now offline Oops: general protection fault, probably for non-canonical address 0xdffffc0000000000: 0000 [#1] SMP KASAN NOPTI KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007] CPU: 2 UID: 0 PID: 6237 Comm: syz.2.92 Not tainted syzkaller #0 PREEMPT(full) Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014 RIP: 0010:__hlist_del include/linux/list.h:1029 [inline] RIP: 0010:detach_timer kernel/time/timer.c:891 [inline] RIP: 0010:migrate_timer_list kernel/time/timer.c:2493 [inline] RIP: 0010:timers_dead_cpu+0x326/0x860 kernel/time/timer.c:2541 Code: 0f 85 0c 04 00 00 48 8d 7b 08 48 8b 2b 48 89 f8 48 c1 e8 03 42 80 3c 38 00 0f 85 df 03 00 00 48 8b 43 08 48 89 c2 48 c1 ea 03 <42> 80 3c 3a 00 0f 85 b4 03 00 00 48 89 28 48 89 04 24 48 85 ed 74 RSP: 0018:ffffc900040078c0 EFLAGS: 00010046 RAX: 0000000000000000 RBX: ffff88806a523440 RCX: ffffffff81f5d5d8 RDX: 0000000000000000 RSI: ffffffff81f5d420 RDI: ffff88806a523448 RBP: dead000000000122 R08: 0000000000000001 R09: 0000000000000000 R10: 0000000000000001 R11: 0000000000000000 R12: ffff88806a5258d0 R13: ffffed100d4a4b1a R14: ffff88806a6251c0 R15: dffffc0000000000 FS: 00007f01d4dfe6c0(0000) GS:ffff8880d5fec000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 00007f9c1ca70000 CR3: 000000003de42000 CR4: 0000000000352ef0 Call Trace: <TASK> cpuhp_invoke_callback+0x3b4/0x9a0 kernel/cpu.c:194 __cpuhp_invoke_callback_range+0x158/0x1d0 kernel/cpu.c:972 cpuhp_invoke_callback_range kernel/cpu.c:996 [inline] cpuhp_down_callbacks kernel/cpu.c:1386 [inline] _cpu_down+0x523/0x1020 kernel/cpu.c:1457 cpu_down_maps_locked kernel/cpu.c:1483 [inline] cpu_down kernel/cpu.c:1491 [inline] cpu_device_down+0x82/0xc0 kernel/cpu.c:1508 device_offline drivers/base/core.c:4279 [inline] device_offline+0x2f4/0x400 drivers/base/core.c:4263 online_store+0xd1/0x180 drivers/base/core.c:2879 dev_attr_store+0x58/0x80 drivers/base/core.c:2505 sysfs_kf_write+0xf2/0x150 fs/sysfs/file.c:145 kernfs_fop_write_iter+0x3e0/0x5f0 fs/kernfs/file.c:345 new_sync_write fs/read_write.c:595 [inline] vfs_write+0x6ac/0x1050 fs/read_write.c:687 ksys_write+0x12a/0x250 fs/read_write.c:739 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x115/0x870 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f RIP: 0033:0x7f01d579e0d9 Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 e8 ff ff ff f7 d8 64 89 01 48 RSP: 002b:00007f01d4dfe028 EFLAGS: 00000246 ORIG_RAX: 0000000000000001 RAX: ffffffffffffffda RBX: 00007f01d5a25fa0 RCX: 00007f01d579e0d9 RDX: 0000000000000002 RSI: 0000200000000180 RDI: 0000000000000003 RBP: 00007f01d5835024 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000 R13: 00007f01d5a26038 R14: 00007f01d5a25fa0 R15: 00007ffe4c5cd718 </TASK> Modules linked in: ---[ end trace 0000000000000000 ]--- RIP: 0010:__hlist_del include/linux/list.h:1029 [inline] RIP: 0010:detach_timer kernel/time/timer.c:891 [inline] RIP: 0010:migrate_timer_list kernel/time/timer.c:2493 [inline] RIP: 0010:timers_dead_cpu+0x326/0x860 kernel/time/timer.c:2541 Code: 0f 85 0c 04 00 00 48 8d 7b 08 48 8b 2b 48 89 f8 48 c1 e8 03 42 80 3c 38 00 0f 85 df 03 00 00 48 8b 43 08 48 89 c2 48 c1 ea 03 <42> 80 3c 3a 00 0f 85 b4 03 00 00 48 89 28 48 89 04 24 48 85 ed 74 RSP: 0018:ffffc900040078c0 EFLAGS: 00010046 RAX: 0000000000000000 RBX: ffff88806a523440 RCX: ffffffff81f5d5d8 RDX: 0000000000000000 RSI: ffffffff81f5d420 RDI: ffff88806a523448 RBP: dead000000000122 R08: 0000000000000001 R09: 0000000000000000 R10: 0000000000000001 R11: 0000000000000000 R12: ffff88806a5258d0 R13: ffffed100d4a4b1a R14: ffff88806a6251c0 R15: dffffc0000000000 FS: 00007f01d4dfe6c0(0000) GS:ffff8880d5fec000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 00007f9c1ca70000 CR3: 000000003de42000 CR4: 0000000000352ef0 ---------------- Code disassembly (best guess): 0: 0f 85 0c 04 00 00 jne 0x412 6: 48 8d 7b 08 lea 0x8(%rbx),%rdi a: 48 8b 2b mov (%rbx),%rbp d: 48 89 f8 mov %rdi,%rax 10: 48 c1 e8 03 shr $0x3,%rax 14: 42 80 3c 38 00 cmpb $0x0,(%rax,%r15,1) 19: 0f 85 df 03 00 00 jne 0x3fe 1f: 48 8b 43 08 mov 0x8(%rbx),%rax 23: 48 89 c2 mov %rax,%rdx 26: 48 c1 ea 03 shr $0x3,%rdx * 2a: 42 80 3c 3a 00 cmpb $0x0,(%rdx,%r15,1) <-- trapping instruction 2f: 0f 85 b4 03 00 00 jne 0x3e9 35: 48 89 28 mov %rbp,(%rax) 38: 48 89 04 24 mov %rax,(%rsp) 3c: 48 85 ed test %rbp,%rbp 3f: 74 .byte 0x74 --- This report is generated by a bot. It may contain errors. See https://goo.gl/tpsmEJ for more information about syzbot. syzbot engineers can be reached at syzkaller@googlegroups.com. syzbot will keep track of this issue. See: https://goo.gl/tpsmEJ#status for how to communicate with syzbot. If the report is already addressed, let syzbot know by replying with: #syz fix: exact-commit-title If you want syzbot to run the reproducer, reply with: #syz test: git://repo/address.git branch-or-commit-hash If you attach or paste a git patch, syzbot will apply it before testing. If you want to overwrite report's subsystems, reply with: #syz set subsystems: new-subsystem (See the list of subsystem names on the web dashboard) If the report is a duplicate of another one, reply with: #syz dup: exact-subject-of-another-report If you want to undo deduplication, reply with: #syz undup ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [syzbot] [kernel?] general protection fault in timers_dead_cpu 2026-08-14 7:28 [syzbot] [kernel?] general protection fault in timers_dead_cpu syzbot @ 2026-08-17 21:14 ` Thomas Gleixner 2026-08-17 22:12 ` Thomas Gleixner ` (2 more replies) 0 siblings, 3 replies; 9+ messages in thread From: Thomas Gleixner @ 2026-08-17 21:14 UTC (permalink / raw) To: syzbot, linux-kernel, peterz, syzkaller-bugs Cc: Borislav Petkov, Tony Luck, linux-edac On Fri, Aug 14 2026 at 00:28, syzbot wrote: > HEAD commit: db2ddb871435 Linux 7.2-rc7 > git tree: upstream > console output: https://syzkaller.appspot.com/x/log.txt?x=14ec2149580000 > kernel config: https://syzkaller.appspot.com/x/.config?x=c44651ea7dd2f307 This has NUMA_EMU=y. Either turn it off or add '-smp 2,sockets=2' to the qemu command line. Otherwise the topology code is unhappy. > dashboard link: https://syzkaller.appspot.com/bug?extid=74de56995244fe32ffe2 > compiler: gcc (Debian 14.2.0-19) 14.2.0, GNU ld (GNU Binutils for Debian) 2.44 > syz repro: https://syzkaller.appspot.com/x/repro.syz?x=12ec2149580000 > C reproducer: https://syzkaller.appspot.com/x/repro.c?x=16d962c6580000 > > Downloadable assets: > disk image (non-bootable): https://storage.googleapis.com/syzbot-assets/d900f083ada3/non_bootable_disk-db2ddb87.raw.xz > vmlinux: https://storage.googleapis.com/syzbot-assets/698def9fcf7a/vmlinux-db2ddb87.xz > kernel image: https://storage.googleapis.com/syzbot-assets/fd8b6091a563/bzImage-db2ddb87.xz > > IMPORTANT: if you fix the issue, please add the following tag to the commit: > Reported-by: syzbot+74de56995244fe32ffe2@syzkaller.appspotmail.com > > smpboot: CPU 1 is now offline > Oops: general protection fault, probably for non-canonical address 0xdffffc0000000000: 0000 [#1] SMP KASAN NOPTI > KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007] > CPU: 2 UID: 0 PID: 6237 Comm: syz.2.92 Not tainted syzkaller #0 PREEMPT(full) > Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014 > RIP: 0010:__hlist_del include/linux/list.h:1029 [inline] > RIP: 0010:detach_timer kernel/time/timer.c:891 [inline] > RIP: 0010:migrate_timer_list kernel/time/timer.c:2493 [inline] > RIP: 0010:timers_dead_cpu+0x326/0x860 kernel/time/timer.c:2541 This happens because the reproducer does two things in parallel: 1) Hotplug CPU1 2) Toggle /sys/devices/system/machinecheck/machinecheck0/ignore_ce #2 is not serialized against CPU hotplug so it can end up interfering with the hotplug operation: CPU0 CPU1 hotplug kick_ap() wait_for_ap() hotplug mce_cpu_pre_down() mce_disable_cpu(); timer_delete_sync(); // CPU is still marked online set_ignore_ce() on_each_cpu(mce_enable_ce, (void *)1, 1); IPI timer_start() .... hotplug timers_dead_cpu() migrate timer to CPU0 // Migrates the MCE timer of CPU1, which is a bug in itself ... hotplug bringup_ap() ... identify_secondary_cpu() mcheck_cpu_init() __mcheck_cpu_setup_timer() timer_setup() <- FAIL That re-initializes the active timer, which is now queued on CPU0. What puzzled me was that debugobjects did not catch that issue. It turned out that during some rework the debug_activate() invocation for the timer migration case got lost. So debugobjects carries the wrong state. That's easy to fix: --- a/kernel/time/timer.c +++ b/kernel/time/timer.c @@ -2492,6 +2492,7 @@ static void migrate_timer_list(struct timer_base *new_base, struct hlist_head *h timer = hlist_entry(head->first, struct timer_list, entry); detach_timer(timer, false); timer->flags = (timer->flags & ~TIMER_BASEMASK) | cpu; + debug_timer_activate(timer); internal_add_timer(new_base, timer); } } With that it catches the culprit as expected: ODEBUG: init active (active state 0) object: ffff88827be234a0 object type: timer_list hint: mce_timer_fn+0x0/0x280 WARNING: lib/debugobjects.c:632 at debug_print_object+0xec/0x230, CPU#1: swapper/1/0 CPU: 1 UID: 0 PID: 0 Comm: swapper/1 Not tainted 7.2.0-dirty #274 PREEMPT(full) RIP: 0010:debug_print_object+0x18a/0x230 Call Trace: __debug_object_init+0x230/0x3c0 timer_init_key+0x5c/0x2e0 mcheck_cpu_init+0x3e1/0x600 identify_cpu+0x1e03/0x3660 identify_secondary_cpu+0xaa/0x160 ap_starting+0xa1/0x150 start_secondary+0x66/0x110 common_startup_64+0x13e/0x157 The knee jerk "fix" is to serialize against CPU hotplug in set_ignore_ce() and the other sysfs write functions which can result in exactly the same problem. It's not only the timer. CMCI suffers from the same issue that it can be reenabled via sysfs between the "x86/mce:online" state and going completely offline. Haven't looked further, but that seems to be a general design problem in that code. But guarding against hotplug alone solves it only partially because with partial hotplug the same issue happens when: 1) a partial hotplug goes below the "x86/mce:online" state which disarms the timer, but stops before the CPU is marked offline 2) set_ignore_ce() or one of the other sysfs write functions reenables it 3) a subsequent hotplug operation brings the CPU completely down. The below quick hack, which I'm not proud of, cures it. I let the MCE wizards think about the underlying design problem and let them come up with a hopefully nicer solution. Thanks, tglx --- --- a/arch/x86/kernel/cpu/mce/core.c +++ b/arch/x86/kernel/cpu/mce/core.c @@ -68,6 +68,8 @@ static DEFINE_MUTEX(mce_sysfs_mutex); #define SPINUNIT 100 /* 100ns */ +static struct cpumask mce_active_cpus; + DEFINE_PER_CPU_READ_MOSTLY(unsigned int, mce_num_banks); DEFINE_PER_CPU_READ_MOSTLY(struct mce_bank[MAX_NR_BANKS], mce_banks_array); @@ -2459,6 +2461,8 @@ static void mce_cpu_restart(void *data) { if (!mce_available(raw_cpu_ptr(&cpu_info))) return; + if (!cpumask_test_cpu(smp_processor_id(), &mce_active_cpus)) + return; __mcheck_cpu_init_generic(); __mcheck_cpu_init_prepare_banks(); __mcheck_cpu_init_timer(); @@ -2478,6 +2482,8 @@ static void mce_disable_cmci(void *data) { if (!mce_available(raw_cpu_ptr(&cpu_info))) return; + if (!cpumask_test_cpu(smp_processor_id(), &mce_active_cpus)) + return; cmci_clear(); } @@ -2485,6 +2491,8 @@ static void mce_enable_ce(void *all) { if (!mce_available(raw_cpu_ptr(&cpu_info))) return; + if (!cpumask_test_cpu(smp_processor_id(), &mce_active_cpus)) + return; cmci_reenable(); cmci_recheck(); if (all) @@ -2540,6 +2548,7 @@ static ssize_t set_bank(struct device *s b->ctl = new; mutex_lock(&mce_sysfs_mutex); + guard(cpus_read_lock)(); mce_restart(); mutex_unlock(&mce_sysfs_mutex); @@ -2557,6 +2566,7 @@ static ssize_t set_ignore_ce(struct devi mutex_lock(&mce_sysfs_mutex); if (mca_cfg.ignore_ce ^ !!new) { + guard(cpus_read_lock)(); if (new) { /* disable ce features */ mce_timer_delete_all(); @@ -2584,6 +2594,7 @@ static ssize_t set_cmci_disabled(struct mutex_lock(&mce_sysfs_mutex); if (mca_cfg.cmci_disabled ^ !!new) { + guard(cpus_read_lock)(); if (new) { /* disable cmci */ on_each_cpu(mce_disable_cmci, NULL, 1); @@ -2610,6 +2621,7 @@ static ssize_t store_int_with_restart(st return ret; mutex_lock(&mce_sysfs_mutex); + guard(cpus_read_lock)(); mce_restart(); mutex_unlock(&mce_sysfs_mutex); @@ -2730,6 +2742,8 @@ static void mce_disable_cpu(void) if (!mce_available(raw_cpu_ptr(&cpu_info))) return; + cpumask_clear_cpu(smp_processor_id(), &mce_active_cpus); + if (!cpuhp_tasks_frozen) cmci_clear(); @@ -2752,6 +2766,8 @@ static void mce_reenable_cpu(void) if (b->init) wrmsrq(mca_msr_reg(i, MCA_CTL), b->ctl); } + + cpumask_set_cpu(smp_processor_id(), &mce_active_cpus); } static int mce_cpu_dead(unsigned int cpu) ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [syzbot] [kernel?] general protection fault in timers_dead_cpu 2026-08-17 21:14 ` Thomas Gleixner @ 2026-08-17 22:12 ` Thomas Gleixner 2026-08-18 0:21 ` Borislav Petkov 2026-08-17 22:14 ` [PATCH] timer: Keep debugobjects state consistent in migrate_timer_list() Thomas Gleixner 2026-08-17 23:55 ` [syzbot] [kernel?] general protection fault in timers_dead_cpu Borislav Petkov 2 siblings, 1 reply; 9+ messages in thread From: Thomas Gleixner @ 2026-08-17 22:12 UTC (permalink / raw) To: syzbot, linux-kernel, peterz, syzkaller-bugs Cc: Borislav Petkov, Tony Luck, linux-edac On Mon, Aug 17 2026 at 23:14, Thomas Gleixner wrote: > The below quick hack, which I'm not proud of, cures it. I let the MCE > wizards think about the underlying design problem and let them come up > with a hopefully nicer solution. After talking to Borislav briefly, I came up with less ugly one. Thanks, tglx --- --- a/arch/x86/kernel/cpu/mce/core.c +++ b/arch/x86/kernel/cpu/mce/core.c @@ -1734,8 +1734,13 @@ int memory_failure(unsigned long pfn, in */ static unsigned long check_interval = INITIAL_CHECK_INTERVAL; -static DEFINE_PER_CPU(unsigned long, mce_next_interval); /* in jiffies */ -static DEFINE_PER_CPU(struct timer_list, mce_timer); +struct mce_poll_state { + struct timer_list timer; + unsigned long next_interval; + bool active; +}; + +static DEFINE_PER_CPU(struct mce_poll_state, mce_poll_state); static void __start_timer(struct timer_list *t, unsigned long interval) { @@ -1764,12 +1769,12 @@ static bool should_enable_timer(unsigned static void mce_timer_fn(struct timer_list *t) { - struct timer_list *cpu_t = this_cpu_ptr(&mce_timer); + struct mce_poll_state *pst = this_cpu_ptr(&mce_poll_state); unsigned long iv; - WARN_ON(cpu_t != t); + WARN_ON(&pst->timer != t); - iv = __this_cpu_read(mce_next_interval); + iv = pst->next_interval; if (mce_available(this_cpu_ptr(&cpu_info))) mc_poll_banks(); @@ -1786,7 +1791,7 @@ static void mce_timer_fn(struct timer_li if (mce_get_storm_mode()) { __start_timer(t, HZ); } else if (should_enable_timer(iv)) { - __this_cpu_write(mce_next_interval, iv); + pst->next_interval = iv; __start_timer(t, iv); } } @@ -1798,14 +1803,14 @@ static void mce_timer_fn(struct timer_li */ void mce_timer_kick(bool storm) { - struct timer_list *t = this_cpu_ptr(&mce_timer); + struct mce_poll_state *pst = this_cpu_ptr(&mce_poll_state); mce_set_storm_mode(storm); if (storm) - __start_timer(t, HZ); + __start_timer(&pst->timer, HZ); else - __this_cpu_write(mce_next_interval, check_interval * HZ); + pst->next_interval = check_interval * HZ; } /* Must not be called in IRQ context where timer_delete_sync() can deadlock */ @@ -1814,7 +1819,7 @@ static void mce_timer_delete_all(void) int cpu; for_each_online_cpu(cpu) - timer_delete_sync(&per_cpu(mce_timer, cpu)); + timer_delete_sync(&per_cpu(mce_poll_state.timer, cpu)); } static void __mcheck_cpu_mce_banks_init(void) @@ -2070,29 +2075,29 @@ static void __mcheck_cpu_clear_vendor(st } } -static void mce_start_timer(struct timer_list *t) +static void mce_start_timer(struct mce_poll_state *pst) { unsigned long iv = check_interval * HZ; if (should_enable_timer(iv)) { - this_cpu_write(mce_next_interval, iv); - __start_timer(t, iv); + pst->next_interval = iv; + __start_timer(&pst->timer, iv); } } static void __mcheck_cpu_setup_timer(void) { - struct timer_list *t = this_cpu_ptr(&mce_timer); + struct mce_poll_state *pst = this_cpu_ptr(&mce_poll_state); - timer_setup(t, mce_timer_fn, TIMER_PINNED); + timer_setup(&pst->timer, mce_timer_fn, TIMER_PINNED); } static void __mcheck_cpu_init_timer(void) { - struct timer_list *t = this_cpu_ptr(&mce_timer); + struct mce_poll_state *pst = this_cpu_ptr(&mce_poll_state); - timer_setup(t, mce_timer_fn, TIMER_PINNED); - mce_start_timer(t); + timer_setup(&pst->timer, mce_timer_fn, TIMER_PINNED); + mce_start_timer(pst); } bool filter_mce(struct mce *m) @@ -2459,6 +2464,8 @@ static void mce_cpu_restart(void *data) { if (!mce_available(raw_cpu_ptr(&cpu_info))) return; + if (!this_cpu_read(mce_poll_state.active)) + return; __mcheck_cpu_init_generic(); __mcheck_cpu_init_prepare_banks(); __mcheck_cpu_init_timer(); @@ -2478,6 +2485,8 @@ static void mce_disable_cmci(void *data) { if (!mce_available(raw_cpu_ptr(&cpu_info))) return; + if (!this_cpu_read(mce_poll_state.active)) + return; cmci_clear(); } @@ -2485,6 +2494,8 @@ static void mce_enable_ce(void *all) { if (!mce_available(raw_cpu_ptr(&cpu_info))) return; + if (!this_cpu_read(mce_poll_state.active)) + return; cmci_reenable(); cmci_recheck(); if (all) @@ -2540,6 +2551,7 @@ static ssize_t set_bank(struct device *s b->ctl = new; mutex_lock(&mce_sysfs_mutex); + guard(cpus_read_lock)(); mce_restart(); mutex_unlock(&mce_sysfs_mutex); @@ -2557,6 +2569,7 @@ static ssize_t set_ignore_ce(struct devi mutex_lock(&mce_sysfs_mutex); if (mca_cfg.ignore_ce ^ !!new) { + guard(cpus_read_lock)(); if (new) { /* disable ce features */ mce_timer_delete_all(); @@ -2584,6 +2597,7 @@ static ssize_t set_cmci_disabled(struct mutex_lock(&mce_sysfs_mutex); if (mca_cfg.cmci_disabled ^ !!new) { + guard(cpus_read_lock)(); if (new) { /* disable cmci */ on_each_cpu(mce_disable_cmci, NULL, 1); @@ -2610,6 +2624,7 @@ static ssize_t store_int_with_restart(st return ret; mutex_lock(&mce_sysfs_mutex); + guard(cpus_read_lock)(); mce_restart(); mutex_unlock(&mce_sysfs_mutex); @@ -2764,21 +2779,23 @@ static int mce_cpu_dead(unsigned int cpu static int mce_cpu_online(unsigned int cpu) { - struct timer_list *t = this_cpu_ptr(&mce_timer); + struct mce_poll_state *pst = this_cpu_ptr(&mce_poll_state); mce_device_create(cpu); mce_threshold_create_device(cpu); mce_reenable_cpu(); - mce_start_timer(t); + mce_start_timer(pst); + pst->active = true; return 0; } static int mce_cpu_pre_down(unsigned int cpu) { - struct timer_list *t = this_cpu_ptr(&mce_timer); + struct mce_poll_state *pst = this_cpu_ptr(&mce_poll_state); + pst->active = false; mce_disable_cpu(); - timer_delete_sync(t); + timer_delete_sync(&pst->timer); mce_threshold_remove_device(cpu); mce_device_remove(cpu); return 0; ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [syzbot] [kernel?] general protection fault in timers_dead_cpu 2026-08-17 22:12 ` Thomas Gleixner @ 2026-08-18 0:21 ` Borislav Petkov 0 siblings, 0 replies; 9+ messages in thread From: Borislav Petkov @ 2026-08-18 0:21 UTC (permalink / raw) To: Thomas Gleixner Cc: syzbot, linux-kernel, peterz, syzkaller-bugs, Tony Luck, linux-edac On Tue, Aug 18, 2026 at 12:12:45AM +0200, Thomas Gleixner wrote: > On Mon, Aug 17 2026 at 23:14, Thomas Gleixner wrote: > > The below quick hack, which I'm not proud of, cures it. I let the MCE > > wizards think about the underlying design problem and let them come up > > with a hopefully nicer solution. > > After talking to Borislav briefly, I came up with less ugly one. Yap, LGTM. Lemme try to unify the whole MCE percpu gunk as we spoke. -- Regards/Gruss, Boris. https://people.kernel.org/tglx/notes-about-netiquette ^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH] timer: Keep debugobjects state consistent in migrate_timer_list() 2026-08-17 21:14 ` Thomas Gleixner 2026-08-17 22:12 ` Thomas Gleixner @ 2026-08-17 22:14 ` Thomas Gleixner 2026-08-18 8:56 ` [tip: core/urgent] " tip-bot2 for Thomas Gleixner 2026-08-17 23:55 ` [syzbot] [kernel?] general protection fault in timers_dead_cpu Borislav Petkov 2 siblings, 1 reply; 9+ messages in thread From: Thomas Gleixner @ 2026-08-17 22:14 UTC (permalink / raw) To: syzbot, linux-kernel, peterz, syzkaller-bugs Cc: Borislav Petkov, Tony Luck, linux-edac, Frederic Weisbecker When timers are migrated away from an offline CPU the debugobjects state gets corrupted. The timer is accounted as inactive on deletion, but the enqueue on the alive CPU lacks the activation call. That used to work, but got broken when the trace point and the debug objects call got separated. That change missed to fixup migrate_timer_list(). Add the missing debug_timer_activate() invocation to fix it. Fixes: dc1e7dc5ac62 ("timer: Move trace point to get proper index") Signed-off-by: Thomas Gleixner <tglx@kernel.org> Cc: stable@vger.kernel.org --- kernel/time/timer.c | 1 + 1 file changed, 1 insertion(+) --- a/kernel/time/timer.c +++ b/kernel/time/timer.c @@ -2492,6 +2492,7 @@ static void migrate_timer_list(struct ti timer = hlist_entry(head->first, struct timer_list, entry); detach_timer(timer, false); timer->flags = (timer->flags & ~TIMER_BASEMASK) | cpu; + debug_timer_activate(timer); internal_add_timer(new_base, timer); } } ^ permalink raw reply [flat|nested] 9+ messages in thread
* [tip: core/urgent] timer: Keep debugobjects state consistent in migrate_timer_list() 2026-08-17 22:14 ` [PATCH] timer: Keep debugobjects state consistent in migrate_timer_list() Thomas Gleixner @ 2026-08-18 8:56 ` tip-bot2 for Thomas Gleixner 0 siblings, 0 replies; 9+ messages in thread From: tip-bot2 for Thomas Gleixner @ 2026-08-18 8:56 UTC (permalink / raw) To: linux-tip-commits; +Cc: Thomas Gleixner, stable, x86, linux-kernel The following commit has been merged into the core/urgent branch of tip: Commit-ID: c793bbfc4a0a9f5a66978fc91559e9681748dbeb Gitweb: https://git.kernel.org/tip/c793bbfc4a0a9f5a66978fc91559e9681748dbeb Author: Thomas Gleixner <tglx@kernel.org> AuthorDate: Tue, 18 Aug 2026 00:14:57 +02:00 Committer: Thomas Gleixner <tglx@kernel.org> CommitterDate: Tue, 18 Aug 2026 10:51:43 +02:00 timer: Keep debugobjects state consistent in migrate_timer_list() When timers are migrated away from an offline CPU the debugobjects state gets corrupted. The timer is accounted as inactive on deletion, but the enqueue on the alive CPU lacks the activation call. That used to work, but got broken when the trace point and the debug objects call got separated. That change missed to fixup migrate_timer_list(). Add the missing debug_timer_activate() invocation to fix it. Fixes: dc1e7dc5ac62 ("timer: Move trace point to get proper index") Signed-off-by: Thomas Gleixner <tglx@kernel.org> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/87bjb0l7ha.ffs@fw13 --- kernel/time/timer.c | 1 + 1 file changed, 1 insertion(+) diff --git a/kernel/time/timer.c b/kernel/time/timer.c index 655a8c6..ae9abf1 100644 --- a/kernel/time/timer.c +++ b/kernel/time/timer.c @@ -2492,6 +2492,7 @@ static void migrate_timer_list(struct timer_base *new_base, struct hlist_head *h timer = hlist_entry(head->first, struct timer_list, entry); detach_timer(timer, false); timer->flags = (timer->flags & ~TIMER_BASEMASK) | cpu; + debug_timer_activate(timer); internal_add_timer(new_base, timer); } } ^ permalink raw reply related [flat|nested] 9+ messages in thread
* Re: [syzbot] [kernel?] general protection fault in timers_dead_cpu 2026-08-17 21:14 ` Thomas Gleixner 2026-08-17 22:12 ` Thomas Gleixner 2026-08-17 22:14 ` [PATCH] timer: Keep debugobjects state consistent in migrate_timer_list() Thomas Gleixner @ 2026-08-17 23:55 ` Borislav Petkov 2026-08-18 0:18 ` Borislav Petkov 2 siblings, 1 reply; 9+ messages in thread From: Borislav Petkov @ 2026-08-17 23:55 UTC (permalink / raw) To: Thomas Gleixner Cc: syzbot, linux-kernel, peterz, syzkaller-bugs, Tony Luck, linux-edac On Mon, Aug 17, 2026 at 11:14:35PM +0200, Thomas Gleixner wrote: > timers_dead_cpu() > migrate timer to CPU0 > > // Migrates the MCE timer of CPU1, which is a bug in itself Stupid question: can we prevent this? As in, this timer is not migratable, do not migrate it. But then what do you do with a timer which is not migratable and its CPU goes offline? Perhaps cancel it... It won't matter in the MCE case, that's for sure. Anyway, just some musings from reading this... -- Regards/Gruss, Boris. https://people.kernel.org/tglx/notes-about-netiquette ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [syzbot] [kernel?] general protection fault in timers_dead_cpu 2026-08-17 23:55 ` [syzbot] [kernel?] general protection fault in timers_dead_cpu Borislav Petkov @ 2026-08-18 0:18 ` Borislav Petkov 2026-08-18 9:09 ` Thomas Gleixner 0 siblings, 1 reply; 9+ messages in thread From: Borislav Petkov @ 2026-08-18 0:18 UTC (permalink / raw) To: Thomas Gleixner Cc: syzbot, linux-kernel, peterz, syzkaller-bugs, Tony Luck, linux-edac On Mon, Aug 17, 2026 at 04:55:53PM -0700, Borislav Petkov wrote: > On Mon, Aug 17, 2026 at 11:14:35PM +0200, Thomas Gleixner wrote: > > timers_dead_cpu() > > migrate timer to CPU0 > > > > // Migrates the MCE timer of CPU1, which is a bug in itself > > Stupid question: can we prevent this? > > As in, this timer is not migratable, do not migrate it. > > But then what do you do with a timer which is not migratable and its CPU goes > offline? > > Perhaps cancel it... > > It won't matter in the MCE case, that's for sure. > > Anyway, just some musings from reading this... Hmm, the down path does timer_delete_sync() so I guess I'm missing an aspect here about the timer migration. -- Regards/Gruss, Boris. https://people.kernel.org/tglx/notes-about-netiquette ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [syzbot] [kernel?] general protection fault in timers_dead_cpu 2026-08-18 0:18 ` Borislav Petkov @ 2026-08-18 9:09 ` Thomas Gleixner 0 siblings, 0 replies; 9+ messages in thread From: Thomas Gleixner @ 2026-08-18 9:09 UTC (permalink / raw) To: Borislav Petkov Cc: syzbot, linux-kernel, peterz, syzkaller-bugs, Tony Luck, linux-edac On Mon, Aug 17 2026 at 17:18, Borislav Petkov wrote: > On Mon, Aug 17, 2026 at 04:55:53PM -0700, Borislav Petkov wrote: >> On Mon, Aug 17, 2026 at 11:14:35PM +0200, Thomas Gleixner wrote: >> > timers_dead_cpu() >> > migrate timer to CPU0 >> > >> > // Migrates the MCE timer of CPU1, which is a bug in itself >> >> Stupid question: can we prevent this? >> >> As in, this timer is not migratable, do not migrate it. We could do that, but that's just papering over the underlying issues. >> But then what do you do with a timer which is not migratable and its CPU goes >> offline? >> >> Perhaps cancel it... Yes, but that's not really well defined. >> It won't matter in the MCE case, that's for sure. Correct. You still have CMCI ... >> Anyway, just some musings from reading this... > > Hmm, the down path does timer_delete_sync() so I guess I'm missing an aspect > here about the timer migration. Care to read my first reply where I described exactly how that happens? The timer is rearmed by that sysfs muck _after_ the down callback deleted it. And the same happens to CMCI. The down callback stops it and the sysfs muck reenables it. Alternatively we can split the hotplug callbacks and have one in the late stage of hotplug after the point of no return, which stops the timer and CMCI. Then let the existing one only care about the device stuff which requires task context. Something like the untested below. Thanks, tglx --- --- a/arch/x86/kernel/cpu/mce/core.c +++ b/arch/x86/kernel/cpu/mce/core.c @@ -2762,23 +2762,33 @@ static int mce_cpu_dead(unsigned int cpu return 0; } -static int mce_cpu_online(unsigned int cpu) +static int mce_cpu_starting(unsigned int cpu) { struct timer_list *t = this_cpu_ptr(&mce_timer); - mce_device_create(cpu); - mce_threshold_create_device(cpu); mce_reenable_cpu(); mce_start_timer(t); return 0; } -static int mce_cpu_pre_down(unsigned int cpu) +static int mce_cpu_dying(unsigned int cpu) { struct timer_list *t = this_cpu_ptr(&mce_timer); mce_disable_cpu(); timer_delete_sync(t); + return 0; +} + +static int mce_cpu_online(unsigned int cpu) +{ + mce_device_create(cpu); + mce_threshold_create_device(cpu); + return 0; +} + +static int mce_cpu_pre_down(unsigned int cpu) +{ mce_threshold_remove_device(cpu); mce_device_remove(cpu); return 0; @@ -2841,6 +2851,14 @@ static __init int mcheck_init_device(voi mce_cpu_dead); if (err) goto err_out_mem; + /* + * Invokes mce_cpu_starting() on all CPUs which are online when + * the state is installed. + */ + err = cpuhp_setup_state(CPUHP_AP_X86_MCE_STARTING, "x86/mce:starting", + mce_cpu_starting, mce_cpu_dying); + if (err < 0) + goto err_out_starting; /* * Invokes mce_cpu_online() on all CPUs which are online when @@ -2856,6 +2874,9 @@ static __init int mcheck_init_device(voi return 0; err_out_online: + cpuhp_remove_state(CPUHP_AP_X86_MCE_STARTING); + +err_out_starting: cpuhp_remove_state(CPUHP_X86_MCE_DEAD); err_out_mem: --- a/include/linux/cpuhotplug.h +++ b/include/linux/cpuhotplug.h @@ -186,6 +186,7 @@ enum cpuhp_state { CPUHP_AP_HRTIMERS_DYING, CPUHP_AP_TICK_DYING, CPUHP_AP_X86_TBOOT_DYING, + CPUHP_AP_X86_MCE_STARTING, CPUHP_AP_ARM_CACHE_B15_RAC_DYING, CPUHP_AP_ONLINE, CPUHP_TEARDOWN_CPU, ^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-08-18 9:09 UTC | newest] Thread overview: 9+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-08-14 7:28 [syzbot] [kernel?] general protection fault in timers_dead_cpu syzbot 2026-08-17 21:14 ` Thomas Gleixner 2026-08-17 22:12 ` Thomas Gleixner 2026-08-18 0:21 ` Borislav Petkov 2026-08-17 22:14 ` [PATCH] timer: Keep debugobjects state consistent in migrate_timer_list() Thomas Gleixner 2026-08-18 8:56 ` [tip: core/urgent] " tip-bot2 for Thomas Gleixner 2026-08-17 23:55 ` [syzbot] [kernel?] general protection fault in timers_dead_cpu Borislav Petkov 2026-08-18 0:18 ` Borislav Petkov 2026-08-18 9:09 ` Thomas Gleixner
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.