The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3)
@ 2026-08-05  0:01 syzbot
  2026-08-05 19:29 ` Andrew Morton
  0 siblings, 1 reply; 7+ messages in thread
From: syzbot @ 2026-08-05  0:01 UTC (permalink / raw)
  To: akpm, hannes, jackmanb, linux-kernel, linux-mm, mhocko, surenb,
	syzkaller-bugs, vbabka, ziy

Hello,

syzbot found the following issue on:

HEAD commit:    3708dd948844 Merge tag 'pm-7.2-rc6' of git://git.kernel.or..
git tree:       upstream
console output: https://syzkaller.appspot.com/x/log.txt?x=11ac703e580000
kernel config:  https://syzkaller.appspot.com/x/.config?x=4e38b15c29e6a1d9
dashboard link: https://syzkaller.appspot.com/bug?extid=d2401aeb74cc84adba04
compiler:       Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8

Unfortunately, I don't have any reproducer for this issue yet.

Downloadable assets:
disk image: https://storage.googleapis.com/syzbot-assets/c0390423374e/disk-3708dd94.raw.xz
vmlinux: https://storage.googleapis.com/syzbot-assets/3130bf9c5dbf/vmlinux-3708dd94.xz
kernel image: https://storage.googleapis.com/syzbot-assets/2c8fbf8aeba4/bzImage-3708dd94.xz

IMPORTANT: if you fix the issue, please add the following tag to the commit:
Reported-by: syzbot+d2401aeb74cc84adba04@syzkaller.appspotmail.com

rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
rcu: 	Tasks blocked on level-0 rcu_node (CPUs 0-1): P37/1:b..l
rcu: 	(detected by 1, t=10503 jiffies, g=32265, q=1190 ncpus=2)
task:khugepaged      state:R  running task     stack:26864 pid:37    tgid:37    ppid:2      task_flags:0x200040 flags:0x00080000
Call Trace:
 <TASK>
 context_switch kernel/sched/core.c:5510 [inline]
 __schedule+0x17d9/0x56c0 kernel/sched/core.c:7234
 preempt_schedule_irq+0x4d/0xa0 kernel/sched/core.c:7556
 irqentry_exit_to_kernel_mode include/linux/irq-entry-common.h:539 [inline]
 irqentry_exit+0x14f/0x8f0 kernel/entry/common.c:167
 asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:674
RIP: 0010:lock_release+0x2d7/0x3c0 kernel/locking/lockdep.c:5893
Code: 7d ca 11 00 00 00 00 eb b5 e8 55 ad 2e 0a f7 c3 00 02 00 00 74 b9 65 48 8b 05 45 38 ca 11 48 3b 44 24 28 75 44 fb 48 83 c4 30 <5b> 41 5c 41 5d 41 5e 41 5f 5d c3 cc cc cc cc cc 48 8d 3d a2 33 b8
RSP: 0018:ffffc90000ad7028 EFLAGS: 00000286
RAX: 623a32d217845600 RBX: 0000000000000206 RCX: 0000000000000046
RDX: 0000000000000000 RSI: ffffffff8e4b3351 RDI: ffffffff8c4bdd80
RBP: ffff888020ea0ba0 R08: ffffc90000ad74d0 R09: 0000000000000000
R10: ffffc90000ad7158 R11: fffff5200015ae2d R12: 0000000000000000
R13: 0000000000000000 R14: ffffffff8eb59c60 R15: ffff888020ea0000
 rcu_lock_release include/linux/rcupdate.h:310 [inline]
 rcu_read_unlock include/linux/rcupdate.h:871 [inline]
 class_rcu_destructor include/linux/rcupdate.h:1183 [inline]
 unwind_next_frame+0x1baa/0x2550 arch/x86/kernel/unwind_orc.c:709
 arch_stack_walk+0x11b/0x150 arch/x86/kernel/stacktrace.c:25
 stack_trace_save+0xa9/0x100 kernel/stacktrace.c:122
 save_stack+0x122/0x230 mm/page_owner.c:165
 __reset_page_owner+0x71/0x1f0 mm/page_owner.c:320
 reset_page_owner include/linux/page_owner.h:25 [inline]
 __free_pages_prepare mm/page_alloc.c:1406 [inline]
 __free_frozen_pages+0xc1e/0xd10 mm/page_alloc.c:2950
 __folio_put+0x4b3/0x590 mm/swap.c:112
 folio_put_refs include/linux/mm.h:2144 [inline]
 collapse_file mm/khugepaged.c:2643 [inline]
 collapse_scan_file+0x4285/0x5210 mm/khugepaged.c:2773
 collapse_single_pmd+0x2b1/0x3da0 mm/khugepaged.c:2808
 collapse_scan_mm_slot mm/khugepaged.c:2913 [inline]
 khugepaged_do_scan mm/khugepaged.c:2993 [inline]
 khugepaged+0xa00/0x1780 mm/khugepaged.c:3048
 kthread+0x388/0x470 kernel/kthread.c:436
 ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
 </TASK>


---
This report is generated by a bot. It may contain errors.
See https://goo.gl/tpsmEJ for more information about syzbot.
syzbot engineers can be reached at syzkaller@googlegroups.com.

syzbot will keep track of this issue. See:
https://goo.gl/tpsmEJ#status for how to communicate with syzbot.

If the report is already addressed, let syzbot know by replying with:
#syz fix: exact-commit-title

If you want to overwrite report's subsystems, reply with:
#syz set subsystems: new-subsystem
(See the list of subsystem names on the web dashboard)

If the report is a duplicate of another one, reply with:
#syz dup: exact-subject-of-another-report

If you want to undo deduplication, reply with:
#syz undup

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3)
  2026-08-05  0:01 [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3) syzbot
@ 2026-08-05 19:29 ` Andrew Morton
  2026-08-05 20:28   ` Paul E. McKenney
  0 siblings, 1 reply; 7+ messages in thread
From: Andrew Morton @ 2026-08-05 19:29 UTC (permalink / raw)
  To: syzbot
  Cc: hannes, jackmanb, linux-kernel, linux-mm, mhocko, surenb,
	syzkaller-bugs, vbabka, ziy, Paul E. McKenney

On Tue, 04 Aug 2026 17:01:48 -0700 syzbot <syzbot+d2401aeb74cc84adba04@syzkaller.appspotmail.com> wrote:

> Hello,
> 
> syzbot found the following issue on:
> 
> HEAD commit:    3708dd948844 Merge tag 'pm-7.2-rc6' of git://git.kernel.or..
> git tree:       upstream
> console output: https://syzkaller.appspot.com/x/log.txt?x=11ac703e580000
> kernel config:  https://syzkaller.appspot.com/x/.config?x=4e38b15c29e6a1d9
> dashboard link: https://syzkaller.appspot.com/bug?extid=d2401aeb74cc84adba04
> compiler:       Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
> 
> Unfortunately, I don't have any reproducer for this issue yet.

Thanks.

Lazy optimists (ahem) paste this gunk into Gemini and ask "what the
heck just happened".  The results are often useful, but should be
treated with skepticism.  In this case I think it came usably close.

	https://share.gemini.google/vq4TLhTiLBih


tl;dr: khugepaged's collapse_scan_file() is taking too long and RCU got
starved.  I don't think khugepaged is doing anything wrong here,
per-se.  There's a lot of work to do and we're doing it.

An appropriate fix would be to take a break, let RCU do its thing then
get back to work.  But I don't think RCU offers interfaces for that?

collapse_scan_file()'s main loop has

		if (need_resched()) {
			xas_pause(&xas);
			cond_resched_rcu();
		}

but that won't help with the RCU stall detector(?).

I suggest that a suitable fix here would be to add the analogous

	if (rcu_i_need_to_take_a_break()) {
		rcu_read_unlock();
		rcu_take_a_break())	
		rcu_read_lock();
	}

(iirc rcu_read_unlock() does an rcu run, so rcu_take_a_break() isn't
needed here)

Paul, wdyt?


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3)
  2026-08-05 19:29 ` Andrew Morton
@ 2026-08-05 20:28   ` Paul E. McKenney
  2026-08-06  6:03     ` Andrew Morton
  2026-08-06  8:18     ` Vlastimil Babka (SUSE)
  0 siblings, 2 replies; 7+ messages in thread
From: Paul E. McKenney @ 2026-08-05 20:28 UTC (permalink / raw)
  To: Andrew Morton
  Cc: syzbot, hannes, jackmanb, linux-kernel, linux-mm, mhocko, surenb,
	syzkaller-bugs, vbabka, ziy

On Wed, Aug 05, 2026 at 12:29:52PM -0700, Andrew Morton wrote:
> On Tue, 04 Aug 2026 17:01:48 -0700 syzbot <syzbot+d2401aeb74cc84adba04@syzkaller.appspotmail.com> wrote:
> 
> > Hello,
> > 
> > syzbot found the following issue on:
> > 
> > HEAD commit:    3708dd948844 Merge tag 'pm-7.2-rc6' of git://git.kernel.or..
> > git tree:       upstream
> > console output: https://syzkaller.appspot.com/x/log.txt?x=11ac703e580000
> > kernel config:  https://syzkaller.appspot.com/x/.config?x=4e38b15c29e6a1d9
> > dashboard link: https://syzkaller.appspot.com/bug?extid=d2401aeb74cc84adba04
> > compiler:       Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
> > 
> > Unfortunately, I don't have any reproducer for this issue yet.
> 
> Thanks.
> 
> Lazy optimists (ahem) paste this gunk into Gemini and ask "what the
> heck just happened".  The results are often useful, but should be
> treated with skepticism.  In this case I think it came usably close.
> 
> 	https://share.gemini.google/vq4TLhTiLBih
> 
> 
> tl;dr: khugepaged's collapse_scan_file() is taking too long and RCU got
> starved.  I don't think khugepaged is doing anything wrong here,
> per-se.  There's a lot of work to do and we're doing it.
> 
> An appropriate fix would be to take a break, let RCU do its thing then
> get back to work.  But I don't think RCU offers interfaces for that?
> 
> collapse_scan_file()'s main loop has
> 
> 		if (need_resched()) {
> 			xas_pause(&xas);
> 			cond_resched_rcu();
> 		}
> 
> but that won't help with the RCU stall detector(?).
> 
> I suggest that a suitable fix here would be to add the analogous
> 
> 	if (rcu_i_need_to_take_a_break()) {
> 		rcu_read_unlock();
> 		rcu_take_a_break())	
> 		rcu_read_lock();
> 	}
> 
> (iirc rcu_read_unlock() does an rcu run, so rcu_take_a_break() isn't
> needed here)
> 
> Paul, wdyt?

Let's see...

The console log says "rcu_preempt detected stalls on CPUs/tasks",
which means that cond_resched() is a no-op, but it also means that
the rcu_read_unlock() in cond_resched_rcu() will directly take care of
informing RCU of the pause.

But that is clearly not happening.  Why?

Well, we have this:

rcu: 	Tasks blocked on level-0 rcu_node (CPUs 0-1): P37/1:b..l

This means that the task whose RCU read-side critical section is blocking
the current RCU grace period isn't even running, and thus cannot invoke
cond_resched_rcu(), let alone the rcu_read_unlock() within that function.
So an RCU CPU stall warning is expected behavior.  Or at least it is not
in any way ruled out.

What we need is RCU priority boosting.  Except that the .config file
does not enable this.  Not only is there no CONFIG_RCU_BOOST=y, there
is also no CONFIG_RCU_EXPERT=y and no CONFIG_PREEMPT_RT=y.  But there
is CONFIG_RT_MUTEX=y and CONFIG_RCU_EXPERT=y.

Because we don't have RCU priority boosting, if the load on the system
is heavy enough to prevent our poor preempted RCU reader (PID 37) from
running, the grace period cannot end.

I am not sure why this task is saving its stack, but maybe that is normal
for this code path?

My bemusement aside, I recommend running this test either with
non-preemptible RCU (CONFIG_PREEMPT_LAZY=y these days) or enabling RCU
priority boosting (CONFIG_RCU_EXPERT=y and CONFIG_RCU_BOOST=y).

Maybe RCU_BOOST should no longer depend on RCU_EXPERT?  I would of
course need ot remove the prompt ("Enable RCU priority boosting") to
avoid annoying Linus.  Maybe as shown below.

Thoughts?

							Thanx, Paul

------------------------------------------------------------------------

diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig
index 1a5fb3156c062a..5141ad8d1cd029 100644
--- a/kernel/rcu/Kconfig
+++ b/kernel/rcu/Kconfig
@@ -237,17 +237,16 @@ config RCU_FANOUT_LEAF
 	  Take the default if unsure.
 
 config RCU_BOOST
-	bool "Enable RCU priority boosting"
-	depends on (RT_MUTEXES && PREEMPT_RCU && RCU_EXPERT) || PREEMPT_RT
+	bool
+	depends on (RT_MUTEXES && PREEMPT_RCU) || PREEMPT_RT
 	default y if PREEMPT_RT
 	help
 	  This option boosts the priority of preempted RCU readers that
 	  block the current preemptible RCU grace period for too long.
 	  This option also prevents heavy loads from blocking RCU
-	  callback invocation.
+	  callback invocation.  It is now automatically enabled in
+	  any kernel that can benefit from it and that can support it.
 
-	  Say Y here if you are working with real-time apps or heavy loads
-	  Say N here if you are unsure.
 
 config RCU_BOOST_DELAY
 	int "Milliseconds to delay boosting after RCU grace-period start"

^ permalink raw reply related	[flat|nested] 7+ messages in thread

* Re: [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3)
  2026-08-05 20:28   ` Paul E. McKenney
@ 2026-08-06  6:03     ` Andrew Morton
  2026-08-06 16:27       ` Paul E. McKenney
  2026-08-06  8:18     ` Vlastimil Babka (SUSE)
  1 sibling, 1 reply; 7+ messages in thread
From: Andrew Morton @ 2026-08-06  6:03 UTC (permalink / raw)
  To: paulmck
  Cc: syzbot, hannes, jackmanb, linux-kernel, linux-mm, mhocko, surenb,
	syzkaller-bugs, vbabka, ziy

On Wed, 5 Aug 2026 13:28:16 -0700 "Paul E. McKenney" <paulmck@kernel.org> wrote:

> > collapse_scan_file()'s main loop has
> > 
> > 		if (need_resched()) {
> > 			xas_pause(&xas);
> > 			cond_resched_rcu();
> > 		}
> > 
> > but that won't help with the RCU stall detector(?).
> > 
> > I suggest that a suitable fix here would be to add the analogous
> > 
> > 	if (rcu_i_need_to_take_a_break()) {
> > 		rcu_read_unlock();
> > 		rcu_take_a_break())	
> > 		rcu_read_lock();
> > 	}
> > 
> > (iirc rcu_read_unlock() does an rcu run, so rcu_take_a_break() isn't
> > needed here)
> > 
> > Paul, wdyt?
> 
> Let's see...
> 
> The console log says "rcu_preempt detected stalls on CPUs/tasks",
> which means that cond_resched() is a no-op, but it also means that
> the rcu_read_unlock() in cond_resched_rcu() will directly take care of
> informing RCU of the pause.
> 
> But that is clearly not happening.  Why?
> 
> Well, we have this:
> 
> rcu: 	Tasks blocked on level-0 rcu_node (CPUs 0-1): P37/1:b..l
> 
> This means that the task whose RCU read-side critical section is blocking
> the current RCU grace period isn't even running, and thus cannot invoke
> cond_resched_rcu(), let alone the rcu_read_unlock() within that function.
> So an RCU CPU stall warning is expected behavior.  Or at least it is not
> in any way ruled out.
> 
> What we need is RCU priority boosting.

Do we?  I'm suggesting we need need_resched_rcu()!

> Except that the .config file
> does not enable this.  Not only is there no CONFIG_RCU_BOOST=y, there
> is also no CONFIG_RCU_EXPERT=y and no CONFIG_PREEMPT_RT=y.  But there
> is CONFIG_RT_MUTEX=y and CONFIG_RCU_EXPERT=y.
> 
> Because we don't have RCU priority boosting, if the load on the system
> is heavy enough to prevent our poor preempted RCU reader (PID 37) from
> running, the grace period cannot end.
> 
> I am not sure why this task is saving its stack, but maybe that is normal
> for this code path?
> 
> My bemusement aside, I recommend running this test either with
> non-preemptible RCU (CONFIG_PREEMPT_LAZY=y these days) or enabling RCU
> priority boosting (CONFIG_RCU_EXPERT=y and CONFIG_RCU_BOOST=y).
> 
> Maybe RCU_BOOST should no longer depend on RCU_EXPERT?  I would of
> course need ot remove the prompt ("Enable RCU priority boosting") to
> avoid annoying Linus.  Maybe as shown below.
> 
> Thoughts?

If I'm understanding correctly, this workload is busted with this
config and the proposed fix is to alter the config?  Well, why are we
permitting that config at all?

Seems to me that a solution to permit this config to work is very
simple.  Something like:

	time_t start;

	rcu_read_lock();
	start = current_time();

	for (lots of work) {
		...
		if (need_resched_rcu(start)) {
			cond_resched_rcu();
			start = current_time();
	}

Where need_resched_rcu() tests to see if we're getting close to hitting
the watchdog timeout.

No?

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3)
  2026-08-05 20:28   ` Paul E. McKenney
  2026-08-06  6:03     ` Andrew Morton
@ 2026-08-06  8:18     ` Vlastimil Babka (SUSE)
  2026-08-06 17:19       ` Paul E. McKenney
  1 sibling, 1 reply; 7+ messages in thread
From: Vlastimil Babka (SUSE) @ 2026-08-06  8:18 UTC (permalink / raw)
  To: paulmck, Andrew Morton
  Cc: syzbot, hannes, jackmanb, linux-kernel, linux-mm, mhocko, surenb,
	syzkaller-bugs, ziy

On 8/5/26 22:28, Paul E. McKenney wrote:
> On Wed, Aug 05, 2026 at 12:29:52PM -0700, Andrew Morton wrote:
>> On Tue, 04 Aug 2026 17:01:48 -0700 syzbot <syzbot+d2401aeb74cc84adba04@syzkaller.appspotmail.com> wrote:
>> 
>> > Hello,
>> > 
>> > syzbot found the following issue on:
>> > 
>> > HEAD commit:    3708dd948844 Merge tag 'pm-7.2-rc6' of git://git.kernel.or..
>> > git tree:       upstream
>> > console output: https://syzkaller.appspot.com/x/log.txt?x=11ac703e580000
>> > kernel config:  https://syzkaller.appspot.com/x/.config?x=4e38b15c29e6a1d9
>> > dashboard link: https://syzkaller.appspot.com/bug?extid=d2401aeb74cc84adba04
>> > compiler:       Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
>> > 
>> > Unfortunately, I don't have any reproducer for this issue yet.
>> 
>> Thanks.
>> 
>> Lazy optimists (ahem) paste this gunk into Gemini and ask "what the
>> heck just happened".  The results are often useful, but should be
>> treated with skepticism.  In this case I think it came usably close.
>> 
>> 	https://share.gemini.google/vq4TLhTiLBih
>> 
>> 
>> tl;dr: khugepaged's collapse_scan_file() is taking too long and RCU got
>> starved.  I don't think khugepaged is doing anything wrong here,
>> per-se.  There's a lot of work to do and we're doing it.
>> 
>> An appropriate fix would be to take a break, let RCU do its thing then
>> get back to work.  But I don't think RCU offers interfaces for that?
>> 
>> collapse_scan_file()'s main loop has
>> 
>> 		if (need_resched()) {
>> 			xas_pause(&xas);
>> 			cond_resched_rcu();
>> 		}
>> 
>> but that won't help with the RCU stall detector(?).
>> 
>> I suggest that a suitable fix here would be to add the analogous
>> 
>> 	if (rcu_i_need_to_take_a_break()) {
>> 		rcu_read_unlock();
>> 		rcu_take_a_break())	
>> 		rcu_read_lock();
>> 	}
>> 
>> (iirc rcu_read_unlock() does an rcu run, so rcu_take_a_break() isn't
>> needed here)
>> 
>> Paul, wdyt?
> 
> Let's see...
> 
> The console log says "rcu_preempt detected stalls on CPUs/tasks",
> which means that cond_resched() is a no-op, but it also means that
> the rcu_read_unlock() in cond_resched_rcu() will directly take care of
> informing RCU of the pause.
> 
> But that is clearly not happening.  Why?
> 
> Well, we have this:
> 
> rcu: 	Tasks blocked on level-0 rcu_node (CPUs 0-1): P37/1:b..l
> 
> This means that the task whose RCU read-side critical section is blocking
> the current RCU grace period isn't even running, and thus cannot invoke
> cond_resched_rcu(), let alone the rcu_read_unlock() within that function.
> So an RCU CPU stall warning is expected behavior.  Or at least it is not
> in any way ruled out.
> 
> What we need is RCU priority boosting.  Except that the .config file
> does not enable this.  Not only is there no CONFIG_RCU_BOOST=y, there
> is also no CONFIG_RCU_EXPERT=y and no CONFIG_PREEMPT_RT=y.  But there
> is CONFIG_RT_MUTEX=y and CONFIG_RCU_EXPERT=y.

It comes from syzbot so might be likely a randconfig and there's no point in
trying to find any sense in that combination :)

> Because we don't have RCU priority boosting, if the load on the system
> is heavy enough to prevent our poor preempted RCU reader (PID 37) from
> running, the grace period cannot end.
> 
> I am not sure why this task is saving its stack, but maybe that is normal
> for this code path?

That's because page_owner is also enabled so it's saving the freeing stack
for the page it's freeing. That's not a normal production config, only when
debugging.

> My bemusement aside, I recommend running this test either with
> non-preemptible RCU (CONFIG_PREEMPT_LAZY=y these days) or enabling RCU
> priority boosting (CONFIG_RCU_EXPERT=y and CONFIG_RCU_BOOST=y).
> 
> Maybe RCU_BOOST should no longer depend on RCU_EXPERT?  I would of
> course need ot remove the prompt ("Enable RCU priority boosting") to
> avoid annoying Linus.  Maybe as shown below.

The "no longer depend" part alone would make no difference with randconfigs.
Removing the prompt too should help indeed.

Maybe a possible strategy in general would be indeed to unconditionally
select what's the expected config, like you did below, and only make it
possible to override that with RCU_EXPERT. So here with RCU_EXPERT you could
disable RCU_BOOST even if it was automatically enabled - assuming this is
useful for development or internal rcu testing by people who know what they
are doing (not syzbot randconfig) or whatnot.

But then RCU_EXPERT should be excluded from (impossible to be enabled by)
randconfig to indicate it's not valid for this kind of testing.
I don't know if there's any precedent for such a strategy.

Specifically for the proposal below, could the problem still happen with
PREEMPT_RCU without RT_MUTEXES? If yes, it wouldn't be enough?

> Thoughts?
> 
> 							Thanx, Paul
> 
> ------------------------------------------------------------------------
> 
> diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig
> index 1a5fb3156c062a..5141ad8d1cd029 100644
> --- a/kernel/rcu/Kconfig
> +++ b/kernel/rcu/Kconfig
> @@ -237,17 +237,16 @@ config RCU_FANOUT_LEAF
>  	  Take the default if unsure.
>  
>  config RCU_BOOST
> -	bool "Enable RCU priority boosting"
> -	depends on (RT_MUTEXES && PREEMPT_RCU && RCU_EXPERT) || PREEMPT_RT
> +	bool
> +	depends on (RT_MUTEXES && PREEMPT_RCU) || PREEMPT_RT
>  	default y if PREEMPT_RT
>  	help
>  	  This option boosts the priority of preempted RCU readers that
>  	  block the current preemptible RCU grace period for too long.
>  	  This option also prevents heavy loads from blocking RCU
> -	  callback invocation.
> +	  callback invocation.  It is now automatically enabled in
> +	  any kernel that can benefit from it and that can support it.
>  
> -	  Say Y here if you are working with real-time apps or heavy loads
> -	  Say N here if you are unsure.
>  
>  config RCU_BOOST_DELAY
>  	int "Milliseconds to delay boosting after RCU grace-period start"


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3)
  2026-08-06  6:03     ` Andrew Morton
@ 2026-08-06 16:27       ` Paul E. McKenney
  0 siblings, 0 replies; 7+ messages in thread
From: Paul E. McKenney @ 2026-08-06 16:27 UTC (permalink / raw)
  To: Andrew Morton
  Cc: syzbot, hannes, jackmanb, linux-kernel, linux-mm, mhocko, surenb,
	syzkaller-bugs, vbabka, ziy

On Wed, Aug 05, 2026 at 11:03:11PM -0700, Andrew Morton wrote:
> On Wed, 5 Aug 2026 13:28:16 -0700 "Paul E. McKenney" <paulmck@kernel.org> wrote:
> 
> > > collapse_scan_file()'s main loop has
> > > 
> > > 		if (need_resched()) {
> > > 			xas_pause(&xas);
> > > 			cond_resched_rcu();
> > > 		}
> > > 
> > > but that won't help with the RCU stall detector(?).
> > > 
> > > I suggest that a suitable fix here would be to add the analogous
> > > 
> > > 	if (rcu_i_need_to_take_a_break()) {
> > > 		rcu_read_unlock();
> > > 		rcu_take_a_break())	
> > > 		rcu_read_lock();
> > > 	}
> > > 
> > > (iirc rcu_read_unlock() does an rcu run, so rcu_take_a_break() isn't
> > > needed here)
> > > 
> > > Paul, wdyt?
> > 
> > Let's see...
> > 
> > The console log says "rcu_preempt detected stalls on CPUs/tasks",
> > which means that cond_resched() is a no-op, but it also means that
> > the rcu_read_unlock() in cond_resched_rcu() will directly take care of
> > informing RCU of the pause.
> > 
> > But that is clearly not happening.  Why?
> > 
> > Well, we have this:
> > 
> > rcu: 	Tasks blocked on level-0 rcu_node (CPUs 0-1): P37/1:b..l
> > 
> > This means that the task whose RCU read-side critical section is blocking
> > the current RCU grace period isn't even running, and thus cannot invoke
> > cond_resched_rcu(), let alone the rcu_read_unlock() within that function.
> > So an RCU CPU stall warning is expected behavior.  Or at least it is not
> > in any way ruled out.
> > 
> > What we need is RCU priority boosting.
> 
> Do we?  I'm suggesting we need need_resched_rcu()!

Understood, and I initially agreed with you.  Except that I then
found that the poor preempted task isn't executing anything at all.
Which means that an added need_resched_rcu() cannot possibly help.

> > Except that the .config file
> > does not enable this.  Not only is there no CONFIG_RCU_BOOST=y, there
> > is also no CONFIG_RCU_EXPERT=y and no CONFIG_PREEMPT_RT=y.  But there
> > is CONFIG_RT_MUTEX=y and CONFIG_RCU_EXPERT=y.
> > 
> > Because we don't have RCU priority boosting, if the load on the system
> > is heavy enough to prevent our poor preempted RCU reader (PID 37) from
> > running, the grace period cannot end.
> > 
> > I am not sure why this task is saving its stack, but maybe that is normal
> > for this code path?
> > 
> > My bemusement aside, I recommend running this test either with
> > non-preemptible RCU (CONFIG_PREEMPT_LAZY=y these days) or enabling RCU
> > priority boosting (CONFIG_RCU_EXPERT=y and CONFIG_RCU_BOOST=y).
> > 
> > Maybe RCU_BOOST should no longer depend on RCU_EXPERT?  I would of
> > course need ot remove the prompt ("Enable RCU priority boosting") to
> > avoid annoying Linus.  Maybe as shown below.
> > 
> > Thoughts?
> 
> If I'm understanding correctly, this workload is busted with this
> config and the proposed fix is to alter the config?  Well, why are we
> permitting that config at all?

Agreed, and my proposal is in fact to remove that config.  This assumes
that CONFIG_RCU_BOOST is ready for prime time, and given that I have
been testing it for years and that CONFIG_PREEMPT_RT has been using it
for years, I am cautiously optimistic.

> Seems to me that a solution to permit this config to work is very
> simple.  Something like:
> 
> 	time_t start;
> 
> 	rcu_read_lock();
> 	start = current_time();
> 
> 	for (lots of work) {
> 		...
> 		if (need_resched_rcu(start)) {
> 			cond_resched_rcu();
> 			start = current_time();
> 	}
> 
> Where need_resched_rcu() tests to see if we're getting close to hitting
> the watchdog timeout.
> 
> No?

No.

Added code doesn't help a task that has been preempted for some seconds
and thus isn't executing any code at all.

The task has been preempted while within an RCU read-side critical
section.  The sequence of events is as follows:

	rcu_read_lock();
	// preempted for many tens of seconds.
	cond_resched_rcu(); // doesn't help because it is never executed.
	rcu_read_unlock();

Or am I missing your point?

							Thanx, Paul

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3)
  2026-08-06  8:18     ` Vlastimil Babka (SUSE)
@ 2026-08-06 17:19       ` Paul E. McKenney
  0 siblings, 0 replies; 7+ messages in thread
From: Paul E. McKenney @ 2026-08-06 17:19 UTC (permalink / raw)
  To: Vlastimil Babka (SUSE)
  Cc: Andrew Morton, syzbot, hannes, jackmanb, linux-kernel, linux-mm,
	mhocko, surenb, syzkaller-bugs, ziy

On Thu, Aug 06, 2026 at 10:18:57AM +0200, Vlastimil Babka (SUSE) wrote:
> On 8/5/26 22:28, Paul E. McKenney wrote:
> > On Wed, Aug 05, 2026 at 12:29:52PM -0700, Andrew Morton wrote:
> >> On Tue, 04 Aug 2026 17:01:48 -0700 syzbot <syzbot+d2401aeb74cc84adba04@syzkaller.appspotmail.com> wrote:
> >> 
> >> > Hello,
> >> > 
> >> > syzbot found the following issue on:
> >> > 
> >> > HEAD commit:    3708dd948844 Merge tag 'pm-7.2-rc6' of git://git.kernel.or..
> >> > git tree:       upstream
> >> > console output: https://syzkaller.appspot.com/x/log.txt?x=11ac703e580000
> >> > kernel config:  https://syzkaller.appspot.com/x/.config?x=4e38b15c29e6a1d9
> >> > dashboard link: https://syzkaller.appspot.com/bug?extid=d2401aeb74cc84adba04
> >> > compiler:       Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
> >> > 
> >> > Unfortunately, I don't have any reproducer for this issue yet.
> >> 
> >> Thanks.
> >> 
> >> Lazy optimists (ahem) paste this gunk into Gemini and ask "what the
> >> heck just happened".  The results are often useful, but should be
> >> treated with skepticism.  In this case I think it came usably close.
> >> 
> >> 	https://share.gemini.google/vq4TLhTiLBih
> >> 
> >> 
> >> tl;dr: khugepaged's collapse_scan_file() is taking too long and RCU got
> >> starved.  I don't think khugepaged is doing anything wrong here,
> >> per-se.  There's a lot of work to do and we're doing it.
> >> 
> >> An appropriate fix would be to take a break, let RCU do its thing then
> >> get back to work.  But I don't think RCU offers interfaces for that?
> >> 
> >> collapse_scan_file()'s main loop has
> >> 
> >> 		if (need_resched()) {
> >> 			xas_pause(&xas);
> >> 			cond_resched_rcu();
> >> 		}
> >> 
> >> but that won't help with the RCU stall detector(?).
> >> 
> >> I suggest that a suitable fix here would be to add the analogous
> >> 
> >> 	if (rcu_i_need_to_take_a_break()) {
> >> 		rcu_read_unlock();
> >> 		rcu_take_a_break())	
> >> 		rcu_read_lock();
> >> 	}
> >> 
> >> (iirc rcu_read_unlock() does an rcu run, so rcu_take_a_break() isn't
> >> needed here)
> >> 
> >> Paul, wdyt?
> > 
> > Let's see...
> > 
> > The console log says "rcu_preempt detected stalls on CPUs/tasks",
> > which means that cond_resched() is a no-op, but it also means that
> > the rcu_read_unlock() in cond_resched_rcu() will directly take care of
> > informing RCU of the pause.
> > 
> > But that is clearly not happening.  Why?
> > 
> > Well, we have this:
> > 
> > rcu: 	Tasks blocked on level-0 rcu_node (CPUs 0-1): P37/1:b..l
> > 
> > This means that the task whose RCU read-side critical section is blocking
> > the current RCU grace period isn't even running, and thus cannot invoke
> > cond_resched_rcu(), let alone the rcu_read_unlock() within that function.
> > So an RCU CPU stall warning is expected behavior.  Or at least it is not
> > in any way ruled out.
> > 
> > What we need is RCU priority boosting.  Except that the .config file
> > does not enable this.  Not only is there no CONFIG_RCU_BOOST=y, there
> > is also no CONFIG_RCU_EXPERT=y and no CONFIG_PREEMPT_RT=y.  But there
> > is CONFIG_RT_MUTEX=y and CONFIG_RCU_EXPERT=y.
> 
> It comes from syzbot so might be likely a randconfig and there's no point in
> trying to find any sense in that combination :)

;-) ;-) ;-)

> > Because we don't have RCU priority boosting, if the load on the system
> > is heavy enough to prevent our poor preempted RCU reader (PID 37) from
> > running, the grace period cannot end.
> > 
> > I am not sure why this task is saving its stack, but maybe that is normal
> > for this code path?
> 
> That's because page_owner is also enabled so it's saving the freeing stack
> for the page it's freeing. That's not a normal production config, only when
> debugging.

Ah, OK, I feel much better now.

> > My bemusement aside, I recommend running this test either with
> > non-preemptible RCU (CONFIG_PREEMPT_LAZY=y these days) or enabling RCU
> > priority boosting (CONFIG_RCU_EXPERT=y and CONFIG_RCU_BOOST=y).
> > 
> > Maybe RCU_BOOST should no longer depend on RCU_EXPERT?  I would of
> > course need ot remove the prompt ("Enable RCU priority boosting") to
> > avoid annoying Linus.  Maybe as shown below.
> 
> The "no longer depend" part alone would make no difference with randconfigs.
> Removing the prompt too should help indeed.

Agreed!

> Maybe a possible strategy in general would be indeed to unconditionally
> select what's the expected config, like you did below, and only make it
> possible to override that with RCU_EXPERT. So here with RCU_EXPERT you could
> disable RCU_BOOST even if it was automatically enabled - assuming this is
> useful for development or internal rcu testing by people who know what they
> are doing (not syzbot randconfig) or whatnot.
> 
> But then RCU_EXPERT should be excluded from (impossible to be enabled by)
> randconfig to indicate it's not valid for this kind of testing.
> I don't know if there's any precedent for such a strategy.

Good point!  For the first cut, I will just force it, so that someone
wanting to do that sort of testing gets to edit the Kconfig file, but
it is good to have a trick like that in my back pocket, so thank you!

> Specifically for the proposal below, could the problem still happen with
> PREEMPT_RCU without RT_MUTEXES? If yes, it wouldn't be enough?

Quite true!  Maybe I should make PREEMPT_RCU select RT_MUTEXES?

But that might need a bit of discussion, so if I take that approach
it needs to be a separate patch.  ;-)

							Thanx, Paul

> > ------------------------------------------------------------------------
> > 
> > diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig
> > index 1a5fb3156c062a..5141ad8d1cd029 100644
> > --- a/kernel/rcu/Kconfig
> > +++ b/kernel/rcu/Kconfig
> > @@ -237,17 +237,16 @@ config RCU_FANOUT_LEAF
> >  	  Take the default if unsure.
> >  
> >  config RCU_BOOST
> > -	bool "Enable RCU priority boosting"
> > -	depends on (RT_MUTEXES && PREEMPT_RCU && RCU_EXPERT) || PREEMPT_RT
> > +	bool
> > +	depends on (RT_MUTEXES && PREEMPT_RCU) || PREEMPT_RT
> >  	default y if PREEMPT_RT
> >  	help
> >  	  This option boosts the priority of preempted RCU readers that
> >  	  block the current preemptible RCU grace period for too long.
> >  	  This option also prevents heavy loads from blocking RCU
> > -	  callback invocation.
> > +	  callback invocation.  It is now automatically enabled in
> > +	  any kernel that can benefit from it and that can support it.
> >  
> > -	  Say Y here if you are working with real-time apps or heavy loads
> > -	  Say N here if you are unsure.
> >  
> >  config RCU_BOOST_DELAY
> >  	int "Milliseconds to delay boosting after RCU grace-period start"
> 

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-08-06 17:19 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-05  0:01 [syzbot] [mm?] INFO: rcu detected stall in khugepaged (3) syzbot
2026-08-05 19:29 ` Andrew Morton
2026-08-05 20:28   ` Paul E. McKenney
2026-08-06  6:03     ` Andrew Morton
2026-08-06 16:27       ` Paul E. McKenney
2026-08-06  8:18     ` Vlastimil Babka (SUSE)
2026-08-06 17:19       ` Paul E. McKenney

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox