BPF List
 help / color / mirror / Atom feed
* [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan
@ 2026-07-20 13:23 Breno Leitao
  2026-07-20 13:37 ` sashiko-bot
                   ` (2 more replies)
  0 siblings, 3 replies; 4+ messages in thread
From: Breno Leitao @ 2026-07-20 13:23 UTC (permalink / raw)
  To: Catalin Marinas, Andrew Morton
  Cc: paulmck, puranjay, linux-mm, linux-kernel, bpf, kernel-team,
	Breno Leitao

kmemleak_scan() can run for ages on large debug kernels. It was
causing some soft-lockups which I got fixed with commit
3175fcfec8b16baeb ("mm/kmemleak: avoid soft lockup when scanning task
stacks") with our beloved cond_resched().

I've got the fix above deployed in the Meta fleet, and now I am seeing:

  INFO: rcu_tasks detected stalls on tasks:
    task:kmemleak state:R ... nvcsw: 274/274 holdout: 1 idle_cpu: -1/3
     scan_block
     scan_gray_list
     kmemleak_scan

and, worse, blocks the callers waiting on that grace period. Here a BPF
struct_ops map free, which waits via synchronize_rcu_mult(call_rcu,
call_rcu_tasks), is stuck long enough to also trip the hung task check:

  INFO: task kworker/...:bpf_map_free_deferred blocked for 122 seconds
     __wait_rcu_gp
     bpf_struct_ops_map_free

Then I've learned that cond_resched() is not an RCU-tasks quiescent
state, so, we need to use stronger primitives.

Use cond_resched_tasks_rcu_qs() at the scan reschedule points so the scan
reports an RCU-tasks quiescent state as it proceeds.

Inspired by commit b96285e10aad ("tracing: Have osnoise_main() add a
quiescent state for task rcu").

Signed-off-by: Breno Leitao <leitao@debian.org>
---
 mm/kmemleak.c | 14 +++++++-------
 1 file changed, 7 insertions(+), 7 deletions(-)

diff --git a/mm/kmemleak.c b/mm/kmemleak.c
index 85f18b17e79c4..f63dfacee7ca1 100644
--- a/mm/kmemleak.c
+++ b/mm/kmemleak.c
@@ -1586,7 +1586,7 @@ static int scan_large_block(void *start, void *end)
 		if (scan_block(start, next, NULL))
 			return 1;
 		start = next;
-		cond_resched();
+		cond_resched_tasks_rcu_qs();
 	}
 
 	return 0;
@@ -1623,7 +1623,7 @@ static void scan_object(struct kmemleak_object *object)
 			scan_block(start, end, object);
 
 			raw_spin_unlock_irqrestore(&object->lock, flags);
-			cond_resched();
+			cond_resched_tasks_rcu_qs();
 			raw_spin_lock_irqsave(&object->lock, flags);
 			if (!(object->flags & OBJECT_ALLOCATED))
 				break;
@@ -1645,7 +1645,7 @@ static void scan_object(struct kmemleak_object *object)
 				break;
 
 			raw_spin_unlock_irqrestore(&object->lock, flags);
-			cond_resched();
+			cond_resched_tasks_rcu_qs();
 			raw_spin_lock_irqsave(&object->lock, flags);
 		} while (object->flags & OBJECT_ALLOCATED);
 	} else {
@@ -1673,7 +1673,7 @@ static void scan_gray_list(void)
 	 */
 	object = list_entry(gray_list.next, typeof(*object), gray_list);
 	while (&object->gray_list != &gray_list) {
-		cond_resched();
+		cond_resched_tasks_rcu_qs();
 
 		/* may add new objects to the list */
 		if (!scan_should_stop())
@@ -1708,7 +1708,7 @@ static void kmemleak_cond_resched(struct kmemleak_object *object)
 	raw_spin_unlock_irq(&kmemleak_lock);
 
 	rcu_read_unlock();
-	cond_resched();
+	cond_resched_tasks_rcu_qs();
 	rcu_read_lock();
 
 	raw_spin_lock_irq(&kmemleak_lock);
@@ -1753,7 +1753,7 @@ static void kmemleak_scan_task_stacks(void)
 			}
 			put_task_struct(p);
 		}
-		cond_resched();
+		cond_resched_tasks_rcu_qs();
 	} while (pid && !stop);
 }
 
@@ -1935,7 +1935,7 @@ static int __kmemleak_scan(bool full)
 			struct page *page = pfn_to_online_page(pfn);
 
 			if (!(pfn & 63))
-				cond_resched();
+				cond_resched_tasks_rcu_qs();
 
 			if (!page)
 				continue;

---
base-commit: 0718283ab28bc3907e10b61a6b4be6fefa1cbb2f
change-id: 20260720-kmemleak_rcu_task-6d3ea111b947

Best regards,
--  
Breno Leitao <leitao@debian.org>


^ permalink raw reply related	[flat|nested] 4+ messages in thread

* Re: [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan
  2026-07-20 13:23 [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan Breno Leitao
@ 2026-07-20 13:37 ` sashiko-bot
  2026-07-20 15:25 ` Paul E. McKenney
  2026-07-20 22:39 ` Andrew Morton
  2 siblings, 0 replies; 4+ messages in thread
From: sashiko-bot @ 2026-07-20 13:37 UTC (permalink / raw)
  To: Breno Leitao; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

Pre-existing issues:
- [High] Incomplete fix: Tasks RCU stalls persist in __kmemleak_do_cleanup() and __kmemleak_scan().
- [High] Unbounded RCU critical section in kmemleak_clear() stalls Normal RCU.
--

--- Patch [1]: [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan ---
commit 558fef0ac4a016de42546b6d38536df76b835edb
Author: Breno Leitao <leitao@debian.org>

mm/kmemleak: report RCU-tasks quiescent states during the scan

This patch replaces cond_resched() with cond_resched_tasks_rcu_qs() at various
reschedule points during the kmemleak scan. This aims to report Tasks RCU
quiescent states and prevent Tasks RCU stalls and hung tasks, since
cond_resched() by itself is not a Tasks RCU quiescent state.

> diff --git a/mm/kmemleak.c b/mm/kmemleak.c
> index 85f18b17e79c4..f63dfacee7ca1 100644
> --- a/mm/kmemleak.c
> +++ b/mm/kmemleak.c
[ ... ]
> @@ -1708,7 +1708,7 @@ static void kmemleak_cond_resched(struct kmemleak_object *object)
>  	raw_spin_unlock_irq(&kmemleak_lock);
>  
>  	rcu_read_unlock();
> -	cond_resched();
> +	cond_resched_tasks_rcu_qs();
>  	rcu_read_lock();
>  
>  	raw_spin_lock_irq(&kmemleak_lock);

[Severity: High]
This is a pre-existing issue, but modifying kmemleak_cond_resched() here
might not fully resolve the Tasks RCU stall.

kmemleak_cond_resched() is only called in __kmemleak_scan() when
need_resched() is true:

	if (need_resched())
		kmemleak_cond_resched(object);

Since Tasks RCU does not actively set the need_resched flag on running tasks
(unlike normal RCU, which uses inter-processor interrupts during a stall),
need_resched() could remain false on an otherwise idle system.

Does this mean the Tasks RCU quiescent state update is never reached,
bypassing this new logic and leaving the stall bug intact?

[ ... ]

[Severity: High]
Additionally, this is a pre-existing issue, but does __kmemleak_do_cleanup()
need a similar fix?

It still uses the bare cond_resched() in its loop:

mm/kmemleak.c:__kmemleak_do_cleanup() {
	list_for_each_entry_safe(object, tmp, &object_list, object_list) {
		...
		/* Call cond_resched() once per 64 iterations to avoid soft lockup */
		if (!(++cnt & 0x3f))
			cond_resched();
	}
}

Will this cause Tasks RCU grace periods to stall during cleanup work, just
like they were stalling during the scan?

[Severity: High]
This is another pre-existing issue, but can the unyielding loop in
kmemleak_clear() stall normal RCU?

When a privileged user writes 'clear' to /sys/kernel/debug/kmemleak, it
invokes kmemleak_clear(), which iterates over the entire object_list:

mm/kmemleak.c:kmemleak_clear() {
	rcu_read_lock();
	list_for_each_entry_rcu(object, &object_list, object_list) {
		raw_spin_lock_irq(&object->lock);
		...
		raw_spin_unlock_irq(&object->lock);
	}
	rcu_read_unlock();
}

Since this loop runs fully inside an RCU read-side critical section with no
yielding mechanisms whatsoever, could iterating millions of objects
uninterrupted exceed the 21-second RCU stall timeout and trigger normal
RCU stall warnings?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260720-kmemleak_rcu_task-v1-1-5b460ade777d@debian.org?part=1

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan
  2026-07-20 13:23 [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan Breno Leitao
  2026-07-20 13:37 ` sashiko-bot
@ 2026-07-20 15:25 ` Paul E. McKenney
  2026-07-20 22:39 ` Andrew Morton
  2 siblings, 0 replies; 4+ messages in thread
From: Paul E. McKenney @ 2026-07-20 15:25 UTC (permalink / raw)
  To: Breno Leitao
  Cc: Catalin Marinas, Andrew Morton, puranjay, linux-mm, linux-kernel,
	bpf, kernel-team

On Mon, Jul 20, 2026 at 06:23:45AM -0700, Breno Leitao wrote:
> kmemleak_scan() can run for ages on large debug kernels. It was
> causing some soft-lockups which I got fixed with commit
> 3175fcfec8b16baeb ("mm/kmemleak: avoid soft lockup when scanning task
> stacks") with our beloved cond_resched().
> 
> I've got the fix above deployed in the Meta fleet, and now I am seeing:
> 
>   INFO: rcu_tasks detected stalls on tasks:
>     task:kmemleak state:R ... nvcsw: 274/274 holdout: 1 idle_cpu: -1/3
>      scan_block
>      scan_gray_list
>      kmemleak_scan
> 
> and, worse, blocks the callers waiting on that grace period. Here a BPF
> struct_ops map free, which waits via synchronize_rcu_mult(call_rcu,
> call_rcu_tasks), is stuck long enough to also trip the hung task check:
> 
>   INFO: task kworker/...:bpf_map_free_deferred blocked for 122 seconds
>      __wait_rcu_gp
>      bpf_struct_ops_map_free
> 
> Then I've learned that cond_resched() is not an RCU-tasks quiescent
> state, so, we need to use stronger primitives.
> 
> Use cond_resched_tasks_rcu_qs() at the scan reschedule points so the scan
> reports an RCU-tasks quiescent state as it proceeds.
> 
> Inspired by commit b96285e10aad ("tracing: Have osnoise_main() add a
> quiescent state for task rcu").
> 
> Signed-off-by: Breno Leitao <leitao@debian.org>

Reviewed-by: Paul E. McKenney <paulmck@kernel.org>

> ---
>  mm/kmemleak.c | 14 +++++++-------
>  1 file changed, 7 insertions(+), 7 deletions(-)
> 
> diff --git a/mm/kmemleak.c b/mm/kmemleak.c
> index 85f18b17e79c4..f63dfacee7ca1 100644
> --- a/mm/kmemleak.c
> +++ b/mm/kmemleak.c
> @@ -1586,7 +1586,7 @@ static int scan_large_block(void *start, void *end)
>  		if (scan_block(start, next, NULL))
>  			return 1;
>  		start = next;
> -		cond_resched();
> +		cond_resched_tasks_rcu_qs();
>  	}
>  
>  	return 0;
> @@ -1623,7 +1623,7 @@ static void scan_object(struct kmemleak_object *object)
>  			scan_block(start, end, object);
>  
>  			raw_spin_unlock_irqrestore(&object->lock, flags);
> -			cond_resched();
> +			cond_resched_tasks_rcu_qs();
>  			raw_spin_lock_irqsave(&object->lock, flags);
>  			if (!(object->flags & OBJECT_ALLOCATED))
>  				break;
> @@ -1645,7 +1645,7 @@ static void scan_object(struct kmemleak_object *object)
>  				break;
>  
>  			raw_spin_unlock_irqrestore(&object->lock, flags);
> -			cond_resched();
> +			cond_resched_tasks_rcu_qs();
>  			raw_spin_lock_irqsave(&object->lock, flags);
>  		} while (object->flags & OBJECT_ALLOCATED);
>  	} else {
> @@ -1673,7 +1673,7 @@ static void scan_gray_list(void)
>  	 */
>  	object = list_entry(gray_list.next, typeof(*object), gray_list);
>  	while (&object->gray_list != &gray_list) {
> -		cond_resched();
> +		cond_resched_tasks_rcu_qs();
>  
>  		/* may add new objects to the list */
>  		if (!scan_should_stop())
> @@ -1708,7 +1708,7 @@ static void kmemleak_cond_resched(struct kmemleak_object *object)
>  	raw_spin_unlock_irq(&kmemleak_lock);
>  
>  	rcu_read_unlock();
> -	cond_resched();
> +	cond_resched_tasks_rcu_qs();
>  	rcu_read_lock();
>  
>  	raw_spin_lock_irq(&kmemleak_lock);
> @@ -1753,7 +1753,7 @@ static void kmemleak_scan_task_stacks(void)
>  			}
>  			put_task_struct(p);
>  		}
> -		cond_resched();
> +		cond_resched_tasks_rcu_qs();
>  	} while (pid && !stop);
>  }
>  
> @@ -1935,7 +1935,7 @@ static int __kmemleak_scan(bool full)
>  			struct page *page = pfn_to_online_page(pfn);
>  
>  			if (!(pfn & 63))
> -				cond_resched();
> +				cond_resched_tasks_rcu_qs();
>  
>  			if (!page)
>  				continue;
> 
> ---
> base-commit: 0718283ab28bc3907e10b61a6b4be6fefa1cbb2f
> change-id: 20260720-kmemleak_rcu_task-6d3ea111b947
> 
> Best regards,
> --  
> Breno Leitao <leitao@debian.org>
> 

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan
  2026-07-20 13:23 [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan Breno Leitao
  2026-07-20 13:37 ` sashiko-bot
  2026-07-20 15:25 ` Paul E. McKenney
@ 2026-07-20 22:39 ` Andrew Morton
  2 siblings, 0 replies; 4+ messages in thread
From: Andrew Morton @ 2026-07-20 22:39 UTC (permalink / raw)
  To: Breno Leitao
  Cc: Catalin Marinas, paulmck, puranjay, linux-mm, linux-kernel, bpf,
	kernel-team

On Mon, 20 Jul 2026 06:23:45 -0700 Breno Leitao <leitao@debian.org> wrote:

> kmemleak_scan() can run for ages on large debug kernels. It was
> causing some soft-lockups which I got fixed with commit
> 3175fcfec8b16baeb ("mm/kmemleak: avoid soft lockup when scanning task
> stacks") with our beloved cond_resched().
> 
> I've got the fix above deployed in the Meta fleet, and now I am seeing:
> 
>   INFO: rcu_tasks detected stalls on tasks:
>     task:kmemleak state:R ... nvcsw: 274/274 holdout: 1 idle_cpu: -1/3
>      scan_block
>      scan_gray_list
>      kmemleak_scan
> 
> and, worse, blocks the callers waiting on that grace period. Here a BPF
> struct_ops map free, which waits via synchronize_rcu_mult(call_rcu,
> call_rcu_tasks), is stuck long enough to also trip the hung task check:
> 
>   INFO: task kworker/...:bpf_map_free_deferred blocked for 122 seconds
>      __wait_rcu_gp
>      bpf_struct_ops_map_free
> 
> Then I've learned that cond_resched() is not an RCU-tasks quiescent
> state, so, we need to use stronger primitives.
> 
> Use cond_resched_tasks_rcu_qs() at the scan reschedule points so the scan
> reports an RCU-tasks quiescent state as it proceeds.
> 
> Inspired by commit b96285e10aad ("tracing: Have osnoise_main() add a
> quiescent state for task rcu").

I'll add

Fixes: c4b28963fd79 ("mm/kmemleak: rely on rcu for task stack scanning")
Cc: <stable@vger.kernel.org>
.
.


^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-07-20 22:39 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-20 13:23 [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan Breno Leitao
2026-07-20 13:37 ` sashiko-bot
2026-07-20 15:25 ` Paul E. McKenney
2026-07-20 22:39 ` Andrew Morton

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox