The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
@ 2026-08-10  9:57 Breno Leitao
  2026-08-10 15:38 ` Paul E. McKenney
                   ` (2 more replies)
  0 siblings, 3 replies; 7+ messages in thread
From: Breno Leitao @ 2026-08-10  9:57 UTC (permalink / raw)
  To: Andrew Morton, Johannes Weiner, David Hildenbrand, Michal Hocko,
	Qi Zheng, Shakeel Butt, Lorenzo Stoakes, Kairui Song, Barry Song,
	Axel Rasmussen, Yuanchu Xie, Wei Xu, paulmck
  Cc: linux-mm, linux-kernel, kernel-team, stable, Breno Leitao

I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.

  INFO: rcu_tasks detected stalls on tasks:
	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
  Call Trace:
   shrink_lruvec
   mem_cgroup_iter
   shrink_node
   do_try_to_free_pages
   try_to_free_pages
   __alloc_frozen_pages_noprof
   alloc_pages_noprof
   pte_alloc_one
   __pte_alloc
   handle_mm_fault

Nothing promises direct reclaim returns in bounded time, and the scan
loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
quiescent state, so the reclaiming task never reports one and becomes a
holdout.

Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
state even when cond_resched() does nothing.

PS: This has been discussed in [1]

Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
Cc: stable@vger.kernel.org
Signed-off-by: Breno Leitao <leitao@debian.org>
---
 mm/vmscan.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/mm/vmscan.c b/mm/vmscan.c
index 26436059ea394..6ac2fde137b89 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -6023,7 +6023,7 @@ static void shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc)
 			}
 		}
 
-		cond_resched();
+		cond_resched_tasks_rcu_qs();
 
 		if (nr_reclaimed < nr_to_reclaim || proportional_reclaim)
 			continue;

---
base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727
change-id: 20260810-rcu_task_shrink_lruvec-711112e87de0

Best regards,
--  
Breno Leitao <leitao@debian.org>


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
  2026-08-10  9:57 [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec() Breno Leitao
@ 2026-08-10 15:38 ` Paul E. McKenney
  2026-08-10 20:02 ` Johannes Weiner
  2026-08-11  4:43 ` Shakeel Butt
  2 siblings, 0 replies; 7+ messages in thread
From: Paul E. McKenney @ 2026-08-10 15:38 UTC (permalink / raw)
  To: Breno Leitao
  Cc: Andrew Morton, Johannes Weiner, David Hildenbrand, Michal Hocko,
	Qi Zheng, Shakeel Butt, Lorenzo Stoakes, Kairui Song, Barry Song,
	Axel Rasmussen, Yuanchu Xie, Wei Xu, linux-mm, linux-kernel,
	kernel-team, stable

On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> 
>   INFO: rcu_tasks detected stalls on tasks:
> 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
>   Call Trace:
>    shrink_lruvec
>    mem_cgroup_iter
>    shrink_node
>    do_try_to_free_pages
>    try_to_free_pages
>    __alloc_frozen_pages_noprof
>    alloc_pages_noprof
>    pte_alloc_one
>    __pte_alloc
>    handle_mm_fault
> 
> Nothing promises direct reclaim returns in bounded time, and the scan
> loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> quiescent state, so the reclaiming task never reports one and becomes a
> holdout.
> 
> Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
> state even when cond_resched() does nothing.
> 
> PS: This has been discussed in [1]
> 
> Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
> Cc: stable@vger.kernel.org
> Signed-off-by: Breno Leitao <leitao@debian.org>

Reviewed-by: Paul E. McKenney <paulmck@kernel.org>

> ---
>  mm/vmscan.c | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
> 
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 26436059ea394..6ac2fde137b89 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -6023,7 +6023,7 @@ static void shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc)
>  			}
>  		}
>  
> -		cond_resched();
> +		cond_resched_tasks_rcu_qs();
>  
>  		if (nr_reclaimed < nr_to_reclaim || proportional_reclaim)
>  			continue;
> 
> ---
> base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727
> change-id: 20260810-rcu_task_shrink_lruvec-711112e87de0
> 
> Best regards,
> --  
> Breno Leitao <leitao@debian.org>
> 

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
  2026-08-10  9:57 [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec() Breno Leitao
  2026-08-10 15:38 ` Paul E. McKenney
@ 2026-08-10 20:02 ` Johannes Weiner
  2026-08-11  4:43 ` Shakeel Butt
  2 siblings, 0 replies; 7+ messages in thread
From: Johannes Weiner @ 2026-08-10 20:02 UTC (permalink / raw)
  To: Breno Leitao
  Cc: Andrew Morton, David Hildenbrand, Michal Hocko, Qi Zheng,
	Shakeel Butt, Lorenzo Stoakes, Kairui Song, Barry Song,
	Axel Rasmussen, Yuanchu Xie, Wei Xu, paulmck, linux-mm,
	linux-kernel, kernel-team, stable

On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> 
>   INFO: rcu_tasks detected stalls on tasks:
> 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
>   Call Trace:
>    shrink_lruvec
>    mem_cgroup_iter
>    shrink_node
>    do_try_to_free_pages
>    try_to_free_pages
>    __alloc_frozen_pages_noprof
>    alloc_pages_noprof
>    pte_alloc_one
>    __pte_alloc
>    handle_mm_fault
> 
> Nothing promises direct reclaim returns in bounded time, and the scan
> loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> quiescent state, so the reclaiming task never reports one and becomes a
> holdout.
> 
> Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
> state even when cond_resched() does nothing.
> 
> PS: This has been discussed in [1]
> 
> Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
> Cc: stable@vger.kernel.org
> Signed-off-by: Breno Leitao <leitao@debian.org>

Acked-by: Johannes Weiner <hannes@cmpxchg.org>

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
  2026-08-10  9:57 [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec() Breno Leitao
  2026-08-10 15:38 ` Paul E. McKenney
  2026-08-10 20:02 ` Johannes Weiner
@ 2026-08-11  4:43 ` Shakeel Butt
  2026-08-11 17:38   ` Paul E. McKenney
  2 siblings, 1 reply; 7+ messages in thread
From: Shakeel Butt @ 2026-08-11  4:43 UTC (permalink / raw)
  To: Breno Leitao
  Cc: Andrew Morton, Johannes Weiner, David Hildenbrand, Michal Hocko,
	Qi Zheng, Lorenzo Stoakes, Kairui Song, Barry Song,
	Axel Rasmussen, Yuanchu Xie, Wei Xu, paulmck, linux-mm,
	linux-kernel, kernel-team, stable

On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> 
>   INFO: rcu_tasks detected stalls on tasks:
> 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
>   Call Trace:
>    shrink_lruvec
>    mem_cgroup_iter
>    shrink_node
>    do_try_to_free_pages
>    try_to_free_pages
>    __alloc_frozen_pages_noprof
>    alloc_pages_noprof
>    pte_alloc_one
>    __pte_alloc
>    handle_mm_fault
> 
> Nothing promises direct reclaim returns in bounded time, and the scan
> loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> quiescent state, so the reclaiming task never reports one and becomes a
> holdout.

I still don't understand why cond_resched() is being treated as involuntary
preemption but that is orthogonal to this patch.

> 
> Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
> state even when cond_resched() does nothing.
> 
> PS: This has been discussed in [1]
> 
> Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
> Cc: stable@vger.kernel.org
> Signed-off-by: Breno Leitao <leitao@debian.org>

Acked-by: Shakeel Butt <shakeel.butt@linux.dev>


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
  2026-08-11  4:43 ` Shakeel Butt
@ 2026-08-11 17:38   ` Paul E. McKenney
  2026-08-11 17:54     ` Shakeel Butt
  0 siblings, 1 reply; 7+ messages in thread
From: Paul E. McKenney @ 2026-08-11 17:38 UTC (permalink / raw)
  To: Shakeel Butt
  Cc: Breno Leitao, Andrew Morton, Johannes Weiner, David Hildenbrand,
	Michal Hocko, Qi Zheng, Lorenzo Stoakes, Kairui Song, Barry Song,
	Axel Rasmussen, Yuanchu Xie, Wei Xu, linux-mm, linux-kernel,
	kernel-team, stable

On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote:
> On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> > 
> >   INFO: rcu_tasks detected stalls on tasks:
> > 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> > 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
> >   Call Trace:
> >    shrink_lruvec
> >    mem_cgroup_iter
> >    shrink_node
> >    do_try_to_free_pages
> >    try_to_free_pages
> >    __alloc_frozen_pages_noprof
> >    alloc_pages_noprof
> >    pte_alloc_one
> >    __pte_alloc
> >    handle_mm_fault
> > 
> > Nothing promises direct reclaim returns in bounded time, and the scan
> > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> > PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> > quiescent state, so the reclaiming task never reports one and becomes a
> > holdout.
> 
> I still don't understand why cond_resched() is being treated as involuntary
> preemption but that is orthogonal to this patch.

The history is that cond_resched() was originally intended to be a
preemption point in any otherwise non-preemptible kernel.  Therefore,
because it is a preemption point, it counts as an involuntary context
switch.

							Thanx, Paul

> > Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
> > state even when cond_resched() does nothing.
> > 
> > PS: This has been discussed in [1]
> > 
> > Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Breno Leitao <leitao@debian.org>
> 
> Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
> 

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
  2026-08-11 17:38   ` Paul E. McKenney
@ 2026-08-11 17:54     ` Shakeel Butt
  2026-08-11 18:39       ` Paul E. McKenney
  0 siblings, 1 reply; 7+ messages in thread
From: Shakeel Butt @ 2026-08-11 17:54 UTC (permalink / raw)
  To: Paul E. McKenney
  Cc: Breno Leitao, Andrew Morton, Johannes Weiner, David Hildenbrand,
	Michal Hocko, Qi Zheng, Lorenzo Stoakes, Kairui Song, Barry Song,
	Axel Rasmussen, Yuanchu Xie, Wei Xu, linux-mm, linux-kernel,
	kernel-team, stable

On Tue, Aug 11, 2026 at 10:38:40AM -0700, Paul E. McKenney wrote:
> On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote:
> > On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> > > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> > > 
> > >   INFO: rcu_tasks detected stalls on tasks:
> > > 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> > > 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
> > >   Call Trace:
> > >    shrink_lruvec
> > >    mem_cgroup_iter
> > >    shrink_node
> > >    do_try_to_free_pages
> > >    try_to_free_pages
> > >    __alloc_frozen_pages_noprof
> > >    alloc_pages_noprof
> > >    pte_alloc_one
> > >    __pte_alloc
> > >    handle_mm_fault
> > > 
> > > Nothing promises direct reclaim returns in bounded time, and the scan
> > > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> > > PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> > > quiescent state, so the reclaiming task never reports one and becomes a
> > > holdout.
> > 
> > I still don't understand why cond_resched() is being treated as involuntary
> > preemption but that is orthogonal to this patch.
> 
> The history is that cond_resched() was originally intended to be a
> preemption point in any otherwise non-preemptible kernel.  Therefore,
> because it is a preemption point, it counts as an involuntary context
> switch.

Thanks for the explanation. Is cond_resched_tasks_rcu_qs() voluntary or
involuntary context switch?


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
  2026-08-11 17:54     ` Shakeel Butt
@ 2026-08-11 18:39       ` Paul E. McKenney
  0 siblings, 0 replies; 7+ messages in thread
From: Paul E. McKenney @ 2026-08-11 18:39 UTC (permalink / raw)
  To: Shakeel Butt
  Cc: Breno Leitao, Andrew Morton, Johannes Weiner, David Hildenbrand,
	Michal Hocko, Qi Zheng, Lorenzo Stoakes, Kairui Song, Barry Song,
	Axel Rasmussen, Yuanchu Xie, Wei Xu, linux-mm, linux-kernel,
	kernel-team, stable

On Tue, Aug 11, 2026 at 10:54:23AM -0700, Shakeel Butt wrote:
> On Tue, Aug 11, 2026 at 10:38:40AM -0700, Paul E. McKenney wrote:
> > On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote:
> > > On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> > > > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> > > > 
> > > >   INFO: rcu_tasks detected stalls on tasks:
> > > > 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> > > > 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
> > > >   Call Trace:
> > > >    shrink_lruvec
> > > >    mem_cgroup_iter
> > > >    shrink_node
> > > >    do_try_to_free_pages
> > > >    try_to_free_pages
> > > >    __alloc_frozen_pages_noprof
> > > >    alloc_pages_noprof
> > > >    pte_alloc_one
> > > >    __pte_alloc
> > > >    handle_mm_fault
> > > > 
> > > > Nothing promises direct reclaim returns in bounded time, and the scan
> > > > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> > > > PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> > > > quiescent state, so the reclaiming task never reports one and becomes a
> > > > holdout.
> > > 
> > > I still don't understand why cond_resched() is being treated as involuntary
> > > preemption but that is orthogonal to this patch.
> > 
> > The history is that cond_resched() was originally intended to be a
> > preemption point in any otherwise non-preemptible kernel.  Therefore,
> > because it is a preemption point, it counts as an involuntary context
> > switch.
> 
> Thanks for the explanation. Is cond_resched_tasks_rcu_qs() voluntary or
> involuntary context switch?

That one is voluntary in order to take care of Tasks RCU's need for a
voluntary context switch.

Except that it doesn't bother involving the scheduler at all, but
instead just updates the current task's state with respect to Tasks RCU.
Much faster that way.  ;-)

							Thanx, Paul

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-08-11 18:39 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-10  9:57 [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec() Breno Leitao
2026-08-10 15:38 ` Paul E. McKenney
2026-08-10 20:02 ` Johannes Weiner
2026-08-11  4:43 ` Shakeel Butt
2026-08-11 17:38   ` Paul E. McKenney
2026-08-11 17:54     ` Shakeel Butt
2026-08-11 18:39       ` Paul E. McKenney

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox