Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch()
@ 2026-07-27 13:50 Breno Leitao
  2026-07-27 14:24 ` Zi Yan
  0 siblings, 1 reply; 2+ messages in thread
From: Breno Leitao @ 2026-07-27 13:50 UTC (permalink / raw)
  To: Andrew Morton, David Hildenbrand, Zi Yan, Matthew Brost,
	Joshua Hahn, Rakie Kim, Byungchul Park, Gregory Price, Ying Huang,
	Alistair Popple, paulmck
  Cc: linux-mm, linux-kernel, kernel-team, Breno Leitao

migrate_pages_batch() unmaps each folio before moving it, and every
unmap runs the mmu_notifier invalidate callbacks.  On KVM hosts
try_to_migrate() ends up in kvm_mmu_notifier_invalidate_range_start() ->
tdp_mmu_zap_leafs(), which is expensive, so unmapping a large batch keeps
the CPU busy for a long time.

The loop already calls cond_resched(), but on PREEMPTION kernels that is
a no-op, and involuntary preemption is not a Tasks-RCU quiescent state.

A long batch therefore never reports a quiescent state, and the
migrating task (e.g. kcompactd) becomes a Tasks-RCU holdout, stalling the
Tasks-RCU grace period for minutes, which is common at Meta fleet:

  INFO: rcu_tasks detected stalls on tasks:
  0000000055349ecc: .. nvcsw: 1157401/1157401 holdout: 1 idle_cpu: -1/56 task:kcompactd0      state:R  running task
  Call Trace:
   tdp_mmu_zap_leafs
   tdp_mmu_next_root
   gfn_to_pfn_cache_invalidate_start
   kvm_mmu_notifier_invalidate_range_start
   __mmu_notifier_invalidate_range_start
   try_to_migrate_one
   try_to_migrate
   migrate_pages_batch
   migrate_pages
   compact_zone
   compact_node
   kcompactd
   kthread

Use cond_resched_tasks_rcu_qs() so a quiescent state is reported even
when cond_resched() does nothing.

PS: This has also been discussed at [1]

Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
Signed-off-by: Breno Leitao <leitao@debian.org>
---
 mm/migrate.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/mm/migrate.c b/mm/migrate.c
index ab15a4dddd047..b937cbd764808 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1843,7 +1843,7 @@ static int migrate_pages_batch(struct list_head *from,
 			is_thp = folio_test_pmd_mappable(folio);
 			nr_pages = folio_nr_pages(folio);
 
-			cond_resched();
+			cond_resched_tasks_rcu_qs();
 
 			/*
 			 * The rare folio on the deferred split list should

---
base-commit: c5e32e86ca02b003f86e095d379b38148999293d
change-id: 20260727-kcompact-ff0550fc0ebb

Best regards,
--  
Breno Leitao <leitao@debian.org>



^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH] mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch()
  2026-07-27 13:50 [PATCH] mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch() Breno Leitao
@ 2026-07-27 14:24 ` Zi Yan
  0 siblings, 0 replies; 2+ messages in thread
From: Zi Yan @ 2026-07-27 14:24 UTC (permalink / raw)
  To: Breno Leitao
  Cc: Andrew Morton, David Hildenbrand, Matthew Brost, Joshua Hahn,
	Rakie Kim, Byungchul Park, Gregory Price, Ying Huang,
	Alistair Popple, paulmck, linux-mm, linux-kernel, kernel-team

On 27 Jul 2026, at 9:50, Breno Leitao wrote:

> migrate_pages_batch() unmaps each folio before moving it, and every
> unmap runs the mmu_notifier invalidate callbacks.  On KVM hosts
> try_to_migrate() ends up in kvm_mmu_notifier_invalidate_range_start() ->
> tdp_mmu_zap_leafs(), which is expensive, so unmapping a large batch keeps
> the CPU busy for a long time.
>
> The loop already calls cond_resched(), but on PREEMPTION kernels that is
> a no-op, and involuntary preemption is not a Tasks-RCU quiescent state.
>
> A long batch therefore never reports a quiescent state, and the
> migrating task (e.g. kcompactd) becomes a Tasks-RCU holdout, stalling the
> Tasks-RCU grace period for minutes, which is common at Meta fleet:
>
>   INFO: rcu_tasks detected stalls on tasks:
>   0000000055349ecc: .. nvcsw: 1157401/1157401 holdout: 1 idle_cpu: -1/56 task:kcompactd0      state:R  running task
>   Call Trace:
>    tdp_mmu_zap_leafs
>    tdp_mmu_next_root
>    gfn_to_pfn_cache_invalidate_start
>    kvm_mmu_notifier_invalidate_range_start
>    __mmu_notifier_invalidate_range_start
>    try_to_migrate_one
>    try_to_migrate
>    migrate_pages_batch
>    migrate_pages
>    compact_zone
>    compact_node
>    kcompactd
>    kthread
>
> Use cond_resched_tasks_rcu_qs() so a quiescent state is reported even
> when cond_resched() does nothing.

The explanation looks good to me.

Acked-by: Zi Yan <ziy@nvidia.com>

BTW, Sashiko complained about hugetlb path, but hugetlb does not batch
migration, so there should not be an issue like this one.

>
> PS: This has also been discussed at [1]
>
> Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
> Signed-off-by: Breno Leitao <leitao@debian.org>
> ---
>  mm/migrate.c | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/mm/migrate.c b/mm/migrate.c
> index ab15a4dddd047..b937cbd764808 100644
> --- a/mm/migrate.c
> +++ b/mm/migrate.c
> @@ -1843,7 +1843,7 @@ static int migrate_pages_batch(struct list_head *from,
>  			is_thp = folio_test_pmd_mappable(folio);
>  			nr_pages = folio_nr_pages(folio);
>
> -			cond_resched();
> +			cond_resched_tasks_rcu_qs();
>
>  			/*
>  			 * The rare folio on the deferred split list should
>
> ---
> base-commit: c5e32e86ca02b003f86e095d379b38148999293d
> change-id: 20260727-kcompact-ff0550fc0ebb
>
> Best regards,
> --
> Breno Leitao <leitao@debian.org>


Best Regards,
Yan, Zi


^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-07-27 14:25 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-27 13:50 [PATCH] mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch() Breno Leitao
2026-07-27 14:24 ` Zi Yan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox