All of lore.kernel.org
 help / color / mirror / Atom feed
From: Rik van Riel <riel@surriel.com>
To: "Paul E. McKenney" <paulmck@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>,
	Qi Zheng <zhengqi.arch@bytedance.com>,
	Peter Zijlstra <peterz@infradead.org>,
	David Hildenbrand <david@redhat.com>,
	kernel test robot <oliver.sang@intel.com>,
	oe-lkp@lists.linux.dev, lkp@intel.com,
	linux-kernel@vger.kernel.org,
	Andrew Morton <akpm@linux-foundation.org>,
	Dave Hansen <dave.hansen@linux.intel.com>,
	Andy Lutomirski <luto@kernel.org>,
	Catalin Marinas <catalin.marinas@arm.com>,
	David Rientjes <rientjes@google.com>,
	Hugh Dickins <hughd@google.com>, Jann Horn <jannh@google.com>,
	Lorenzo Stoakes <lorenzo.stoakes@oracle.com>,
	Mel Gorman <mgorman@suse.de>, Muchun Song <muchun.song@linux.dev>,
	Peter Xu <peterx@redhat.com>, Will Deacon <will@kernel.org>,
	Zach O'Keefe <zokeefe@google.com>,
	Dan Carpenter <dan.carpenter@linaro.org>,
	Frederic Weisbecker <frederic@kernel.org>,
	Neeraj Upadhyay <neeraj.upadhyay@kernel.org>
Subject: Re: [linus:master] [x86] 4817f70c25: stress-ng.mmapaddr.ops_per_sec 63.0% regression
Date: Wed, 29 Jan 2025 11:53:20 -0500	[thread overview]
Message-ID: <20250129115320.1334ad5f@fangorn> (raw)
In-Reply-To: <05da0ae9-073e-4578-b65c-f837c28eead8@paulmck-laptop>

On Wed, 29 Jan 2025 08:36:12 -0800
"Paul E. McKenney" <paulmck@kernel.org> wrote:
> On Wed, Jan 29, 2025 at 11:14:29AM -0500, Rik van Riel wrote:

> > Paul, does this look like it could do the trick,
> > or do we need something else to make RCU freeing
> > happy again?  
> 
> I don't claim to fully understand the issue, but this would prevent
> any RCU grace periods starting subsequently from completing.  It would
> not prevent RCU callbacks from being invoked for RCU grace periods that
> started earlier.
> 
> So it won't prevent RCU callbacks from being invoked.

That makes things clear! I guess we need a different approach.

Qi, does the patch below resolve the regression for you?

---8<---

From 5de4fa686fca15678a7e0a186852f921166854a3 Mon Sep 17 00:00:00 2001
From: Rik van Riel <riel@surriel.com>
Date: Wed, 29 Jan 2025 10:51:51 -0500
Subject: [PATCH 2/2] mm,rcu: prevent RCU callbacks from running with pcp lock
 held

Enabling MMU_GATHER_RCU_TABLE_FREE can create contention on the
zone->lock.  This turns out to be because in some configurations
RCU callbacks are called when IRQs are re-enabled inside
rmqueue_bulk, while the CPU is still holding the per-cpu pages lock.

That results in the RCU callbacks being unable to grab the
PCP lock, and taking the slow path with the zone->lock for
each item freed.

Speed things up by blocking RCU callbacks while holding the
PCP lock.

Signed-off-by: Rik van Riel <riel@surriel.com>
Suggested-by: Paul McKenney <paulmck@kernel.org>
Reported-by: Qi Zheng <zhengqi.arch@bytedance.com>
---
 mm/page_alloc.c | 10 +++++++---
 1 file changed, 7 insertions(+), 3 deletions(-)

diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 6e469c7ef9a4..73e334f403fd 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -94,11 +94,15 @@ static DEFINE_MUTEX(pcp_batch_high_lock);
 
 #if defined(CONFIG_SMP) || defined(CONFIG_PREEMPT_RT)
 /*
- * On SMP, spin_trylock is sufficient protection.
+ * On SMP, spin_trylock is sufficient protection against recursion.
  * On PREEMPT_RT, spin_trylock is equivalent on both SMP and UP.
+ *
+ * Block softirq execution to prevent RCU frees from running in softirq
+ * context while this CPU holds the PCP lock, which could result in a whole
+ * bunch of frees contending on the zone->lock.
  */
-#define pcp_trylock_prepare(flags)	do { } while (0)
-#define pcp_trylock_finish(flag)	do { } while (0)
+#define pcp_trylock_prepare(flags)	local_bh_disable()
+#define pcp_trylock_finish(flag)	local_bh_enable()
 #else
 
 /* UP spin_trylock always succeeds so disable IRQs to prevent re-entrancy. */
-- 
2.47.1


  reply	other threads:[~2025-01-29 16:53 UTC|newest]

Thread overview: 26+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2025-01-28  9:57 [linus:master] [x86] 4817f70c25: stress-ng.mmapaddr.ops_per_sec 63.0% regression kernel test robot
2025-01-28 10:05 ` David Hildenbrand
2025-01-28 11:31   ` Peter Zijlstra
2025-01-28 11:39     ` David Hildenbrand
2025-01-28 13:28       ` Peter Zijlstra
2025-01-28 13:42         ` David Hildenbrand
2025-01-28 15:59           ` Qi Zheng
2025-01-28 17:06           ` Qi Zheng
2025-01-28 17:51             ` Qi Zheng
2025-01-28 18:35             ` Rik van Riel
2025-01-29  8:14               ` Qi Zheng
2025-01-29 15:23                 ` Rik van Riel
2025-01-29 15:59                 ` Rik van Riel
2025-01-29 16:12                   ` Matthew Wilcox
2025-01-29 16:14                     ` Rik van Riel
2025-01-29 16:36                       ` Paul E. McKenney
2025-01-29 16:53                         ` Rik van Riel [this message]
2025-01-29 17:33                           ` Qi Zheng
2025-01-29 17:53                             ` Qi Zheng
2025-01-29 19:19                               ` Paul E. McKenney
2025-01-31 21:11                             ` Rik van Riel
2025-02-01  3:44                               ` Qi Zheng
2025-01-29 16:53                         ` Frederic Weisbecker
2025-01-29 16:57                           ` Rik van Riel
2025-01-29 17:23                             ` Frederic Weisbecker
2025-01-29 17:28                         ` Paul E. McKenney

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20250129115320.1334ad5f@fangorn \
    --to=riel@surriel.com \
    --cc=akpm@linux-foundation.org \
    --cc=catalin.marinas@arm.com \
    --cc=dan.carpenter@linaro.org \
    --cc=dave.hansen@linux.intel.com \
    --cc=david@redhat.com \
    --cc=frederic@kernel.org \
    --cc=hughd@google.com \
    --cc=jannh@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lkp@intel.com \
    --cc=lorenzo.stoakes@oracle.com \
    --cc=luto@kernel.org \
    --cc=mgorman@suse.de \
    --cc=muchun.song@linux.dev \
    --cc=neeraj.upadhyay@kernel.org \
    --cc=oe-lkp@lists.linux.dev \
    --cc=oliver.sang@intel.com \
    --cc=paulmck@kernel.org \
    --cc=peterx@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rientjes@google.com \
    --cc=will@kernel.org \
    --cc=willy@infradead.org \
    --cc=zhengqi.arch@bytedance.com \
    --cc=zokeefe@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.