From: Shakeel Butt <shakeel.butt@linux.dev>
To: Usama Arif <usama.arif@linux.dev>
Cc: "kernel test robot" <yi1.lai@intel.com>,
oe-lkp@lists.linux.dev, lkp@intel.com,
"Andrew Morton" <akpm@linux-foundation.org>,
"David Hildenbrand" <david@kernel.org>,
"Johannes Weiner" <hannes@cmpxchg.org>,
"Liam R. Howlett" <liam@infradead.org>,
"Lorenzo Stoakes" <ljs@kernel.org>,
"Michal Hocko" <mhocko@suse.com>,
"Michal Koutný" <mkoutny@suse.com>,
"Mike Rapoport" <rppt@kernel.org>,
"Roman Gushchin" <roman.gushchin@linux.dev>,
"Suren Baghdasaryan" <surenb@google.com>,
"Tejun Heo" <tj@kernel.org>,
"Vlastimil Babka" <vbabka@kernel.org>,
linux-mm@kvack.org, cgroups@vger.kernel.org,
oliver.sang@intel.com
Subject: Re: [linux-next:master] [mm/vmpressure] ea928e9e18: stress-ng.mremap.ops_per_sec 36.2% regression
Date: Fri, 14 Aug 2026 17:25:39 -0700 [thread overview]
Message-ID: <an-w9UJy6y9Wg11J@linux.dev> (raw)
In-Reply-To: <f330dd75-d9b8-4259-a02f-9e3ff4bcdc1e@linux.dev>
On Fri, Aug 14, 2026 at 11:32:09AM +0100, Usama Arif wrote:
>
>
> On 13/08/2026 18:29, Shakeel Butt wrote:
> > On Thu, Aug 13, 2026 at 06:14:50PM +0100, Usama Arif wrote:
> >>
> >>
> >> On 13/08/2026 18:01, Shakeel Butt wrote:
> >>> On Thu, Aug 13, 2026 at 09:16:22PM +0800, kernel test robot wrote:
> >>>>
> >>>>
> >>>> Hello,
> >>>>
> >>>> kernel test robot noticed a 36.2% regression of stress-ng.mremap.ops_per_sec on:
> >>>>
> >>>> commit: ea928e9e18da682e9a5bc40aa862bff7ce5ae42e ("mm/vmpressure: move v1 userspace eventfd code into memcontrol-v1.c") https://git.kernel.org/cgit/linux/kernel/git/next/linux-next.git master
> >>>>
> >>>> in testcase: stress-ng
> >>>> version: stress-ng-x86_64-29ce10a2c-1_20260712
> >>>> with following parameters:
> >>>>
> >>>> nr_threads: 100%
> >>>> testtime: 60s
> >>>> test: mremap
> >>>> cpufreq_governor: performance
> >>>>
> >>>>
> >>>>
> >>>> config: x86_64-rhel-9.4 (CONFIG_MEMCG=y and CONFIG_MEMCG_V1 is not set)
> >>>> compiler: gcc-14
> >>>> test machine: 256 threads 4 sockets INTEL(R) XEON(R) PLATINUM 8592+ (Emerald Rapids) with 256G memory
> >>>>
> >>>> (please refer to attached dmesg/kmsg for entire log/backtrace)
> >>>>
> >>>
> >>> Hi there,
> >>>
> >>> Can you please test the following patch and see if it fixes the regression?
> >>>
> >>>
> >>> From 84c0b05b3bc5cf73ee66ead75aafb1ad684462c3 Mon Sep 17 00:00:00 2001
> >>> From: Shakeel Butt <shakeel.butt@linux.dev>
> >>> Date: Thu, 13 Aug 2026 09:38:28 -0700
> >>> Subject: [PATCH] memcg: keep vmstats_percpu off the memory_events[] cacheline
> >>>
> >>> Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
> >>> ---
> >>> include/linux/memcontrol.h | 10 ++++++----
> >>> 1 file changed, 6 insertions(+), 4 deletions(-)
> >>>
> >>> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
> >>> index e78bc98ab229..e25d5b9a1db8 100644
> >>> --- a/include/linux/memcontrol.h
> >>> +++ b/include/linux/memcontrol.h
> >>> @@ -246,8 +246,13 @@ struct mem_cgroup {
> >>> /* handle for "memory.swap.events" */
> >>> struct cgroup_file swap_events_file;
> >>>
> >>> - /* memory.stat */
> >>> + /* Read-mostly. */
> >>> struct memcg_vmstats *vmstats;
> >>> + struct memcg_vmstats_percpu __percpu *vmstats_percpu;
> >>> + int kmemcg_id;
> >>> +
> >>> + /* Write-hot from here on; do not let it share with the above. */
> >>> + CACHELINE_PADDING(_pad_);
> >>>
> >>> /* memory.events */
> >>> atomic_long_t memory_events[MEMCG_NR_MEMORY_EVENTS];
> >>> @@ -266,9 +271,6 @@ struct mem_cgroup {
> >>> #if BITS_PER_LONG < 64
> >>> seqlock_t socket_pressure_seqlock;
> >>> #endif
> >>> - int kmemcg_id;
> >>> -
> >>> - struct memcg_vmstats_percpu __percpu *vmstats_percpu;
> >>>
> >>> #ifdef CONFIG_CGROUP_WRITEBACK
> >>> struct list_head cgwb_list;
> >>
> >>
> >> I was currently testing this diff, not sure which one would be better.
> >
> > I was just checking if false sharing of vmstats_percpu is the cause. If your
> > patch does not increase the struct size, we can go with that as a backportable
> > fix. I am planning to rearrange fields of struct mem_cgroup more drastically and
> > have it more stable as future work as we continuously see these regressions keep
> > popping up.
> >
>
>
> Yes this makes sense. I did not expect such a big change in a benchmark
> with my patch, although I feel like the microbenchmark is probably not
> that realistic.
>
> I think another issue is that its a 4 socket system.
> I only have access to a single socket system, and I see a 4.38% regression.
>
> Do you know if there a way for kernel test robot to test the below patch
> on its host?
+Oliver
I used to request Oliver to test a patch for performance regressions reported by
intel test robot.
>
>
> From b862e84e7bd6a54b1546b7f21a6cf991253def18 Mon Sep 17 00:00:00 2001
> From: Usama Arif <usama.arif@linux.dev>
> Date: Thu, 13 Aug 2026 11:42:05 -0700
> Subject: [PATCH] mm/memcontrol: avoid false sharing between vmstats and events
>
> Moving v1 userspace eventfd handling into memcontrol-v1.c shrank
> struct vmpressure from 112 to 24 bytes when CONFIG_MEMCG_V1 is disabled.
> This moved memory_events_local[MEMCG_SWAP_FAIL] and the hot
> vmstats_percpu pointer onto the same cacheline.
>
> The stress-ng mremap stressor exercises MADV_PAGEOUT with swap
> disabled, generating about 20 million MEMCG_SWAP_FAIL updates per
> 60-second run on a 176-CPU test system. Those writes bounce the line
> while memcg statistics paths load vmstats_percpu.
>
> Move cgwb_list into the existing alignment gap and cacheline-align
> vmstats_percpu. This separates the pointer from the event counters
> without increasing the size of struct mem_cgroup in the tested
> configuration.
>
> The blamed commit reduced median mremap throughput by 4.38% on the
> test system with one socket. The patched kernel brings the performance
> to within 0.5% of the parent which is within the observed boot-to-boot
> spread (up to 1.2%).
>
> Fixes: ea928e9e18da ("mm/vmpressure: move v1 userspace eventfd code into memcontrol-v1.c")
> Reported-by: kernel test robot <yi1.lai@intel.com>
> Closes: https://lore.kernel.org/oe-lkp/202608131743.c6a7dda4-lkp@intel.com
> Signed-off-by: Usama Arif <usama.arif@linux.dev>
> ---
> include/linux/memcontrol.h | 9 +++++++--
> 1 file changed, 7 insertions(+), 2 deletions(-)
>
> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
> index e78bc98ab229b..215e2e87f42b2 100644
> --- a/include/linux/memcontrol.h
> +++ b/include/linux/memcontrol.h
> @@ -268,10 +268,15 @@ struct mem_cgroup {
> #endif
> int kmemcg_id;
>
> - struct memcg_vmstats_percpu __percpu *vmstats_percpu;
> -
> #ifdef CONFIG_CGROUP_WRITEBACK
> struct list_head cgwb_list;
> +#endif
> +
> + /* Keep the hot per-CPU stats pointer away from memory event counters. */
> + struct memcg_vmstats_percpu __percpu *vmstats_percpu
> + ____cacheline_aligned_in_smp;
> +
> +#ifdef CONFIG_CGROUP_WRITEBACK
> struct wb_domain cgwb_domain;
> struct memcg_cgwb_frn cgwb_frn[MEMCG_CGWB_FRN_CNT];
> #endif
> --
> 2.53.0-Meta
>
>
>
> >>
> >> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
> >> index e78bc98ab229b..215e2e87f42b2 100644
> >> --- a/include/linux/memcontrol.h
> >> +++ b/include/linux/memcontrol.h
> >> @@ -268,10 +268,15 @@ struct mem_cgroup {
> >> #endif
> >> int kmemcg_id;
> >>
> >> - struct memcg_vmstats_percpu __percpu *vmstats_percpu;
> >> -
> >> #ifdef CONFIG_CGROUP_WRITEBACK
> >> struct list_head cgwb_list;
> >> +#endif
> >> +
> >> + /* Keep the hot per-CPU stats pointer away from memory event counters. */
> >> + struct memcg_vmstats_percpu __percpu *vmstats_percpu
> >> + ____cacheline_aligned_in_smp;
> >> +
> >> +#ifdef CONFIG_CGROUP_WRITEBACK
> >> struct wb_domain cgwb_domain;
> >> struct memcg_cgwb_frn cgwb_frn[MEMCG_CGWB_FRN_CNT];
> >> #endif
> >>
>
prev parent reply other threads:[~2026-08-15 0:25 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-13 13:16 [linux-next:master] [mm/vmpressure] ea928e9e18: stress-ng.mremap.ops_per_sec 36.2% regression kernel test robot
2026-08-13 17:01 ` Shakeel Butt
2026-08-13 17:14 ` Usama Arif
[not found] ` <an3-aksL0MZPM3ZJ@linux.dev>
2026-08-14 10:32 ` Usama Arif
2026-08-15 0:25 ` Shakeel Butt [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=an-w9UJy6y9Wg11J@linux.dev \
--to=shakeel.butt@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=cgroups@vger.kernel.org \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=liam@infradead.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=lkp@intel.com \
--cc=mhocko@suse.com \
--cc=mkoutny@suse.com \
--cc=oe-lkp@lists.linux.dev \
--cc=oliver.sang@intel.com \
--cc=roman.gushchin@linux.dev \
--cc=rppt@kernel.org \
--cc=surenb@google.com \
--cc=tj@kernel.org \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=yi1.lai@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox