* [linux-next:master] [mm/vmpressure] ea928e9e18: stress-ng.mremap.ops_per_sec 36.2% regression
@ 2026-08-13 13:16 kernel test robot
2026-08-13 17:01 ` Shakeel Butt
0 siblings, 1 reply; 4+ messages in thread
From: kernel test robot @ 2026-08-13 13:16 UTC (permalink / raw)
To: Usama Arif
Cc: oe-lkp, lkp, Andrew Morton, Shakeel Butt, David Hildenbrand,
Johannes Weiner, Liam R. Howlett, Lorenzo Stoakes, Michal Hocko,
Michal Koutný, Mike Rapoport, Roman Gushchin,
Suren Baghdasaryan, Tejun Heo, Vlastimil Babka, linux-mm, cgroups,
yi1.lai
Hello,
kernel test robot noticed a 36.2% regression of stress-ng.mremap.ops_per_sec on:
commit: ea928e9e18da682e9a5bc40aa862bff7ce5ae42e ("mm/vmpressure: move v1 userspace eventfd code into memcontrol-v1.c") https://git.kernel.org/cgit/linux/kernel/git/next/linux-next.git master
in testcase: stress-ng
version: stress-ng-x86_64-29ce10a2c-1_20260712
with following parameters:
nr_threads: 100%
testtime: 60s
test: mremap
cpufreq_governor: performance
config: x86_64-rhel-9.4 (CONFIG_MEMCG=y and CONFIG_MEMCG_V1 is not set)
compiler: gcc-14
test machine: 256 threads 4 sockets INTEL(R) XEON(R) PLATINUM 8592+ (Emerald Rapids) with 256G memory
(please refer to attached dmesg/kmsg for entire log/backtrace)
If you fix the issue in a separate patch/commit (i.e. not just a new version of the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <yi1.lai@intel.com>
| Closes: https://lore.kernel.org/oe-lkp/202608131743.c6a7dda4-lkp@intel.com
2026-08-11 09:13:23 sysctl -w kernel.sched_autogroup_enabled=0 kernel.sched_autogroup_enabled = 0
2026-08-11 09:13:23 stress-ng --timeout 60 --times --verify --metrics --no-rand-seed --mremap 256
stress-ng: info: [10886] setting to a 1 min run per stressor
stress-ng: info: [10886] dispatching hogs: 256 mremap
stress-ng: info: [11033] mremap: using 1MB per stressor instance (total 256MB of 243.06GB available memory)
stress-ng: metrc: [10886] stressor bogo ops real time usr time sys time bogo ops/s bogo ops/s CPU used per RSS Max
stress-ng: metrc: [10886] (secs) (secs) (secs) (real time) (usr+sys time) instance (%) (KB)
stress-ng: metrc: [10886] mremap 2212529 59.96 209.34 15014.92 36901.22 145.33 99.19 6684
stress-ng: metrc: [10886] miscellaneous metrics:
stress-ng: metrc: [10886] mremap 151372.07 nanosecs per mremap call (harmonic mean of 256 instances)
stress-ng: info: [10886] for a 60.01s run time:
stress-ng: info: [10886] 15363.58s available CPU time
stress-ng: info: [10886] 209.34s user time ( 1.36%)
stress-ng: info: [10886] 15015.19s system time ( 97.73%)
stress-ng: info: [10886] 15224.53s total time ( 99.09%)
stress-ng: info: [10886] load average: 163.54 47.28 16.30
stress-ng: info: [10886] skipped: 0
stress-ng: info: [10886] passed: 256: mremap (256)
stress-ng: info: [10886] failed: 0
stress-ng: info: [10886] metrics untrustworthy: 0
stress-ng: info: [10886] successful run completed in 1 min
The kernel config and materials to reproduce are available at:
https://download.01.org/0day-ci/archive/20260813/202608131743.c6a7dda4-lkp@intel.com
--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [linux-next:master] [mm/vmpressure] ea928e9e18: stress-ng.mremap.ops_per_sec 36.2% regression
2026-08-13 13:16 [linux-next:master] [mm/vmpressure] ea928e9e18: stress-ng.mremap.ops_per_sec 36.2% regression kernel test robot
@ 2026-08-13 17:01 ` Shakeel Butt
2026-08-13 17:14 ` Usama Arif
0 siblings, 1 reply; 4+ messages in thread
From: Shakeel Butt @ 2026-08-13 17:01 UTC (permalink / raw)
To: kernel test robot
Cc: Usama Arif, oe-lkp, lkp, Andrew Morton, David Hildenbrand,
Johannes Weiner, Liam R. Howlett, Lorenzo Stoakes, Michal Hocko,
Michal Koutný, Mike Rapoport, Roman Gushchin,
Suren Baghdasaryan, Tejun Heo, Vlastimil Babka, linux-mm, cgroups
On Thu, Aug 13, 2026 at 09:16:22PM +0800, kernel test robot wrote:
>
>
> Hello,
>
> kernel test robot noticed a 36.2% regression of stress-ng.mremap.ops_per_sec on:
>
> commit: ea928e9e18da682e9a5bc40aa862bff7ce5ae42e ("mm/vmpressure: move v1 userspace eventfd code into memcontrol-v1.c") https://git.kernel.org/cgit/linux/kernel/git/next/linux-next.git master
>
> in testcase: stress-ng
> version: stress-ng-x86_64-29ce10a2c-1_20260712
> with following parameters:
>
> nr_threads: 100%
> testtime: 60s
> test: mremap
> cpufreq_governor: performance
>
>
>
> config: x86_64-rhel-9.4 (CONFIG_MEMCG=y and CONFIG_MEMCG_V1 is not set)
> compiler: gcc-14
> test machine: 256 threads 4 sockets INTEL(R) XEON(R) PLATINUM 8592+ (Emerald Rapids) with 256G memory
>
> (please refer to attached dmesg/kmsg for entire log/backtrace)
>
Hi there,
Can you please test the following patch and see if it fixes the regression?
From 84c0b05b3bc5cf73ee66ead75aafb1ad684462c3 Mon Sep 17 00:00:00 2001
From: Shakeel Butt <shakeel.butt@linux.dev>
Date: Thu, 13 Aug 2026 09:38:28 -0700
Subject: [PATCH] memcg: keep vmstats_percpu off the memory_events[] cacheline
Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
---
include/linux/memcontrol.h | 10 ++++++----
1 file changed, 6 insertions(+), 4 deletions(-)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index e78bc98ab229..e25d5b9a1db8 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -246,8 +246,13 @@ struct mem_cgroup {
/* handle for "memory.swap.events" */
struct cgroup_file swap_events_file;
- /* memory.stat */
+ /* Read-mostly. */
struct memcg_vmstats *vmstats;
+ struct memcg_vmstats_percpu __percpu *vmstats_percpu;
+ int kmemcg_id;
+
+ /* Write-hot from here on; do not let it share with the above. */
+ CACHELINE_PADDING(_pad_);
/* memory.events */
atomic_long_t memory_events[MEMCG_NR_MEMORY_EVENTS];
@@ -266,9 +271,6 @@ struct mem_cgroup {
#if BITS_PER_LONG < 64
seqlock_t socket_pressure_seqlock;
#endif
- int kmemcg_id;
-
- struct memcg_vmstats_percpu __percpu *vmstats_percpu;
#ifdef CONFIG_CGROUP_WRITEBACK
struct list_head cgwb_list;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [linux-next:master] [mm/vmpressure] ea928e9e18: stress-ng.mremap.ops_per_sec 36.2% regression
2026-08-13 17:01 ` Shakeel Butt
@ 2026-08-13 17:14 ` Usama Arif
2026-08-13 17:29 ` Shakeel Butt
0 siblings, 1 reply; 4+ messages in thread
From: Usama Arif @ 2026-08-13 17:14 UTC (permalink / raw)
To: Shakeel Butt, kernel test robot
Cc: oe-lkp, lkp, Andrew Morton, David Hildenbrand, Johannes Weiner,
Liam R. Howlett, Lorenzo Stoakes, Michal Hocko,
Michal Koutný, Mike Rapoport, Roman Gushchin,
Suren Baghdasaryan, Tejun Heo, Vlastimil Babka, linux-mm, cgroups
On 13/08/2026 18:01, Shakeel Butt wrote:
> On Thu, Aug 13, 2026 at 09:16:22PM +0800, kernel test robot wrote:
>>
>>
>> Hello,
>>
>> kernel test robot noticed a 36.2% regression of stress-ng.mremap.ops_per_sec on:
>>
>> commit: ea928e9e18da682e9a5bc40aa862bff7ce5ae42e ("mm/vmpressure: move v1 userspace eventfd code into memcontrol-v1.c") https://git.kernel.org/cgit/linux/kernel/git/next/linux-next.git master
>>
>> in testcase: stress-ng
>> version: stress-ng-x86_64-29ce10a2c-1_20260712
>> with following parameters:
>>
>> nr_threads: 100%
>> testtime: 60s
>> test: mremap
>> cpufreq_governor: performance
>>
>>
>>
>> config: x86_64-rhel-9.4 (CONFIG_MEMCG=y and CONFIG_MEMCG_V1 is not set)
>> compiler: gcc-14
>> test machine: 256 threads 4 sockets INTEL(R) XEON(R) PLATINUM 8592+ (Emerald Rapids) with 256G memory
>>
>> (please refer to attached dmesg/kmsg for entire log/backtrace)
>>
>
> Hi there,
>
> Can you please test the following patch and see if it fixes the regression?
>
>
> From 84c0b05b3bc5cf73ee66ead75aafb1ad684462c3 Mon Sep 17 00:00:00 2001
> From: Shakeel Butt <shakeel.butt@linux.dev>
> Date: Thu, 13 Aug 2026 09:38:28 -0700
> Subject: [PATCH] memcg: keep vmstats_percpu off the memory_events[] cacheline
>
> Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
> ---
> include/linux/memcontrol.h | 10 ++++++----
> 1 file changed, 6 insertions(+), 4 deletions(-)
>
> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
> index e78bc98ab229..e25d5b9a1db8 100644
> --- a/include/linux/memcontrol.h
> +++ b/include/linux/memcontrol.h
> @@ -246,8 +246,13 @@ struct mem_cgroup {
> /* handle for "memory.swap.events" */
> struct cgroup_file swap_events_file;
>
> - /* memory.stat */
> + /* Read-mostly. */
> struct memcg_vmstats *vmstats;
> + struct memcg_vmstats_percpu __percpu *vmstats_percpu;
> + int kmemcg_id;
> +
> + /* Write-hot from here on; do not let it share with the above. */
> + CACHELINE_PADDING(_pad_);
>
> /* memory.events */
> atomic_long_t memory_events[MEMCG_NR_MEMORY_EVENTS];
> @@ -266,9 +271,6 @@ struct mem_cgroup {
> #if BITS_PER_LONG < 64
> seqlock_t socket_pressure_seqlock;
> #endif
> - int kmemcg_id;
> -
> - struct memcg_vmstats_percpu __percpu *vmstats_percpu;
>
> #ifdef CONFIG_CGROUP_WRITEBACK
> struct list_head cgwb_list;
I was currently testing this diff, not sure which one would be better.
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index e78bc98ab229b..215e2e87f42b2 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -268,10 +268,15 @@ struct mem_cgroup {
#endif
int kmemcg_id;
- struct memcg_vmstats_percpu __percpu *vmstats_percpu;
-
#ifdef CONFIG_CGROUP_WRITEBACK
struct list_head cgwb_list;
+#endif
+
+ /* Keep the hot per-CPU stats pointer away from memory event counters. */
+ struct memcg_vmstats_percpu __percpu *vmstats_percpu
+ ____cacheline_aligned_in_smp;
+
+#ifdef CONFIG_CGROUP_WRITEBACK
struct wb_domain cgwb_domain;
struct memcg_cgwb_frn cgwb_frn[MEMCG_CGWB_FRN_CNT];
#endif
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [linux-next:master] [mm/vmpressure] ea928e9e18: stress-ng.mremap.ops_per_sec 36.2% regression
2026-08-13 17:14 ` Usama Arif
@ 2026-08-13 17:29 ` Shakeel Butt
0 siblings, 0 replies; 4+ messages in thread
From: Shakeel Butt @ 2026-08-13 17:29 UTC (permalink / raw)
To: Usama Arif
Cc: kernel test robot, oe-lkp, lkp, Andrew Morton, David Hildenbrand,
Johannes Weiner, Liam R. Howlett, Lorenzo Stoakes, Michal Hocko,
Michal Koutný, Mike Rapoport, Roman Gushchin,
Suren Baghdasaryan, Tejun Heo, Vlastimil Babka, linux-mm, cgroups
On Thu, Aug 13, 2026 at 06:14:50PM +0100, Usama Arif wrote:
>
>
> On 13/08/2026 18:01, Shakeel Butt wrote:
> > On Thu, Aug 13, 2026 at 09:16:22PM +0800, kernel test robot wrote:
> >>
> >>
> >> Hello,
> >>
> >> kernel test robot noticed a 36.2% regression of stress-ng.mremap.ops_per_sec on:
> >>
> >> commit: ea928e9e18da682e9a5bc40aa862bff7ce5ae42e ("mm/vmpressure: move v1 userspace eventfd code into memcontrol-v1.c") https://git.kernel.org/cgit/linux/kernel/git/next/linux-next.git master
> >>
> >> in testcase: stress-ng
> >> version: stress-ng-x86_64-29ce10a2c-1_20260712
> >> with following parameters:
> >>
> >> nr_threads: 100%
> >> testtime: 60s
> >> test: mremap
> >> cpufreq_governor: performance
> >>
> >>
> >>
> >> config: x86_64-rhel-9.4 (CONFIG_MEMCG=y and CONFIG_MEMCG_V1 is not set)
> >> compiler: gcc-14
> >> test machine: 256 threads 4 sockets INTEL(R) XEON(R) PLATINUM 8592+ (Emerald Rapids) with 256G memory
> >>
> >> (please refer to attached dmesg/kmsg for entire log/backtrace)
> >>
> >
> > Hi there,
> >
> > Can you please test the following patch and see if it fixes the regression?
> >
> >
> > From 84c0b05b3bc5cf73ee66ead75aafb1ad684462c3 Mon Sep 17 00:00:00 2001
> > From: Shakeel Butt <shakeel.butt@linux.dev>
> > Date: Thu, 13 Aug 2026 09:38:28 -0700
> > Subject: [PATCH] memcg: keep vmstats_percpu off the memory_events[] cacheline
> >
> > Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
> > ---
> > include/linux/memcontrol.h | 10 ++++++----
> > 1 file changed, 6 insertions(+), 4 deletions(-)
> >
> > diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
> > index e78bc98ab229..e25d5b9a1db8 100644
> > --- a/include/linux/memcontrol.h
> > +++ b/include/linux/memcontrol.h
> > @@ -246,8 +246,13 @@ struct mem_cgroup {
> > /* handle for "memory.swap.events" */
> > struct cgroup_file swap_events_file;
> >
> > - /* memory.stat */
> > + /* Read-mostly. */
> > struct memcg_vmstats *vmstats;
> > + struct memcg_vmstats_percpu __percpu *vmstats_percpu;
> > + int kmemcg_id;
> > +
> > + /* Write-hot from here on; do not let it share with the above. */
> > + CACHELINE_PADDING(_pad_);
> >
> > /* memory.events */
> > atomic_long_t memory_events[MEMCG_NR_MEMORY_EVENTS];
> > @@ -266,9 +271,6 @@ struct mem_cgroup {
> > #if BITS_PER_LONG < 64
> > seqlock_t socket_pressure_seqlock;
> > #endif
> > - int kmemcg_id;
> > -
> > - struct memcg_vmstats_percpu __percpu *vmstats_percpu;
> >
> > #ifdef CONFIG_CGROUP_WRITEBACK
> > struct list_head cgwb_list;
>
>
> I was currently testing this diff, not sure which one would be better.
I was just checking if false sharing of vmstats_percpu is the cause. If your
patch does not increase the struct size, we can go with that as a backportable
fix. I am planning to rearrange fields of struct mem_cgroup more drastically and
have it more stable as future work as we continuously see these regressions keep
popping up.
>
> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
> index e78bc98ab229b..215e2e87f42b2 100644
> --- a/include/linux/memcontrol.h
> +++ b/include/linux/memcontrol.h
> @@ -268,10 +268,15 @@ struct mem_cgroup {
> #endif
> int kmemcg_id;
>
> - struct memcg_vmstats_percpu __percpu *vmstats_percpu;
> -
> #ifdef CONFIG_CGROUP_WRITEBACK
> struct list_head cgwb_list;
> +#endif
> +
> + /* Keep the hot per-CPU stats pointer away from memory event counters. */
> + struct memcg_vmstats_percpu __percpu *vmstats_percpu
> + ____cacheline_aligned_in_smp;
> +
> +#ifdef CONFIG_CGROUP_WRITEBACK
> struct wb_domain cgwb_domain;
> struct memcg_cgwb_frn cgwb_frn[MEMCG_CGWB_FRN_CNT];
> #endif
>
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-08-13 17:30 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-13 13:16 [linux-next:master] [mm/vmpressure] ea928e9e18: stress-ng.mremap.ops_per_sec 36.2% regression kernel test robot
2026-08-13 17:01 ` Shakeel Butt
2026-08-13 17:14 ` Usama Arif
2026-08-13 17:29 ` Shakeel Butt
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.