From: Andrew Morton <akpm@linux-foundation.org>
To: mm-commits@vger.kernel.org,shakeel.butt@linux.dev,roman.gushchin@linux.dev,muchun.song@linux.dev,mhocko@kernel.org,hannes@cmpxchg.org,joannelkoong@gmail.com,akpm@linux-foundation.org
Subject: + mm-memcontrol-skip-non-hierarchical-memcg-wide-stats-when-v1-is-unavailable.patch added to mm-new branch
Date: Sat, 05 Sep 2026 17:12:38 -0700 [thread overview]
Message-ID: <20260906001239.4D33F1F00A3A@smtp.kernel.org> (raw)
The patch titled
Subject: mm/memcontrol: skip non-hierarchical memcg-wide stats when v1 is unavailable
has been added to the -mm mm-new branch. Its filename is
mm-memcontrol-skip-non-hierarchical-memcg-wide-stats-when-v1-is-unavailable.patch
This patch will shortly appear at
https://git.kernel.org/pub/scm/linux/kernel/git/akpm/25-new.git/tree/patches/mm-memcontrol-skip-non-hierarchical-memcg-wide-stats-when-v1-is-unavailable.patch
This patch will later appear in the mm-new branch at
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Note, mm-new is a provisional staging ground for work-in-progress
patches, and acceptance into mm-new is a notification for others take
notice and to finish up reviews. Please do not hesitate to respond to
review feedback and post updated versions to replace or incrementally
fixup patches in mm-new.
The mm-new branch of mm.git is not included in linux-next
If a few days of testing in mm-new is successful, the patch will me moved
into mm.git's mm-unstable branch, which is included in linux-next
Before you just go and hit "reply", please:
a) Consider who else should be cc'ed
b) Prefer to cc a suitable mailing list as well
c) Ideally: find the original patch on the mailing list and do a
reply-to-all to that, adding suitable additional cc's
*** Remember to use Documentation/process/submit-checklist.rst when testing your code ***
The -mm tree is included into linux-next via various
branches at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
and is updated there most days
------------------------------------------------------
From: Joanne Koong <joannelkoong@gmail.com>
Subject: mm/memcontrol: skip non-hierarchical memcg-wide stats when v1 is unavailable
Date: Thu, 3 Sep 2026 14:56:16 -0700
memcg_vmstats keeps a non-hierarchical copy of every memcg-wide stat item
and event alongside the hierarchical one. The only readers however are
the legacy memory.stat and memory.numa_stat, and reparenting on offline.
All of them live under CONFIG_MEMCG_V1, and their accessors
(memcg_page_state_local() and memcg_events_local()) are already compiled
out with it. This means on a CONFIG_MEMCG_V1=n kernel, memcg_vmstats's
non-hierarchical arrays are written to on every rstat flush, despite their
values never being read / accessed.
The same holds when the kernel does support v1 but the controller has been
blocked from v1 hierarchies with the boot param cgroup_no_v1={memory,all}.
A v1 mount is refused in that case, so the legacy memory.stat can never
exist and the arrays are just as unread / unaccessed.
Compile out the non-hierarchical memcg-wide arrays if CONFIG_MEMCG_V1 is
not set. If it is set but cgroup_no_v1= has blocked the controller, skip
updates on the arrays.
This makes flushes cheaper. mem_cgroup_stat_aggregate() can now skip the
read-modify-write of ac->local[i]. Nothing else accesses state_local or
events_local, so those cachelines were getting pulled in solely for the
writes and they are separate from the ones the loop is already walking.
On an 80-cpu x86_64 machine with 500 cgroups each running a workload that
dirties anon, file, dirty/writeback, slab, kmem, mlock, and reclaim
counters, timing mem_cgroup_css_rstat_flush() in-kernel in TSC ticks per
flush showed roughly
before after delta
memcg-wide aggregation 1231 1180 -4.1%
overall flush function 2452 2397 -2.2%
These numbers are from taking the median of 70 samples, one per 20s window
on each kernel. The 95% intervals observed on the two deltas are [-5.41%,
-1.76%] and [-4.53%, -0.14%]. The values above include the timing
overhead itself, so only the delta is meaningful here.
Counting the items that actually changed, a median of 1.5 of the 77
memcg-wide items (57 state + 20 events) had a non-zero per-cpu delta at
each flush, which means the benchmarks above are with one or two fewer
cachelines pulled in per flush. The count is low because the benchmark
reads memory.stat in a loop to keep the flush rate up. For cases where
flushes are triggered only by the 2s periodic worker, more changes will
have accumulated between flushes, so more cachelines are skipped and the
per-flush saving should be larger.
Link: https://lore.kernel.org/20260903215616.1456239-1-joannelkoong@gmail.com
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
---
include/linux/cgroup.h | 3 +
kernel/cgroup/cgroup-internal.h | 1
mm/memcontrol.c | 48 +++++++++++++++++++++++++-----
3 files changed, 44 insertions(+), 8 deletions(-)
--- a/include/linux/cgroup.h~mm-memcontrol-skip-non-hierarchical-memcg-wide-stats-when-v1-is-unavailable
+++ a/include/linux/cgroup.h
@@ -154,6 +154,9 @@ struct cgroup *cgroup_get_from_path(cons
struct cgroup *cgroup_get_from_fd(int fd);
struct cgroup *cgroup_v1v2_get_from_fd(int fd);
+/* Was this controller blocked from v1 hierarchies by cgroup_no_v1= ? */
+bool cgroup1_ssid_disabled(int ssid);
+
int cgroup_attach_task_all(struct task_struct *from, struct task_struct *);
int cgroup_transfer_tasks(struct cgroup *to, struct cgroup *from);
--- a/kernel/cgroup/cgroup-internal.h~mm-memcontrol-skip-non-hierarchical-memcg-wide-stats-when-v1-is-unavailable
+++ a/kernel/cgroup/cgroup-internal.h
@@ -285,7 +285,6 @@ extern struct kernfs_syscall_ops cgroup1
extern const struct fs_parameter_spec cgroup1_fs_parameters[];
int proc_cgroupstats_show(struct seq_file *m, void *v);
-bool cgroup1_ssid_disabled(int ssid);
void cgroup1_pidlist_destroy_all(struct cgroup *cgrp);
void cgroup1_release_agent(struct work_struct *work);
void cgroup1_check_for_release(struct cgroup *cgrp);
--- a/mm/memcontrol.c~mm-memcontrol-skip-non-hierarchical-memcg-wide-stats-when-v1-is-unavailable
+++ a/mm/memcontrol.c
@@ -676,9 +676,11 @@ struct memcg_vmstats {
long state[MEMCG_VMSTAT_SIZE];
unsigned long events[NR_MEMCG_EVENTS];
+#ifdef CONFIG_MEMCG_V1
/* Non-hierarchical (CPU aggregated) page state & events */
long state_local[MEMCG_VMSTAT_SIZE];
unsigned long events_local[NR_MEMCG_EVENTS];
+#endif
/* Pending child counts during tree propagation */
long state_pending[MEMCG_VMSTAT_SIZE];
@@ -689,6 +691,31 @@ struct memcg_vmstats {
};
/*
+ * The non-hierarchical memcg-wide counters are read back only by the legacy
+ * memory.stat and memory.numa_stat, and by reparenting on offline, all of which
+ * are v1-only. If the kernel is built without CONFIG_MEMCG_V1, or if the boot
+ * param cgroup_no_v1= has blocked the memory controller from v1 hierarchies,
+ * then nothing reads them and writers can skip the updates.
+ */
+static long *memcg_state_local_array(struct mem_cgroup *memcg)
+{
+#ifdef CONFIG_MEMCG_V1
+ if (!cgroup1_ssid_disabled(memory_cgrp_id))
+ return memcg->vmstats->state_local;
+#endif
+ return NULL;
+}
+
+static unsigned long *memcg_events_local_array(struct mem_cgroup *memcg)
+{
+#ifdef CONFIG_MEMCG_V1
+ if (!cgroup1_ssid_disabled(memory_cgrp_id))
+ return memcg->vmstats->events_local;
+#endif
+ return NULL;
+}
+
+/*
* memcg and lruvec stats flushing
*
* Many codepaths leading to stats update or read are performance sensitive and
@@ -4469,7 +4496,10 @@ static void mem_cgroup_css_reset(struct
struct aggregate_control {
/* pointer to the aggregated (CPU and subtree aggregated) counters */
long *aggregate;
- /* pointer to the non-hierarchichal (CPU aggregated) counters */
+ /*
+ * pointer to the non-hierarchical (CPU aggregated) counters or NULL to
+ * skip updating them (see memcg_state_local_array())
+ */
long *local;
/* pointer to the pending child counters during tree propagation */
long *pending;
@@ -4508,7 +4538,7 @@ static void mem_cgroup_stat_aggregate(st
}
/* Aggregate counts on this level and propagate upwards */
- if (delta_cpu)
+ if (delta_cpu && ac->local)
ac->local[i] += delta_cpu;
if (delta) {
@@ -4522,6 +4552,7 @@ static void mem_cgroup_stat_aggregate(st
#ifdef CONFIG_MEMCG_NMI_SAFETY_REQUIRES_ATOMIC
static void flush_nmi_stats(struct mem_cgroup *memcg, struct mem_cgroup *parent)
{
+ long *state_local = memcg_state_local_array(memcg);
int nid;
if (atomic_read(&memcg->kmem_stat)) {
@@ -4529,7 +4560,8 @@ static void flush_nmi_stats(struct mem_c
int index = memcg_stats_index(MEMCG_KMEM);
memcg->vmstats->state[index] += kmem;
- memcg->vmstats->state_local[index] += kmem;
+ if (state_local)
+ state_local[index] += kmem;
if (parent)
parent->vmstats->state_pending[index] += kmem;
}
@@ -4551,7 +4583,8 @@ static void flush_nmi_stats(struct mem_c
if (plstats)
plstats->state_pending[index] += slab;
memcg->vmstats->state[index] += slab;
- memcg->vmstats->state_local[index] += slab;
+ if (state_local)
+ state_local[index] += slab;
if (parent)
parent->vmstats->state_pending[index] += slab;
}
@@ -4564,7 +4597,8 @@ static void flush_nmi_stats(struct mem_c
if (plstats)
plstats->state_pending[index] += slab;
memcg->vmstats->state[index] += slab;
- memcg->vmstats->state_local[index] += slab;
+ if (state_local)
+ state_local[index] += slab;
if (parent)
parent->vmstats->state_pending[index] += slab;
}
@@ -4589,7 +4623,7 @@ static void mem_cgroup_css_rstat_flush(s
ac = (struct aggregate_control) {
.aggregate = memcg->vmstats->state,
- .local = memcg->vmstats->state_local,
+ .local = memcg_state_local_array(memcg),
.pending = memcg->vmstats->state_pending,
.ppending = parent ? parent->vmstats->state_pending : NULL,
.cstat = statc->state,
@@ -4600,7 +4634,7 @@ static void mem_cgroup_css_rstat_flush(s
ac = (struct aggregate_control) {
.aggregate = memcg->vmstats->events,
- .local = memcg->vmstats->events_local,
+ .local = memcg_events_local_array(memcg),
.pending = memcg->vmstats->events_pending,
.ppending = parent ? parent->vmstats->events_pending : NULL,
.cstat = statc->events,
_
Patches currently in -mm which might be from joannelkoong@gmail.com are
mm-memcontrol-skip-non-hierarchical-memcg-wide-stats-when-v1-is-unavailable.patch
reply other threads:[~2026-09-06 0:12 UTC|newest]
Thread overview: [no followups] expand[flat|nested] mbox.gz Atom feed
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260906001239.4D33F1F00A3A@smtp.kernel.org \
--to=akpm@linux-foundation.org \
--cc=hannes@cmpxchg.org \
--cc=joannelkoong@gmail.com \
--cc=mhocko@kernel.org \
--cc=mm-commits@vger.kernel.org \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.