* [PATCH] memcg: keep propagating stats updates to ancestors
@ 2026-09-11 3:29 Runli
2026-09-11 19:36 ` Shakeel Butt
0 siblings, 1 reply; 2+ messages in thread
From: Runli @ 2026-09-11 3:29 UTC (permalink / raw)
To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt,
linux-mm
Cc: Muchun Song, Andrew Morton, Yosry Ahmed, cgroups, linux-kernel,
littleswimmingwhale, Runli
memcg_rstat_updated() stops walking the ancestors when a memcg exceeds
the flush threshold. A concurrent flush can leave an ancestor below the
threshold, so the early exit prevents it from receiving further updates.
Consider root and leaf on a 4-CPU system, with a threshold of 256.
CPU 0 is flushing another CPU's stats after CPU 1's stats have already
been flushed. The counts below are shared stats_updates counters:
CPU 0 (root flush) CPU 1 (leaf updates)
------------------ --------------------
clear leaf->stats_updates
add 64 to leaf and root
clear root->stats_updates
finish flush
leaf = 64, root = 0
4 more batches of 64:
leaf = 320, root = 256
next update:
leaf > threshold
break; root is not updated
Root remains at the threshold, which is insufficient to trigger a flush,
while further leaf updates keep taking the early exit. The periodic
forced flush restores progress, but until then the root's aggregated
statistics can fall far behind. For example, page cache grows at rate of
1 GiB/s, a two-second wait can leave the reported usage about 2 GiB below
the actual usage.
Skip only the flushable node and continue walking its ancestors, so each
ancestor can accumulate updates until it exceeds the flush threshold.
Fixes: 60cada258dfe ("memcg: optimize memcg_rstat_updated")
Signed-off-by: Runli <mingyu.he@shopee.com>
---
Testing:
Based on commit 893e11787f78, with and without this patch,
using 32 concurrent instances of:
netperf -H 127.0.0.1 -p 12875 -t TCP_STREAM -l 60 \
-T <client_cpu>,<server_cpu> -P 0 -v 0 -f m -- -m 1024
Both netperf and netserver ran in the same leaf memcg, directly below
root (root -> leaf), with memory bound to NUMA node 0.
Each kernel was tested for five 60-second runs after a 10-second warmup.
Median aggregate throughput decreased from 107.26 to 105.23 Gbit/s
(1.9%) with this change.
I think fixing the correctness issue is worth the roughly 2% throughput
cost.
mm/memcontrol.c | 7 +++----
1 file changed, 3 insertions(+), 4 deletions(-)
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1271d390b617..4600d9c9244e 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -729,12 +729,11 @@ static inline void memcg_rstat_updated(struct mem_cgroup *memcg, long val,
for (; statc_pcpu; statc_pcpu = statc->parent_pcpu) {
statc = this_cpu_ptr(statc_pcpu);
/*
- * If @memcg is already flushable then all its ancestors are
- * flushable as well and also there is no need to increase
- * stats_updates.
+ * A concurrent flush may have reset an ancestor's counter.
+ * Skip this node if flushable, but keep walking the ancestors.
*/
if (memcg_vmstats_needs_flush(statc->vmstats))
- break;
+ continue;
stats_updates = this_cpu_add_return(statc_pcpu->stats_updates,
abs(val));
--
2.43.0
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [PATCH] memcg: keep propagating stats updates to ancestors
2026-09-11 3:29 [PATCH] memcg: keep propagating stats updates to ancestors Runli
@ 2026-09-11 19:36 ` Shakeel Butt
0 siblings, 0 replies; 2+ messages in thread
From: Shakeel Butt @ 2026-09-11 19:36 UTC (permalink / raw)
To: Runli
Cc: Johannes Weiner, Michal Hocko, Roman Gushchin, linux-mm,
Muchun Song, Andrew Morton, Yosry Ahmed, cgroups, linux-kernel,
littleswimmingwhale
On Fri, Sep 11, 2026 at 11:29:05AM +0800, Runli wrote:
> memcg_rstat_updated() stops walking the ancestors when a memcg exceeds
> the flush threshold. A concurrent flush can leave an ancestor below the
> threshold, so the early exit prevents it from receiving further updates.
>
> Consider root and leaf on a 4-CPU system, with a threshold of 256.
> CPU 0 is flushing another CPU's stats after CPU 1's stats have already
> been flushed. The counts below are shared stats_updates counters:
>
> CPU 0 (root flush) CPU 1 (leaf updates)
> ------------------ --------------------
> clear leaf->stats_updates
> add 64 to leaf and root
> clear root->stats_updates
> finish flush
> leaf = 64, root = 0
> 4 more batches of 64:
> leaf = 320, root = 256
> next update:
> leaf > threshold
> break; root is not updated
>
> Root remains at the threshold, which is insufficient to trigger a flush,
> while further leaf updates keep taking the early exit. The periodic
> forced flush restores progress, but until then the root's aggregated
> statistics can fall far behind. For example, page cache grows at rate of
> 1 GiB/s, a two-second wait can leave the reported usage about 2 GiB below
> the actual usage.
>
> Skip only the flushable node and continue walking its ancestors, so each
> ancestor can accumulate updates until it exceeds the flush threshold.
>
> Fixes: 60cada258dfe ("memcg: optimize memcg_rstat_updated")
> Signed-off-by: Runli <mingyu.he@shopee.com>
> ---
> Testing:
> Based on commit 893e11787f78, with and without this patch,
> using 32 concurrent instances of:
> netperf -H 127.0.0.1 -p 12875 -t TCP_STREAM -l 60 \
> -T <client_cpu>,<server_cpu> -P 0 -v 0 -f m -- -m 1024
>
> Both netperf and netserver ran in the same leaf memcg, directly below
> root (root -> leaf), with memory bound to NUMA node 0.
> Each kernel was tested for five 60-second runs after a 10-second warmup.
>
> Median aggregate throughput decreased from 107.26 to 105.23 Gbit/s
> (1.9%) with this change.
>
> I think fixing the correctness issue is worth the roughly 2% throughput
> cost.
Do you have any real world or production data to show that this is an issue? I
would not change anything unless there is real data. Also we have periodic flush
which will fix this (I have not looked deeper into this to see if this is a real
issue). In addition we have bpf interface for the stats for users who really
care about accuracy and efficiency.
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-11 19:36 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-11 3:29 [PATCH] memcg: keep propagating stats updates to ancestors Runli
2026-09-11 19:36 ` Shakeel Butt
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox