* [PATCH v3 0/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark @ 2026-08-14 1:29 Ridong Chen 2026-08-14 1:29 ` [PATCH v3 1/2] memcg: acquire peaks_lock when reading memory.peak Ridong Chen 2026-08-14 1:29 ` [PATCH v3 2/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark Ridong Chen 0 siblings, 2 replies; 5+ messages in thread From: Ridong Chen @ 2026-08-14 1:29 UTC (permalink / raw) To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton Cc: Muchun Song, Tejun Heo, David Finkel, Michal Koutný, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Tao Cui, Ridong Chen, Ridong Chen From: Ridong Chen <chenridong@xiaomi.com> The memory.peak / memory.swap.peak per-fd watermark tracking has two issues. Each open fd is a watcher and reads back max(its own value, the shared local_watermark); both bugs live in that scheme. Worst case for both is the same and is userspace-visible: a reader of memory.peak (or memory.swap.peak) gets a value lower than the true peak, so a tool that sizes or bills a cgroup by its peak usage under-reports it. Patch 1 (read side) fixes the race Sashiko pointed out in the v1 review [1]: peak_show() inspects local_watermark and the per-fd values without holding peaks_lock, so a reader that races an unrelated peak_write() reset briefly observes the lowered value. Transient. It takes peaks_lock in the show path. Patch 2 (write side) fixes peak_write(): on a reset it stores the current usage into the other watchers instead of the old watermark, so once usage has dropped from a peak a reset on one fd drags every other fd's peak down too, even fds that never reset. --- Changes since v2: - Spell out the worst-case userspace-visible effect, per Andrew's Go back to v1 [2]. Changes since v1: - New patch 1: hold peaks_lock in the peak readers (Sashiko). - Patch 2: floor the peers with max(usage, local_watermark), mirroring peak_show(), and skip the writing fd (Johannes Weiner). [1] https://sashiko.dev/#/patchset/20260730115314.1069089-1-ridong.chen@linux.dev?part=1 [2] https://lore.kernel.org/all/20260730115314.1069089-1-ridong.chen@linux.dev/ Ridong Chen (2): memcg: acquire peaks_lock when reading memory.peak mm, memcg: fix memory.peak reset clobbering other fds' watermark mm/memcontrol.c | 21 ++++++++++++++++----- 1 file changed, 16 insertions(+), 5 deletions(-) -- 2.34.1 ^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 1/2] memcg: acquire peaks_lock when reading memory.peak 2026-08-14 1:29 [PATCH v3 0/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark Ridong Chen @ 2026-08-14 1:29 ` Ridong Chen 2026-08-14 2:12 ` Muchun Song 2026-08-14 1:29 ` [PATCH v3 2/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark Ridong Chen 1 sibling, 1 reply; 5+ messages in thread From: Ridong Chen @ 2026-08-14 1:29 UTC (permalink / raw) To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton Cc: Muchun Song, Tejun Heo, David Finkel, Michal Koutný, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Tao Cui, Ridong Chen, Ridong Chen From: Ridong Chen <chenridong@xiaomi.com> Sashiko reported that a reader can transiently observe a lower peak within a race window [1]. peak_show() returns max(local_watermark, ofp->value), but peak_write() updates those two under peaks_lock while the reader takes no lock. The interleaving is: writer (reset on fd A) reader (fd B) ---------------------- ------------- usage = page_counter_read(pc) WRITE_ONCE(local_watermark, usage) // watermark lowered to usage lw = READ_ONCE(local_watermark) // sees the lowered usage val = READ_ONCE(ofp->value) // B's value not updated yet return max(lw, val) // both low -> low peak WRITE_ONCE(peer_ctx->value, usage) // B updated, but too late Fix it by acquiring peaks_lock when reading the peak, so the reader sees a consistent snapshot of local_watermark and the per-fd values. The same race applies to memory.swap.peak, which shares peaks_lock and the peak_write() path, so take the lock there as well. [1] https://sashiko.dev/#/patchset/20260730115314.1069089-1-ridong.chen@linux.dev?part=1 Fixes: c6f53ed8f213 ("mm, memcg: cg2 memory{.swap,}.peak write handlers") Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ridong Chen <chenridong@xiaomi.com> Acked-by: Johannes Weiner <hannes@cmpxchg.org> Acked-by: Shakeel Butt <shakeel.butt@linux.dev> --- mm/memcontrol.c | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 17da1f43b7d3..6dd8756870ff 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -4713,8 +4713,13 @@ static int peak_show(struct seq_file *sf, void *v, struct page_counter *pc) static int memory_peak_show(struct seq_file *sf, void *v) { struct mem_cgroup *memcg = mem_cgroup_from_css(seq_css(sf)); + int ret; - return peak_show(sf, v, &memcg->memory); + spin_lock(&memcg->peaks_lock); + ret = peak_show(sf, v, &memcg->memory); + spin_unlock(&memcg->peaks_lock); + + return ret; } static int peak_open(struct kernfs_open_file *of) @@ -5858,8 +5863,13 @@ static u64 swap_current_read(struct cgroup_subsys_state *css, static int swap_peak_show(struct seq_file *sf, void *v) { struct mem_cgroup *memcg = mem_cgroup_from_css(seq_css(sf)); + int ret; - return peak_show(sf, v, &memcg->swap); + spin_lock(&memcg->peaks_lock); + ret = peak_show(sf, v, &memcg->swap); + spin_unlock(&memcg->peaks_lock); + + return ret; } static ssize_t swap_peak_write(struct kernfs_open_file *of, char *buf, -- 2.34.1 ^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH v3 1/2] memcg: acquire peaks_lock when reading memory.peak 2026-08-14 1:29 ` [PATCH v3 1/2] memcg: acquire peaks_lock when reading memory.peak Ridong Chen @ 2026-08-14 2:12 ` Muchun Song 2026-08-14 2:25 ` Ridong Chen 0 siblings, 1 reply; 5+ messages in thread From: Muchun Song @ 2026-08-14 2:12 UTC (permalink / raw) To: Ridong Chen Cc: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton, Tejun Heo, David Finkel, Michal Koutný, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Tao Cui, Ridong Chen > On Aug 14, 2026, at 09:29, Ridong Chen <ridong.chen@linux.dev> wrote: > > From: Ridong Chen <chenridong@xiaomi.com> > > Sashiko reported that a reader can transiently observe a lower peak > within a race window [1]. peak_show() returns > max(local_watermark, ofp->value), but peak_write() updates those two > under peaks_lock while the reader takes no lock. The interleaving is: > > writer (reset on fd A) reader (fd B) > ---------------------- ------------- > usage = page_counter_read(pc) > WRITE_ONCE(local_watermark, usage) > // watermark lowered to usage > lw = READ_ONCE(local_watermark) > // sees the lowered usage > val = READ_ONCE(ofp->value) > // B's value not updated yet > return max(lw, val) > // both low -> low peak > WRITE_ONCE(peer_ctx->value, usage) > // B updated, but too late > > Fix it by acquiring peaks_lock when reading the peak, so the reader sees > a consistent snapshot of local_watermark and the per-fd values. The same > race applies to memory.swap.peak, which shares peaks_lock and the > peak_write() path, so take the lock there as well. > > [1] https://sashiko.dev/#/patchset/20260730115314.1069089-1-ridong.chen@linux.dev?part=1 > Fixes: c6f53ed8f213 ("mm, memcg: cg2 memory{.swap,}.peak write handlers") > Assisted-by: Claude:claude-opus-4-8 > Signed-off-by: Ridong Chen <chenridong@xiaomi.com> > Acked-by: Johannes Weiner <hannes@cmpxchg.org> > Acked-by: Shakeel Butt <shakeel.butt@linux.dev> > --- > mm/memcontrol.c | 14 ++++++++++++-- > 1 file changed, 12 insertions(+), 2 deletions(-) > > diff --git a/mm/memcontrol.c b/mm/memcontrol.c > index 17da1f43b7d3..6dd8756870ff 100644 > --- a/mm/memcontrol.c > +++ b/mm/memcontrol.c > @@ -4713,8 +4713,13 @@ static int peak_show(struct seq_file *sf, void *v, struct page_counter *pc) > static int memory_peak_show(struct seq_file *sf, void *v) > { > struct mem_cgroup *memcg = mem_cgroup_from_css(seq_css(sf)); > + int ret; > > - return peak_show(sf, v, &memcg->memory); > + spin_lock(&memcg->peaks_lock); Why not to use guard(spinlock)(&memcg->peaks_lock) to simplify the code. Muchun, Thanks. > + ret = peak_show(sf, v, &memcg->memory); > + spin_unlock(&memcg->peaks_lock); > + > + return ret; > } > > static int peak_open(struct kernfs_open_file *of) > @@ -5858,8 +5863,13 @@ static u64 swap_current_read(struct cgroup_subsys_state *css, > static int swap_peak_show(struct seq_file *sf, void *v) > { > struct mem_cgroup *memcg = mem_cgroup_from_css(seq_css(sf)); > + int ret; > > - return peak_show(sf, v, &memcg->swap); > + spin_lock(&memcg->peaks_lock); > + ret = peak_show(sf, v, &memcg->swap); > + spin_unlock(&memcg->peaks_lock); > + > + return ret; > } > > static ssize_t swap_peak_write(struct kernfs_open_file *of, char *buf, > -- > 2.34.1 > ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3 1/2] memcg: acquire peaks_lock when reading memory.peak 2026-08-14 2:12 ` Muchun Song @ 2026-08-14 2:25 ` Ridong Chen 0 siblings, 0 replies; 5+ messages in thread From: Ridong Chen @ 2026-08-14 2:25 UTC (permalink / raw) To: Muchun Song Cc: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton, Tejun Heo, David Finkel, Michal Koutný, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Tao Cui, Ridong Chen On 8/14/2026 10:12 AM, Muchun Song wrote: > > >> On Aug 14, 2026, at 09:29, Ridong Chen <ridong.chen@linux.dev> wrote: >> >> From: Ridong Chen <chenridong@xiaomi.com> >> >> Sashiko reported that a reader can transiently observe a lower peak >> within a race window [1]. peak_show() returns >> max(local_watermark, ofp->value), but peak_write() updates those two >> under peaks_lock while the reader takes no lock. The interleaving is: >> >> writer (reset on fd A) reader (fd B) >> ---------------------- ------------- >> usage = page_counter_read(pc) >> WRITE_ONCE(local_watermark, usage) >> // watermark lowered to usage >> lw = READ_ONCE(local_watermark) >> // sees the lowered usage >> val = READ_ONCE(ofp->value) >> // B's value not updated yet >> return max(lw, val) >> // both low -> low peak >> WRITE_ONCE(peer_ctx->value, usage) >> // B updated, but too late >> >> Fix it by acquiring peaks_lock when reading the peak, so the reader sees >> a consistent snapshot of local_watermark and the per-fd values. The same >> race applies to memory.swap.peak, which shares peaks_lock and the >> peak_write() path, so take the lock there as well. >> >> [1] https://sashiko.dev/#/patchset/20260730115314.1069089-1-ridong.chen@linux.dev?part=1 >> Fixes: c6f53ed8f213 ("mm, memcg: cg2 memory{.swap,}.peak write handlers") >> Assisted-by: Claude:claude-opus-4-8 >> Signed-off-by: Ridong Chen <chenridong@xiaomi.com> >> Acked-by: Johannes Weiner <hannes@cmpxchg.org> >> Acked-by: Shakeel Butt <shakeel.butt@linux.dev> >> --- >> mm/memcontrol.c | 14 ++++++++++++-- >> 1 file changed, 12 insertions(+), 2 deletions(-) >> >> diff --git a/mm/memcontrol.c b/mm/memcontrol.c >> index 17da1f43b7d3..6dd8756870ff 100644 >> --- a/mm/memcontrol.c >> +++ b/mm/memcontrol.c >> @@ -4713,8 +4713,13 @@ static int peak_show(struct seq_file *sf, void *v, struct page_counter *pc) >> static int memory_peak_show(struct seq_file *sf, void *v) >> { >> struct mem_cgroup *memcg = mem_cgroup_from_css(seq_css(sf)); >> + int ret; >> >> - return peak_show(sf, v, &memcg->memory); >> + spin_lock(&memcg->peaks_lock); > > Why not to use guard(spinlock)(&memcg->peaks_lock) to simplify the code. > > Muchun, > Thanks. > Good idea. Will update. >> + ret = peak_show(sf, v, &memcg->memory); >> + spin_unlock(&memcg->peaks_lock); >> + >> + return ret; >> } >> >> static int peak_open(struct kernfs_open_file *of) >> @@ -5858,8 +5863,13 @@ static u64 swap_current_read(struct cgroup_subsys_state *css, >> static int swap_peak_show(struct seq_file *sf, void *v) >> { >> struct mem_cgroup *memcg = mem_cgroup_from_css(seq_css(sf)); >> + int ret; >> >> - return peak_show(sf, v, &memcg->swap); >> + spin_lock(&memcg->peaks_lock); >> + ret = peak_show(sf, v, &memcg->swap); >> + spin_unlock(&memcg->peaks_lock); >> + >> + return ret; >> } >> >> static ssize_t swap_peak_write(struct kernfs_open_file *of, char *buf, >> -- >> 2.34.1 >> > -- Best regards Ridong ^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 2/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark 2026-08-14 1:29 [PATCH v3 0/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark Ridong Chen 2026-08-14 1:29 ` [PATCH v3 1/2] memcg: acquire peaks_lock when reading memory.peak Ridong Chen @ 2026-08-14 1:29 ` Ridong Chen 1 sibling, 0 replies; 5+ messages in thread From: Ridong Chen @ 2026-08-14 1:29 UTC (permalink / raw) To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton Cc: Muchun Song, Tejun Heo, David Finkel, Michal Koutný, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Tao Cui, Ridong Chen, Ridong Chen From: Ridong Chen <chenridong@xiaomi.com> Writing to memory.peak resets the peak for that fd only. Each fd is a watcher and reads back max(its own value, the shared local_watermark). peak_write() resets by lowering local_watermark to the current usage. To keep the other watchers' peaks it then walks the watcher list, but it stores the current usage into them instead of the old watermark. So once usage has dropped from a peak, a reset on one fd wrongly drags every other fd's peak down too, even fds that never reset. Reproduced on 7.2.0-rc5-next under QEMU, two fds A and B on one cgroup: B sees the peak (410624 KB), usage drops, then A resets -- and B's peak collapses to 1060 KB although B never reset. With this patch B keeps reading 410624 KB. Fix: save the old watermark before lowering it and use that to floor the other watchers, so a reset only affects the fd that issued it. Fixes: c6f53ed8f213 ("mm, memcg: cg2 memory{.swap,}.peak write handlers") Closes: https://sashiko.dev/#/patchset/20260807090000.1532495-1-ridong.chen@linux.dev Assisted-by: Claude:claude-opus-4-8 Acked-by: Tao Cui <cuitao@kylinos.cn> Acked-by: Johannes Weiner <hannes@cmpxchg.org> Signed-off-by: Ridong Chen <chenridong@xiaomi.com> --- mm/memcontrol.c | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 6dd8756870ff..5e5edc988119 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -4748,7 +4748,7 @@ static ssize_t peak_write(struct kernfs_open_file *of, char *buf, size_t nbytes, loff_t off, struct page_counter *pc, struct list_head *watchers) { - unsigned long usage; + unsigned long usage, old_watermark; struct cgroup_of_peak *peer_ctx; struct mem_cgroup *memcg = mem_cgroup_from_css(of_css(of)); struct cgroup_of_peak *ofp = of_peak(of); @@ -4756,11 +4756,12 @@ static ssize_t peak_write(struct kernfs_open_file *of, char *buf, size_t nbytes, spin_lock(&memcg->peaks_lock); usage = page_counter_read(pc); + old_watermark = READ_ONCE(pc->local_watermark); WRITE_ONCE(pc->local_watermark, usage); list_for_each_entry(peer_ctx, watchers, list) - if (usage > peer_ctx->value) - WRITE_ONCE(peer_ctx->value, usage); + if (peer_ctx != ofp && old_watermark > peer_ctx->value) + WRITE_ONCE(peer_ctx->value, old_watermark); /* initial write, register watcher */ if (ofp->value == OFP_PEAK_UNSET) -- 2.34.1 ^ permalink raw reply related [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-08-14 2:25 UTC | newest] Thread overview: 5+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-08-14 1:29 [PATCH v3 0/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark Ridong Chen 2026-08-14 1:29 ` [PATCH v3 1/2] memcg: acquire peaks_lock when reading memory.peak Ridong Chen 2026-08-14 2:12 ` Muchun Song 2026-08-14 2:25 ` Ridong Chen 2026-08-14 1:29 ` [PATCH v3 2/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark Ridong Chen
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox