From: Michal Hocko <mhocko@suse.com>
To: Tao Cui <cui.tao@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>,
akpm@linux-foundation.org, linux-mm@kvack.org,
cgroups@vger.kernel.org, linux-kernel@vger.kernel.org,
hannes@cmpxchg.org, roman.gushchin@linux.dev,
muchun.song@linux.dev, Tao Cui <cuitao@kylinos.cn>
Subject: Re: [PATCH] mm: page_counter: reject empty string in page_counter_memparse()
Date: Wed, 19 Aug 2026 09:00:36 +0200 [thread overview]
Message-ID: <aoVUlFdZYLFn_gvJ@tiehlicka> (raw)
In-Reply-To: <9bd74181-2ce4-4d69-a353-614685559ceb@linux.dev>
On Wed 19-08-26 11:00:09, Tao Cui wrote:
> Hi, Michal, Shakeel
>
> 在 2026/8/19 00:43, Michal Hocko 写道:
> > On Tue 18-08-26 08:06:22, Shakeel Butt wrote:
> >> On Tue, Aug 18, 2026 at 09:45:56AM +0200, Michal Hocko wrote:
> >>> On Mon 17-08-26 09:16:40, Shakeel Butt wrote:
> >>>> On Mon, Aug 17, 2026 at 12:26:52PM +0800, Tao Cui wrote:
> >>>>> From: Tao Cui <cuitao@kylinos.cn>
> >>>>>
> >>>>> memparse() consumes no characters on an empty input and leaves the
> >>>>> end pointer at the terminating NUL. The only validation in
> >>>>> page_counter_memparse() checks for trailing characters, so an empty
> >>>>> input slips through and the limit becomes 0.
> >>>>>
> >>>>> All limit write callbacks of the memory controller strstrip() the
> >>>>> input before calling this helper, so a script that writes an unset
> >>>>> variable hits this path:
> >>>>>
> >>>>> LIMIT=
> >>>>> echo "$LIMIT" > $CG/memory.max
> >>>>> echo $?
> >>>>> 0
> >>>>> cat $CG/memory.max
> >>>>> 0
> >>>>>
> >>>>> Nothing reports the mistake: the limit is now 0 and the OOM killer
> >>>>> goes after every task in the cgroup. The same happens for
> >>>>> memory.min, memory.low, memory.high, memory.swap.high,
> >>>>> memory.swap.max and memory.zswap.max, where 0 silently removes the
> >>>>> protection or disables swap and zswap.
> >>>>>
> >>>>> Reject the input when no characters were consumed, which is the one
> >>>>> case the trailing-character check cannot catch.
> >>>>>
> >>>>> Fixes: 3e32cb2e0a12 ("mm: memcontrol: lockless page counters")
> >>>>> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
> >>>>> ---
> >>>>> mm/page_counter.c | 2 +-
> >>>>> 1 file changed, 1 insertion(+), 1 deletion(-)
> >>>>>
> >>>>> diff --git a/mm/page_counter.c b/mm/page_counter.c
> >>>>> index 661e0f2a5127..d14db705b04f 100644
> >>>>> --- a/mm/page_counter.c
> >>>>> +++ b/mm/page_counter.c
> >>>>> @@ -281,7 +281,7 @@ int page_counter_memparse(const char *buf, const char *max,
> >>>>> }
> >>>>>
> >>>>> bytes = memparse(buf, &end);
> >>>>> - if (*end != '\0')
> >>>>> + if (*end != '\0' || end == buf)
> >>>>> return -EINVAL;
> >>>>
> >>>> I wonder if someone started depending on this behavior. In that case it is
> >>>> better to return error instead of silently ignore, so we will hear complains
> >>>> loudly. This looks good to me.
> >>>
> >>> This is backward incompatible change and I am wondering why should we
> >>> even risk regression.
> >>
> >> Mainly I was wondering if this is intentional or unintentional. If this us
> >> unintentional, can we fix it without anyone noticing?
> >
> > My guess would be this was just omission. Those happen and over years we
> > have learned that userspace is quite creative at using those.
> >
> >> However if we are ok with this then let's make is formal and make this a
> >> documented behavior. I don't have any strong opinion either way but I think you
> >> are saying it safer to just assume this is intentional. Fine with me.
> >
> > My main question is why should we even bother to change this in the
> > first place? Is that reason stronger than a theoretical breakage of
> > userspace that we might learn much later?
>
> Since you asked "why bother", here's how I ran into it.
>
> The patch actually came from a production incident rather than a code
> audit.
>
> A maintenance script on a cluster accidentally wrote an unset variable
> into memory.max of a workload cgroup. The write succeeded, and the
> workload in the cgroup was subsequently OOM-killed. There was no
> indication that the successful write had caused it, so it took quite
> some time to trace the OOMs back to that script.
Understood.
> When I checked the documentation, I noticed that the cpuset controller
> explicitly documents the semantics of empty writes ("An empty value
> indicates that the cgroup is using the same setting as the nearest
> cgroup ancestor..."), while the memory controller documentation says
> nothing about empty input.
yes, this is really unfortunate and mistakes like that happen.
> I reproduced the same behavior in isolation on a Kubernetes cluster
> (v1.29, cgroup v2, two-container pod, 384M pod limit):
>
> # LIMIT=
> # echo "$LIMIT" > $CG/memory.max
> # echo $?
> 0
>
> m6demo 0/2 OOMKilled 0
>
> oom-kill: constraint=CONSTRAINT_MEMCG,
> oom_memcg=/kubepods.slice/.../kubepods-burstable-pod....slice
> Memory cgroup out of memory: Killed process 339529 (sleep) ...
> anon-rss:32kB, file-rss:452kB
>
> The process selected by the OOM killer had less than 1 MB resident
> under a 384M pod limit, but the empty write was accepted as 0,
> immediately triggering a memcg OOM. From the caller's perspective, the
> write simply succeeded.
>
> One caveat is that kubelet reconciles the pod-level memory.max within
> seconds, so the window is short there, although container-level files
> are not reconciled.
>
> So the patch came out of that incident and the documentation gap it
> exposed. My thinking was that rejecting an empty input would make such
> mistakes fail immediately instead of silently changing the limit to 0.
Thanks for sharing the story. Next time I would recommend to make that a
part of the changelog.
> Whether that benefit outweighs the compatibility risk is not mine to
> decide.
And I am still not convinced. I do understand your frustration from the
debugging the issue. In any way, if other maintainers decide this change
is worth I will not stand in the way.
> If it doesn't, then documenting the current empty-write
> behavior, similar to cpuset, would also address the ambiguity that led
> me to investigate this in the first place.
Agreed!
--
Michal Hocko
SUSE Labs
prev parent reply other threads:[~2026-08-19 7:00 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-17 4:26 [PATCH] mm: page_counter: reject empty string in page_counter_memparse() Tao Cui
2026-08-17 16:16 ` Shakeel Butt
2026-08-18 7:45 ` Michal Hocko
2026-08-18 15:06 ` Shakeel Butt
2026-08-18 16:43 ` Michal Hocko
2026-08-19 3:00 ` Tao Cui
2026-08-19 7:00 ` Michal Hocko [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aoVUlFdZYLFn_gvJ@tiehlicka \
--to=mhocko@suse.com \
--cc=akpm@linux-foundation.org \
--cc=cgroups@vger.kernel.org \
--cc=cui.tao@linux.dev \
--cc=cuitao@kylinos.cn \
--cc=hannes@cmpxchg.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox