From: Michal Hocko <mhocko@suse.com>
To: Tao Cui <cui.tao@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>,
akpm@linux-foundation.org, linux-mm@kvack.org,
cgroups@vger.kernel.org, linux-kernel@vger.kernel.org,
hannes@cmpxchg.org, roman.gushchin@linux.dev,
muchun.song@linux.dev, Tao Cui <cuitao@kylinos.cn>
Subject: Re: [PATCH] mm: page_counter: reject empty string in page_counter_memparse()
Date: Wed, 19 Aug 2026 09:00:36 +0200 [thread overview]
Message-ID: <aoVUlFdZYLFn_gvJ@tiehlicka> (raw)
In-Reply-To: <9bd74181-2ce4-4d69-a353-614685559ceb@linux.dev>
On Wed 19-08-26 11:00:09, Tao Cui wrote:
> Hi, Michal, Shakeel
>
> 在 2026/8/19 00:43, Michal Hocko 写道:
> > On Tue 18-08-26 08:06:22, Shakeel Butt wrote:
> >> On Tue, Aug 18, 2026 at 09:45:56AM +0200, Michal Hocko wrote:
> >>> On Mon 17-08-26 09:16:40, Shakeel Butt wrote:
> >>>> On Mon, Aug 17, 2026 at 12:26:52PM +0800, Tao Cui wrote:
> >>>>> From: Tao Cui <cuitao@kylinos.cn>
> >>>>>
> >>>>> memparse() consumes no characters on an empty input and leaves the
> >>>>> end pointer at the terminating NUL. The only validation in
> >>>>> page_counter_memparse() checks for trailing characters, so an empty
> >>>>> input slips through and the limit becomes 0.
> >>>>>
> >>>>> All limit write callbacks of the memory controller strstrip() the
> >>>>> input before calling this helper, so a script that writes an unset
> >>>>> variable hits this path:
> >>>>>
> >>>>> LIMIT=
> >>>>> echo "$LIMIT" > $CG/memory.max
> >>>>> echo $?
> >>>>> 0
> >>>>> cat $CG/memory.max
> >>>>> 0
> >>>>>
> >>>>> Nothing reports the mistake: the limit is now 0 and the OOM killer
> >>>>> goes after every task in the cgroup. The same happens for
> >>>>> memory.min, memory.low, memory.high, memory.swap.high,
> >>>>> memory.swap.max and memory.zswap.max, where 0 silently removes the
> >>>>> protection or disables swap and zswap.
> >>>>>
> >>>>> Reject the input when no characters were consumed, which is the one
> >>>>> case the trailing-character check cannot catch.
> >>>>>
> >>>>> Fixes: 3e32cb2e0a12 ("mm: memcontrol: lockless page counters")
> >>>>> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
> >>>>> ---
> >>>>> mm/page_counter.c | 2 +-
> >>>>> 1 file changed, 1 insertion(+), 1 deletion(-)
> >>>>>
> >>>>> diff --git a/mm/page_counter.c b/mm/page_counter.c
> >>>>> index 661e0f2a5127..d14db705b04f 100644
> >>>>> --- a/mm/page_counter.c
> >>>>> +++ b/mm/page_counter.c
> >>>>> @@ -281,7 +281,7 @@ int page_counter_memparse(const char *buf, const char *max,
> >>>>> }
> >>>>>
> >>>>> bytes = memparse(buf, &end);
> >>>>> - if (*end != '\0')
> >>>>> + if (*end != '\0' || end == buf)
> >>>>> return -EINVAL;
> >>>>
> >>>> I wonder if someone started depending on this behavior. In that case it is
> >>>> better to return error instead of silently ignore, so we will hear complains
> >>>> loudly. This looks good to me.
> >>>
> >>> This is backward incompatible change and I am wondering why should we
> >>> even risk regression.
> >>
> >> Mainly I was wondering if this is intentional or unintentional. If this us
> >> unintentional, can we fix it without anyone noticing?
> >
> > My guess would be this was just omission. Those happen and over years we
> > have learned that userspace is quite creative at using those.
> >
> >> However if we are ok with this then let's make is formal and make this a
> >> documented behavior. I don't have any strong opinion either way but I think you
> >> are saying it safer to just assume this is intentional. Fine with me.
> >
> > My main question is why should we even bother to change this in the
> > first place? Is that reason stronger than a theoretical breakage of
> > userspace that we might learn much later?
>
> Since you asked "why bother", here's how I ran into it.
>
> The patch actually came from a production incident rather than a code
> audit.
>
> A maintenance script on a cluster accidentally wrote an unset variable
> into memory.max of a workload cgroup. The write succeeded, and the
> workload in the cgroup was subsequently OOM-killed. There was no
> indication that the successful write had caused it, so it took quite
> some time to trace the OOMs back to that script.
Understood.
> When I checked the documentation, I noticed that the cpuset controller
> explicitly documents the semantics of empty writes ("An empty value
> indicates that the cgroup is using the same setting as the nearest
> cgroup ancestor..."), while the memory controller documentation says
> nothing about empty input.
yes, this is really unfortunate and mistakes like that happen.
> I reproduced the same behavior in isolation on a Kubernetes cluster
> (v1.29, cgroup v2, two-container pod, 384M pod limit):
>
> # LIMIT=
> # echo "$LIMIT" > $CG/memory.max
> # echo $?
> 0
>
> m6demo 0/2 OOMKilled 0
>
> oom-kill: constraint=CONSTRAINT_MEMCG,
> oom_memcg=/kubepods.slice/.../kubepods-burstable-pod....slice
> Memory cgroup out of memory: Killed process 339529 (sleep) ...
> anon-rss:32kB, file-rss:452kB
>
> The process selected by the OOM killer had less than 1 MB resident
> under a 384M pod limit, but the empty write was accepted as 0,
> immediately triggering a memcg OOM. From the caller's perspective, the
> write simply succeeded.
>
> One caveat is that kubelet reconciles the pod-level memory.max within
> seconds, so the window is short there, although container-level files
> are not reconciled.
>
> So the patch came out of that incident and the documentation gap it
> exposed. My thinking was that rejecting an empty input would make such
> mistakes fail immediately instead of silently changing the limit to 0.
Thanks for sharing the story. Next time I would recommend to make that a
part of the changelog.
> Whether that benefit outweighs the compatibility risk is not mine to
> decide.
And I am still not convinced. I do understand your frustration from the
debugging the issue. In any way, if other maintainers decide this change
is worth I will not stand in the way.
> If it doesn't, then documenting the current empty-write
> behavior, similar to cpuset, would also address the ambiguity that led
> me to investigate this in the first place.
Agreed!
--
Michal Hocko
SUSE Labs
prev parent reply other threads:[~2026-08-19 7:00 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-17 4:26 [PATCH] mm: page_counter: reject empty string in page_counter_memparse() Tao Cui
2026-08-17 16:16 ` Shakeel Butt
2026-08-18 7:45 ` Michal Hocko
2026-08-18 15:06 ` Shakeel Butt
2026-08-18 16:43 ` Michal Hocko
2026-08-19 3:00 ` Tao Cui
2026-08-19 7:00 ` Michal Hocko [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aoVUlFdZYLFn_gvJ@tiehlicka \
--to=mhocko@suse.com \
--cc=akpm@linux-foundation.org \
--cc=cgroups@vger.kernel.org \
--cc=cui.tao@linux.dev \
--cc=cuitao@kylinos.cn \
--cc=hannes@cmpxchg.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.