From: Michal Hocko <mhocko@suse.com>
To: Eric Chanudet <echanude@redhat.com>
Cc: "Andrew Morton" <akpm@linux-foundation.org>,
"David Hildenbrand" <david@kernel.org>,
"Lorenzo Stoakes" <ljs@kernel.org>,
"Liam R. Howlett" <liam@infradead.org>,
"Vlastimil Babka" <vbabka@kernel.org>,
"Mike Rapoport" <rppt@kernel.org>,
"Suren Baghdasaryan" <surenb@google.com>,
"Tejun Heo" <tj@kernel.org>,
"Johannes Weiner" <hannes@cmpxchg.org>,
"Michal Koutný" <mkoutny@suse.com>,
"Jonathan Corbet" <corbet@lwn.net>,
"Shuah Khan" <skhan@linuxfoundation.org>,
"Roman Gushchin" <roman.gushchin@linux.dev>,
"Shakeel Butt" <shakeel.butt@linux.dev>,
"Muchun Song" <muchun.song@linux.dev>,
"Shuah Khan" <shuah@kernel.org>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
cgroups@vger.kernel.org, linux-doc@vger.kernel.org,
linux-kselftest@vger.kernel.org,
"Maxime Ripard" <mripard@redhat.com>,
"Albert Esteve" <aesteve@redhat.com>
Subject: Re: [PATCH 00/11] mm/cma: charge cma allocation to memcg using per area counters
Date: Wed, 26 Aug 2026 10:04:27 +0200 [thread overview]
Message-ID: <ao6eC2SUIGVz3Qct@tiehlicka> (raw)
In-Reply-To: <ao3tFg4JLNlbylTR@echanude-thinkpadx1carbongen13.rmtusma.csb>
On Tue 25-08-26 16:58:51, Eric Chanudet wrote:
> On Tue, Aug 25, 2026 at 09:19:21PM +0200, Michal Hocko wrote:
> > On Tue 25-08-26 14:33:48, Eric Chanudet wrote:
> > > On Tue, Aug 25, 2026 at 04:59:52PM +0200, Michal Hocko wrote:
> > > > On Tue 25-08-26 10:47:53, Eric Chanudet wrote:
> > > > > On Mon, Aug 24, 2026 at 10:58:18AM +0200, Michal Hocko wrote:
> > > > > > On Fri 21-08-26 14:56:52, Eric Chanudet wrote:
> > > > > > > CMA allocations are currently unaccounted for by cgroup memory
> > > > > > > controllers. As system resources, they should fall under memcg, but CMA
> > > > > > > areas partition the available space for different purposes and memcg
> > > > > > > doesn't have a good representation for that.
> > > > > >
> > > > > > Which CMA usecases are covered by this work? It would be also great to
> > > > > > spend more time describing usecases.
> > > > >
> > > > > We would like to offer some usage guaranties to userspace processes
> > > > > ending up doing allocations in CMA.
> > > > >
> > > > > For example, a shared CMA area is described in device-tree for an ARM64
> > > > > platforms. Userspace components could then, for example, allocate from
> > > > > it through the dmabuf heap, or a device or framework-specific ioctl for
> > > > > that matter, to use the buffer with sensors. The dtb may have other CMA
> > > > > areas described additionally that may or may not be used by that
> > > > > component. In this context, we would like the ability to limit one of
> > > > > the userspace component to over-allocate and choke the other(s).
> > > >
> > > > How exactly is this supposed to work? How is the CMA access controled
> > > > and opted in for accounting. What happens when memcg limits are hit. And
> > > > many more details, please.
> > > >
> > >
> > > The administrator opts in by mounting cgroupfs with
> > > memory_cma_accounting. At which point the cma allocator will charge CMA
> > > allocations against memcg and manages a per area counter depending on
> > > what area the allocation was made into.
> >
> > So each CMA area will have its own counter and limits?
>
> Yes, in order to enforce a limit per CMA area this series add a page
> counter for each area. Areas are fixed and discovered early so the
> counters are added to struct mem_cgroup and initialized when the cgroup
> is created.
>
> An admin would then use the cgroupfs entries to assign an area limit to
> a given cgroup, something like the following, using the reserved area
> for example:
> mount -o remount,memory_cma_accounting /sys/fs/cgroup
> echo +memory > /sys/fs/cgroup/cgroup.subtree_control
> mkdir /sys/fs/cgroup/mycg
> echo 16M > /sys/fs/cgroup/mycg/memory.cma.reserved.max
> echo 64M > /sys/fs/cgroup/mycg/memory.max
OK, thanks for the clarification. This confirms my initial suspicion but
it is better to have it clearly articulated. I can see several problems
with this approach. First and formost I do not think dealing with all
cmas this way is manageable. This can become a mess very quickly if we
have one limit per cma and too coarse if there is a single one. I also
have my doubts about space allocation control through a simple limit for
something that is effectively a reserved physical space.
I might be proven wrong but unless cma serves objects of a uniform
size then this will simply not work in practice. Hitting ENOSPC without
hitting limits and thus impractical for shared space management.
[...]
> > > It looked consistent to use memcg since movable pages from regular
> > > allocations may end up in available CMA regions until a CMA allocation
> > > needs the space and has them moved. So in an extreme case, hogging the
> > > CMA space of a large enough area could trigger system memory pressure.
> >
> > I really do not understand what you mean here.
>
> Non-CMA allocations can end up in CMA physical regions when necessary
> (ALLOC_CMA flag).
Correct. But those are a subject of migration so any such placement
should not be blocking real CMA allocations.
> Since both CMA allocations and other system
> allocations are represented the same way, with differences only in
> properties, and they can live in the same regions, it sounds reasonable
> to account for both under the same counter.
From the memcg POV we do account physically consumed memory. So yes,
it makes no difference where the memory comes from. We only care about
the overall capacity you can constrain or protect. Generally speaking it
makes sense to charge heavy memory consumers directly triggerable from
the userspace.
That is all memcg can provide you with. Specific requirements for
specific types of memory is a different story. We currently cannot
control per-numa node for example. There is an ongoing work to make
memcg memory tier aware.
> > > > > memcg
> > > > > looked like a good fit to achieve this, albeit handling the areas, so a
> > > > > cgroup has a quota in a given CMA resource.
> > > >
> > > > Please expand more on why do you think this fits into the memcg model.
> > > > AFAIU we are talking about a unreclaimable memory and reservations of
> > > > CMA areas.
> > >
> > > Since memcg already accounts for some unreclaimable memory (kmem,
> > > hugetlb),
> >
> > hugetlb pages have their own controller
> >
> > > or induces failure if no reclamation is possible, I did not
> > > see CMA allocations being unreclaimable to be a blocker to track what is
> > > otherwise system memory.
> >
> > yes, we can have unreclaimable memory charged to memcg, that is not a
> > real problem. We have all sorts of memory consumers that need to be
> > capped charged to the memcg. If dmabufs are another ones then fine, just
> > charge allocated pages from the cma area. It is the "make all cma users
> > memcg aware and have per cma limits" that I am really struggling with.
>
> CMA is system memory independently from its usage though, and in cases
> with shared CMA areas multiple users can allocate from them. Yet the
> kernel cannot enforce usage limits.
Correct. Those are effectively a shared memory pools without any
control. I do not think memcg is a good method to enfore any usage
limits for that though for reasons mentioned above. Memcg is effective
at capping the overall memory consumption of a workload. Not really
great when it comes to a specific memory pool control.
> > You cannot really assume usecase, requirements, lifetime etc. for an
> > arbitrary cma area. I do not think this is a viable way forward. Focus
> > on your real usecase, which seems to be dmabufs.
>
> While dmabufs are indeed my main use case, they are quite generic and
> may not always have system memory backing them (device memory). Working
> at the CMA allocator alleviated these disparities.
>
> > Explain what do you want to achieve and then we can think whether memcg
> > is the right model for that usecase
>
> Hopefully I expressed this in a better way by now. In short, enforce
> usage limits for concurrent CMA users using shared CMA resources.
Thanks. Yes this is more clear now. And it resembles hugetlb situation
more than memcg. You simply need a memory pool specific access and usage
control. Dispersing that to a global memcg limit seems rather coarse and
I would say impractical. So it really calls for a per pool control with
an understanding of how the specific pool really works.
--
Michal Hocko
SUSE Labs
next prev parent reply other threads:[~2026-08-26 8:04 UTC|newest]
Thread overview: 30+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-21 18:56 [PATCH 00/11] mm/cma: charge cma allocation to memcg using per area counters Eric Chanudet
2026-08-21 18:56 ` [PATCH 01/11] mm/cma: drop const for struct page on release API Eric Chanudet
2026-08-21 18:56 ` [PATCH 02/11] mm/cma: include linux/cma.h in cma.h Eric Chanudet
2026-08-21 18:56 ` [PATCH 03/11] cgroup: add memory_cma_accounting mount option Eric Chanudet
2026-08-24 7:02 ` Maxime Ripard
2026-08-24 22:03 ` Eric Chanudet
2026-08-21 18:56 ` [PATCH 04/11] memcg: add cma charge/uncharge functions for area counters Eric Chanudet
2026-08-21 18:56 ` [PATCH 05/11] mm/cma: charge cma allocation to memcg per " Eric Chanudet
2026-08-21 18:56 ` [PATCH 06/11] memcg: register per-area usage counters in cgroupfs Eric Chanudet
2026-08-21 18:56 ` [PATCH 07/11] selftests: cgroup: add cma configs for cgroup selftest suite Eric Chanudet
2026-08-21 18:57 ` [PATCH 08/11] selftests: cgroup: add memcg cma tests Eric Chanudet
2026-08-21 18:57 ` [PATCH 09/11] selftests: cgroup: add a vmtest script for memcg Eric Chanudet
2026-08-21 18:57 ` [PATCH 10/11] selftests: cgroup: add memcg hugetlb_cma tests Eric Chanudet
2026-08-21 18:57 ` [PATCH 11/11] selftests: cgroup: amend vmtest-memcg to run the hugetlb cma tests Eric Chanudet
2026-08-23 7:02 ` [PATCH 00/11] mm/cma: charge cma allocation to memcg using per area counters Mike Rapoport
2026-08-24 21:58 ` Eric Chanudet
2026-08-25 7:40 ` Michal Hocko
2026-08-24 8:58 ` Michal Hocko
2026-08-25 14:47 ` Eric Chanudet
2026-08-25 14:59 ` Michal Hocko
2026-08-25 18:33 ` Eric Chanudet
2026-08-25 19:19 ` Michal Hocko
2026-08-25 20:58 ` Eric Chanudet
2026-08-26 8:04 ` Michal Hocko [this message]
2026-08-25 15:45 ` David Hildenbrand (Arm)
2026-08-25 16:26 ` Michal Hocko
2026-08-25 18:50 ` Eric Chanudet
2026-08-26 8:02 ` David Hildenbrand (Arm)
2026-08-24 13:18 ` Michal Koutný
2026-08-25 14:54 ` Eric Chanudet
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ao6eC2SUIGVz3Qct@tiehlicka \
--to=mhocko@suse.com \
--cc=aesteve@redhat.com \
--cc=akpm@linux-foundation.org \
--cc=cgroups@vger.kernel.org \
--cc=corbet@lwn.net \
--cc=david@kernel.org \
--cc=echanude@redhat.com \
--cc=hannes@cmpxchg.org \
--cc=liam@infradead.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mkoutny@suse.com \
--cc=mripard@redhat.com \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=rppt@kernel.org \
--cc=shakeel.butt@linux.dev \
--cc=shuah@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=surenb@google.com \
--cc=tj@kernel.org \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox