From: Matthias Goergens <matthias.goergens@gmail.com>
To: Andy Lutomirski <luto@amacapital.net>
Cc: Andrew Morton <akpm@linux-foundation.org>,
Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
Johannes Weiner <hannes@cmpxchg.org>,
David Hildenbrand <david@kernel.org>,
Michal Hocko <mhocko@kernel.org>,
Shakeel Butt <shakeel.butt@linux.dev>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Yosry Ahmed <yosry@kernel.org>,
Youngjun Park <youngjun.park@lge.com>,
Baoquan He <baoquan.he@linux.dev>, Barry Song <baohua@kernel.org>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
cgroups@vger.kernel.org, linux-api@vger.kernel.org,
Alejandro Colomar <alx@kernel.org>,
linux-man@vger.kernel.org, Karel Zak <kzak@redhat.com>,
util-linux@vger.kernel.org,
Jani Nikula <jani.nikula@linux.intel.com>,
Joonas Lahtinen <joonas.lahtinen@linux.intel.com>,
Vivi Rodrigo <rodrigo.vivi@intel.com>,
Tvrtko Ursulin <tursulin@ursulin.net>,
David Airlie <airlied@gmail.com>, Simona Vetter <simona@ffwll.ch>,
intel-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org,
Rafael J Wysocki <rafael@kernel.org>,
Pavel Machek <pavel@kernel.org>,
Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>,
linux-pm@vger.kernel.org, linux-arm-kernel@lists.infradead.org,
Shuah Khan <shuah@kernel.org>,
linux-kselftest@vger.kernel.org
Subject: Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
Date: Mon, 28 Sep 2026 23:24:38 +0800 [thread overview]
Message-ID: <20260928152438.2694783-1-matthias.goergens@gmail.com> (raw)
In-Reply-To: <3FDF3191-94A5-42B3-A33E-12021344D7BD@amacapital.net>
Hi Andy,
Thanks for reading it, and for the direct feedback.
> if the system is under overall memory pressure, why do you care what
> triggered the particular swap operation that is being processed?
You're right, and that question gets at what I actually want better
than the series does. The goal is: when there is no memory pressure,
swap-out may do more work, including allocating memory (compression,
copy-on-write, filesystem-backed swap), because moving cold pages out
to make room for page cache is what swap is for most of the time.
Under real pressure, swap-out must not depend on that.
The RFC was the smallest change I could think of in that direction.
It used "who started the reclaim" as a stand-in for "is there
pressure", and that is wrong both ways: memory.reclaim during global
pressure may allocate, while kswapd with plenty of free memory may
not. Building on Kairui's suggestion in this thread (swap tiers), or
on virtual swap, may well be the better route, and I'm looking at both
for v3.
> Why can't the kernel figure this out itself?
It can: the backend knows whether its writes may allocate (zram knows
it compresses, a filesystem knows at swapon whether the swapfile is
copy-on-write or compressed), so it should declare that, rather than
the administrator setting a flag.
On the practical questions: v2 does nothing to balance fullness
between areas or to empty conventional swap, and does not change OOM
selection. Those are fair gaps and I'll address them, or say
explicitly what is out of scope, in v3.
On the recursion question: within the reclaiming task it can't
happen. memory.reclaim and per-node reclaim both run with PF_MEMALLOC
set (memalloc_noreclaim_save() in try_to_free_mem_cgroup_pages() and
__node_reclaim()), so an allocation the backend makes in that task
never enters direct reclaim; it either fails or, unless it passes
__GFP_NOMEMALLOC, dips into the reserves. Two things do need care. A
backend allocating there can drain the emergency reserves unless it
passes __GFP_NOMEMALLOC or fails fast. And work the backend hands to
another thread, such as a filesystem or zvol worker, runs without
PF_MEMALLOC and can enter reclaim itself; that is only safe if that
reclaim never waits for the write the worker is meant to complete.
> The remainder of the writeup is IMO somewhat incoherent.
Agreed. I cut the v1 cover letter down too far and lost the
definitions it depended on. v3's will start from the goal above and
define its terms.
> Adding "eligible" to a bunch of calls does nothing to explain what's
> going on.
It exists because v2 made "how much swap is free" depend on who asks.
Reclaim itself checks free swap before scanning anonymous pages
(can_reclaim_anon_pages(), MGLRU's get_swappiness()) and before
allocating a slot (folio_alloc_swap()), and callers such as the GPU
shrinkers check it before pushing objects towards swap. With an
offload-only area, each of those would otherwise count space that the
calling reclaim may not use. In v3 I'd keep get_nr_swap_pages()
meaning "usable by ordinary reclaim", so those checks stay as they are
and only the offload path asks for a different count.
Thanks,
Matthias
next prev parent reply other threads:[~2026-09-28 15:24 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
2026-09-26 10:16 ` Kairui Song
2026-09-28 15:19 ` Matthias Goergens
2026-10-04 18:13 ` Youngjun Park
2026-09-26 23:32 ` Chris Li
2026-09-27 0:03 ` Chris Li
2026-09-27 17:33 ` Andy Lutomirski
2026-09-28 15:24 ` Matthias Goergens [this message]
2026-09-28 0:19 ` Chris Li
2026-09-28 15:19 ` Matthias Goergens
2026-09-29 0:27 ` Chris Li
2026-09-28 5:19 ` Christoph Hellwig
2026-09-28 15:19 ` Matthias Goergens
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928152438.2694783-1-matthias.goergens@gmail.com \
--to=matthias.goergens@gmail.com \
--cc=airlied@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=alx@kernel.org \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=catalin.marinas@arm.com \
--cc=cgroups@vger.kernel.org \
--cc=chrisl@kernel.org \
--cc=david@kernel.org \
--cc=dri-devel@lists.freedesktop.org \
--cc=hannes@cmpxchg.org \
--cc=intel-gfx@lists.freedesktop.org \
--cc=jani.nikula@linux.intel.com \
--cc=joonas.lahtinen@linux.intel.com \
--cc=kasong@tencent.com \
--cc=kzak@redhat.com \
--cc=linux-api@vger.kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-man@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-pm@vger.kernel.org \
--cc=luto@amacapital.net \
--cc=mhocko@kernel.org \
--cc=nphamcs@gmail.com \
--cc=pavel@kernel.org \
--cc=rafael@kernel.org \
--cc=rodrigo.vivi@intel.com \
--cc=shakeel.butt@linux.dev \
--cc=shikemeng@huaweicloud.com \
--cc=shuah@kernel.org \
--cc=simona@ffwll.ch \
--cc=tursulin@ursulin.net \
--cc=util-linux@vger.kernel.org \
--cc=will@kernel.org \
--cc=yosry@kernel.org \
--cc=youngjun.park@lge.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox