* [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
@ 2026-09-26 4:55 Matthias Goergens
2026-09-26 10:16 ` Kairui Song
` (4 more replies)
0 siblings, 5 replies; 13+ messages in thread
From: Matthias Goergens @ 2026-09-26 4:55 UTC (permalink / raw)
To: Andrew Morton
Cc: Matthias Goergens, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J . Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest
This RFC adds offload-only swap areas for deliberate cold-page offload to
backends unsuitable for pressure reclaim. ZFS zvol swap has documented
deadlocks under memory pressure [3]; compressed swap is another target
because writes may need memory despite free logical slots.
Explicit proactive reclaim can use these areas alongside conventional
swap, ordered by priority. Ordinary reclaim cannot initiate non-zero
writes to them, including through retained swap entries. Conventional
capacity is not reserved for emergencies. Operators choose the policy;
the kernel does not measure headroom or make an allocating backend safe.
TMO/Senpai [4] and DAMON_RECLAIM [5] are related cold-page reclaim
approaches. This RFC admits offload-only swap through memory.reclaim and
per-node reclaim, but not DAMON. Other related work includes per-cgroup
zswap writeback control [6], Virtual Swap Space v4 [1] and swap tiers v10
[2].
The util-linux companion [8] proposes swapon --offload-only and an fstab
option. Older swapon silently ignores the fstab token, so persistent
activation needs discussion. Apply the separate i915 fix [9] first to avoid
a pre-existing folio-lock leak when shmem writeback is skipped.
The series fixes zram selftest device tracking and error reporting, adds
the swap policy with DRM eligibility checks, and tests routing, workingset
activation and retained-entry write refusal. It is based on mm-new at
995829088503.
Feedback is particularly welcome on:
- Activation-time eligibility versus backend placement or migration:
is retained-entry refusal, with possible reclaim churn or OOM,
acceptable without migration?
- Eligible-capacity accounting, including overcommit limits and OOM
scoring; cache recovery, cluster invalidation and workingset activation.
- Persistent activation and visibility: current swap listings do not
expose the policy.
Validation includes x86 builds, affected arm64 MTE objects and builds
without swap or memory cgroups. QEMU tests covered routing, retained-entry
recovery, data integrity and cleanup, with negative controls for write
refusal and teardown failure.
I have also been using this policy on my own machine with experimental
bcachefs swap support, with no problems observed so far. Deliberate stress
testing is confined to VMs; I do not deliberately stress-test this machine.
Changes since v1, incorporating Sashiko's public review [7]:
- Improve selftest isolation, retained-swap measurement and memlock skips.
- Track allocated zram devices, wait for udev probes and report cleanup
errors without deleting data through a mount that could not be released.
- Clarify activation, capacity reporting and conventional fallback limits.
[1] https://patchew.org/linux/20260825153238.2695446-1-nphamcs@gmail.com/
[2] https://patchew.org/linux/20260713025644.170839-1-youngjun.park@lge.com/
[3] https://openzfs.github.io/openzfs-docs/Project%20and%20Community/FAQ.html#using-a-zvol-for-a-swap-device-on-linux
[4] https://engineering.fb.com/2022/06/20/data-infrastructure/transparent-memory-offloading-more-memory-at-a-fraction-of-the-cost-and-power/
[5] https://www.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
[6] https://www.kernel.org/doc/html/latest/admin-guide/mm/zswap.html
[7] https://sashiko.dev/#/patchset/20260914104501.3960616-1-matthias.goergens@gmail.com
[8] https://github.com/util-linux/util-linux/pull/4633
[9] https://lore.kernel.org/all/20260914105606.3997649-1-matthias.goergens@gmail.com/
v1: https://lore.kernel.org/all/20260914104501.3960616-1-matthias.goergens@gmail.com/
Matthias Goergens (4):
selftests: zram: track owned devices and report cleanup failures
mm: restrict offload-only swap to proactive reclaim
selftests: zram: cover offload-only swap policy
selftests: zram: cover retained offload-only entries
Documentation/mm/swap.rst | 105 ++++
drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
.../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
include/linux/swap.h | 29 +-
include/linux/vm_event_item.h | 1 +
mm/memcontrol.c | 20 +-
mm/page_io.c | 24 +-
mm/swapfile.c | 123 ++++-
mm/vmscan.c | 31 +-
mm/vmstat.c | 1 +
tools/testing/selftests/zram/.gitignore | 2 +
tools/testing/selftests/zram/Makefile | 4 +-
tools/testing/selftests/zram/README | 20 +-
tools/testing/selftests/zram/config | 10 +-
tools/testing/selftests/zram/settings | 1 +
tools/testing/selftests/zram/swap_offload.c | 425 +++++++++++++++++
.../selftests/zram/workingset_offload.c | 207 ++++++++
tools/testing/selftests/zram/zram.sh | 9 +
tools/testing/selftests/zram/zram01.sh | 9 +-
tools/testing/selftests/zram/zram02.sh | 9 +-
tools/testing/selftests/zram/zram03.sh | 174 +++++++
tools/testing/selftests/zram/zram04.sh | 164 +++++++
tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++
tools/testing/selftests/zram/zram_lib.sh | 189 ++++++--
28 files changed, 1926 insertions(+), 94 deletions(-)
create mode 100644 tools/testing/selftests/zram/settings
create mode 100644 tools/testing/selftests/zram/swap_offload.c
create mode 100644 tools/testing/selftests/zram/workingset_offload.c
create mode 100755 tools/testing/selftests/zram/zram03.sh
create mode 100755 tools/testing/selftests/zram/zram04.sh
create mode 100755 tools/testing/selftests/zram/zram05.sh
--
2.55.0
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
@ 2026-09-26 10:16 ` Kairui Song
2026-09-28 15:19 ` Matthias Goergens
2026-09-26 23:32 ` Chris Li
` (3 subsequent siblings)
4 siblings, 1 reply; 13+ messages in thread
From: Kairui Song @ 2026-09-26 10:16 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Chris Li, Youngjun Park, Baoquan He,
Johannes Weiner, David Hildenbrand, Michal Hocko, Shakeel Butt,
Kemeng Shi, Nhat Pham, Yosry Ahmed, Barry Song, linux-mm,
linux-kernel, cgroups, linux-api, Alejandro Colomar, linux-man,
Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx, dri-devel,
Rafael J . Wysocki, Pavel Machek, Catalin Marinas, Will Deacon,
linux-pm, linux-arm-kernel, Shuah Khan, linux-kselftest
On Sat, Sep 26, 2026 at 12:55 PM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
>
> This RFC adds offload-only swap areas for deliberate cold-page offload to
> backends unsuitable for pressure reclaim. ZFS zvol swap has documented
> deadlocks under memory pressure [3]; compressed swap is another target
> because writes may need memory despite free logical slots.
>
> Explicit proactive reclaim can use these areas alongside conventional
> swap, ordered by priority. Ordinary reclaim cannot initiate non-zero
> writes to them, including through retained swap entries. Conventional
> capacity is not reserved for emergencies. Operators choose the policy;
> the kernel does not measure headroom or make an allocating backend safe.
>
> TMO/Senpai [4] and DAMON_RECLAIM [5] are related cold-page reclaim
> approaches. This RFC admits offload-only swap through memory.reclaim and
> per-node reclaim, but not DAMON. Other related work includes per-cgroup
> zswap writeback control [6], Virtual Swap Space v4 [1] and swap tiers v10
> [2].
>
> The util-linux companion [8] proposes swapon --offload-only and an fstab
> option. Older swapon silently ignores the fstab token, so persistent
> activation needs discussion. Apply the separate i915 fix [9] first to avoid
> a pre-existing folio-lock leak when shmem writeback is skipped.
>
> The series fixes zram selftest device tracking and error reporting, adds
> the swap policy with DRM eligibility checks, and tests routing, workingset
> activation and retained-entry write refusal. It is based on mm-new at
> 995829088503.
>
> Feedback is particularly welcome on:
>
> - Activation-time eligibility versus backend placement or migration:
> is retained-entry refusal, with possible reclaim churn or OOM,
> acceptable without migration?
> - Eligible-capacity accounting, including overcommit limits and OOM
> scoring; cache recovery, cluster invalidation and workingset activation.
> - Persistent activation and visibility: current swap listings do not
> expose the policy.
>
> Validation includes x86 builds, affected arm64 MTE objects and builds
> without swap or memory cgroups. QEMU tests covered routing, retained-entry
> recovery, data integrity and cleanup, with negative controls for write
> refusal and teardown failure.
>
> I have also been using this policy on my own machine with experimental
> bcachefs swap support, with no problems observed so far. Deliberate stress
> testing is confined to VMs; I do not deliberately stress-test this machine.
>
> Changes since v1, incorporating Sashiko's public review [7]:
>
> - Improve selftest isolation, retained-swap measurement and memlock skips.
> - Track allocated zram devices, wait for udev probes and report cleanup
> errors without deleting data through a mount that could not be released.
> - Clarify activation, capacity reporting and conventional fallback limits.
>
> [1] https://patchew.org/linux/20260825153238.2695446-1-nphamcs@gmail.com/
> [2] https://patchew.org/linux/20260713025644.170839-1-youngjun.park@lge.com/
> [3] https://openzfs.github.io/openzfs-docs/Project%20and%20Community/FAQ.html#using-a-zvol-for-a-swap-device-on-linux
> [4] https://engineering.fb.com/2022/06/20/data-infrastructure/transparent-memory-offloading-more-memory-at-a-fraction-of-the-cost-and-power/
> [5] https://www.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
> [6] https://www.kernel.org/doc/html/latest/admin-guide/mm/zswap.html
> [7] https://sashiko.dev/#/patchset/20260914104501.3960616-1-matthias.goergens@gmail.com
> [8] https://github.com/util-linux/util-linux/pull/4633
> [9] https://lore.kernel.org/all/20260914105606.3997649-1-matthias.goergens@gmail.com/
>
> v1: https://lore.kernel.org/all/20260914104501.3960616-1-matthias.goergens@gmail.com/
>
> Matthias Goergens (4):
> selftests: zram: track owned devices and report cleanup failures
> mm: restrict offload-only swap to proactive reclaim
> selftests: zram: cover offload-only swap policy
> selftests: zram: cover retained offload-only entries
>
> Documentation/mm/swap.rst | 105 ++++
> drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
> .../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
> drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
> drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
> drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
> drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
> include/linux/swap.h | 29 +-
> include/linux/vm_event_item.h | 1 +
> mm/memcontrol.c | 20 +-
> mm/page_io.c | 24 +-
> mm/swapfile.c | 123 ++++-
> mm/vmscan.c | 31 +-
> mm/vmstat.c | 1 +
> tools/testing/selftests/zram/.gitignore | 2 +
> tools/testing/selftests/zram/Makefile | 4 +-
> tools/testing/selftests/zram/README | 20 +-
> tools/testing/selftests/zram/config | 10 +-
> tools/testing/selftests/zram/settings | 1 +
> tools/testing/selftests/zram/swap_offload.c | 425 +++++++++++++++++
> .../selftests/zram/workingset_offload.c | 207 ++++++++
> tools/testing/selftests/zram/zram.sh | 9 +
> tools/testing/selftests/zram/zram01.sh | 9 +-
> tools/testing/selftests/zram/zram02.sh | 9 +-
> tools/testing/selftests/zram/zram03.sh | 174 +++++++
> tools/testing/selftests/zram/zram04.sh | 164 +++++++
> tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++
> tools/testing/selftests/zram/zram_lib.sh | 189 ++++++--
> 28 files changed, 1926 insertions(+), 94 deletions(-)
Hi Matthias
That's a lot of changes for a rather limited usage. IIUC what it want
to archive is, for different kind of reclaim, swap operations should
only go to certain devices?
This really sounds like another usage of the swap tiering design: A
default per-cgroup tier setup, and a one time tier limit during
proactive reclaim.
Just like we already have a "swappiness=" parameter in the reclaim
interface, perhaps adding a "swap.tier=" would be better and much
cleaner to achieve the same goal? Based on Youngjun's work here:
https://lore.kernel.org/all/20260916183437.2946306-1-youngjun.park@lge.com/
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
2026-09-26 10:16 ` Kairui Song
@ 2026-09-26 23:32 ` Chris Li
2026-09-27 0:03 ` Chris Li
2026-09-27 17:33 ` Andy Lutomirski
` (2 subsequent siblings)
4 siblings, 1 reply; 13+ messages in thread
From: Chris Li @ 2026-09-26 23:32 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx, dri-devel,
Rafael J . Wysocki, Pavel Machek, Catalin Marinas, Will Deacon,
linux-pm, linux-arm-kernel, Shuah Khan, linux-kselftest
On Fri, Sep 25, 2026 at 6:55 PM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
> The series fixes zram selftest device tracking and error reporting, adds
> the swap policy with DRM eligibility checks, and tests routing, workingset
> activation and retained-entry write refusal. It is based on mm-new at
> 995829088503.
Hi Matthias,
I can't find mm-new 995829088503 anymore.
Your patch does not apply cleanly to the current tip of mm-new
mm-unstable, or mm-stable branches.
Do you have a git repo with your patch applied?
mm-new is whatever Andrew can find on the mailing list, regardless of
whether it has been reviewed. That mm-new branch is very unstable,
even more so than mm-unstable, which should at least have some review
already. Maybe consider using mm-stable or mm-unstable as base next
time.
Chris
>
> Feedback is particularly welcome on:
>
> - Activation-time eligibility versus backend placement or migration:
> is retained-entry refusal, with possible reclaim churn or OOM,
> acceptable without migration?
> - Eligible-capacity accounting, including overcommit limits and OOM
> scoring; cache recovery, cluster invalidation and workingset activation.
> - Persistent activation and visibility: current swap listings do not
> expose the policy.
>
> Validation includes x86 builds, affected arm64 MTE objects and builds
> without swap or memory cgroups. QEMU tests covered routing, retained-entry
> recovery, data integrity and cleanup, with negative controls for write
> refusal and teardown failure.
>
> I have also been using this policy on my own machine with experimental
> bcachefs swap support, with no problems observed so far. Deliberate stress
> testing is confined to VMs; I do not deliberately stress-test this machine.
>
> Changes since v1, incorporating Sashiko's public review [7]:
>
> - Improve selftest isolation, retained-swap measurement and memlock skips.
> - Track allocated zram devices, wait for udev probes and report cleanup
> errors without deleting data through a mount that could not be released.
> - Clarify activation, capacity reporting and conventional fallback limits.
>
> [1] https://patchew.org/linux/20260825153238.2695446-1-nphamcs@gmail.com/
> [2] https://patchew.org/linux/20260713025644.170839-1-youngjun.park@lge.com/
> [3] https://openzfs.github.io/openzfs-docs/Project%20and%20Community/FAQ.html#using-a-zvol-for-a-swap-device-on-linux
> [4] https://engineering.fb.com/2022/06/20/data-infrastructure/transparent-memory-offloading-more-memory-at-a-fraction-of-the-cost-and-power/
> [5] https://www.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
> [6] https://www.kernel.org/doc/html/latest/admin-guide/mm/zswap.html
> [7] https://sashiko.dev/#/patchset/20260914104501.3960616-1-matthias.goergens@gmail.com
> [8] https://github.com/util-linux/util-linux/pull/4633
> [9] https://lore.kernel.org/all/20260914105606.3997649-1-matthias.goergens@gmail.com/
>
> v1: https://lore.kernel.org/all/20260914104501.3960616-1-matthias.goergens@gmail.com/
>
> Matthias Goergens (4):
> selftests: zram: track owned devices and report cleanup failures
> mm: restrict offload-only swap to proactive reclaim
> selftests: zram: cover offload-only swap policy
> selftests: zram: cover retained offload-only entries
>
> Documentation/mm/swap.rst | 105 ++++
> drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
> .../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
> drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
> drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
> drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
> drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
> include/linux/swap.h | 29 +-
> include/linux/vm_event_item.h | 1 +
> mm/memcontrol.c | 20 +-
> mm/page_io.c | 24 +-
> mm/swapfile.c | 123 ++++-
> mm/vmscan.c | 31 +-
> mm/vmstat.c | 1 +
> tools/testing/selftests/zram/.gitignore | 2 +
> tools/testing/selftests/zram/Makefile | 4 +-
> tools/testing/selftests/zram/README | 20 +-
> tools/testing/selftests/zram/config | 10 +-
> tools/testing/selftests/zram/settings | 1 +
> tools/testing/selftests/zram/swap_offload.c | 425 +++++++++++++++++
> .../selftests/zram/workingset_offload.c | 207 ++++++++
> tools/testing/selftests/zram/zram.sh | 9 +
> tools/testing/selftests/zram/zram01.sh | 9 +-
> tools/testing/selftests/zram/zram02.sh | 9 +-
> tools/testing/selftests/zram/zram03.sh | 174 +++++++
> tools/testing/selftests/zram/zram04.sh | 164 +++++++
> tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++
> tools/testing/selftests/zram/zram_lib.sh | 189 ++++++--
> 28 files changed, 1926 insertions(+), 94 deletions(-)
> create mode 100644 tools/testing/selftests/zram/settings
> create mode 100644 tools/testing/selftests/zram/swap_offload.c
> create mode 100644 tools/testing/selftests/zram/workingset_offload.c
> create mode 100755 tools/testing/selftests/zram/zram03.sh
> create mode 100755 tools/testing/selftests/zram/zram04.sh
> create mode 100755 tools/testing/selftests/zram/zram05.sh
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 23:32 ` Chris Li
@ 2026-09-27 0:03 ` Chris Li
0 siblings, 0 replies; 13+ messages in thread
From: Chris Li @ 2026-09-27 0:03 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx, dri-devel,
Rafael J . Wysocki, Pavel Machek, Catalin Marinas, Will Deacon,
linux-pm, linux-arm-kernel, Shuah Khan, linux-kselftest
On Sat, Sep 26, 2026 at 1:32 PM Chris Li <chrisl@kernel.org> wrote:
>
> On Fri, Sep 25, 2026 at 6:55 PM Matthias Goergens
> <matthias.goergens@gmail.com> wrote:
> > The series fixes zram selftest device tracking and error reporting, adds
> > the swap policy with DRM eligibility checks, and tests routing, workingset
> > activation and retained-entry write refusal. It is based on mm-new at
> > 995829088503.
>
> Hi Matthias,
>
> I can't find mm-new 995829088503 anymore.
>
> Your patch does not apply cleanly to the current tip of mm-new
> mm-unstable, or mm-stable branches.
> Do you have a git repo with your patch applied?
Never mind, I have the conflict resolved on mm-unstable.
Sorry I usually need to have the patch appliable to use my own review scripts.
Chris
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
2026-09-26 10:16 ` Kairui Song
2026-09-26 23:32 ` Chris Li
@ 2026-09-27 17:33 ` Andy Lutomirski
2026-09-28 15:24 ` Matthias Goergens
2026-09-28 0:19 ` Chris Li
2026-09-28 5:19 ` Christoph Hellwig
4 siblings, 1 reply; 13+ messages in thread
From: Andy Lutomirski @ 2026-09-27 17:33 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Vivi Rodrigo, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest, Matthias Goergens
> On Sep 25, 2026, at 9:55 PM, Matthias Goergens <matthias.goergens@gmail.com> wrote:
>
> This RFC adds offload-only swap areas for deliberate cold-page offload to
> backends unsuitable for pressure reclaim. ZFS zvol swap has documented
> deadlocks under memory pressure [3]; compressed swap is another target
> because writes may need memory despite free logical slots.
This seems unnecessarily annoying to configure. It seems to me that you’re creating a bizarrely named administrative flag that means, roughly, “swapping to this location may allocate memory”. Why can’t the kernel figure this out itself?
For that matter, how well does this even work in practice? If I configure two swap devices, one “offload-only” and one conventional, it seems like the relative fullness of the devices will be mostly an accident of what triggers swap as the system is running. If the conventional swap fills up, is there a means to proactively empty it? Under OOM conditions, will anything preferentially kill tasks that reference memory in conventional swap?
Is the locking and recursion structure of the mm code such that we won’t have deadlocks where cgroup triggers swap to an “offload-only” device, which allocates, which causes memory pressure, which then tries to swap more to conventional swap? (Maybe this works fine.)
For that matter, if the system is under overall memory pressure, why do you care what triggered the particular swap operation that is being processed?
The remainder of the writeup is IMO somewhat incoherent. To the extent that AI was used, can you read it and make sure it makes sense?
And the first few hundred lines of patch I skimmed seem like incomprehensible churn. Adding “eligible” to a bunch of calls does nothing to explain what’s going on.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
` (2 preceding siblings ...)
2026-09-27 17:33 ` Andy Lutomirski
@ 2026-09-28 0:19 ` Chris Li
2026-09-28 15:19 ` Matthias Goergens
2026-09-28 5:19 ` Christoph Hellwig
4 siblings, 1 reply; 13+ messages in thread
From: Chris Li @ 2026-09-28 0:19 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx, dri-devel,
Rafael J . Wysocki, Pavel Machek, Catalin Marinas, Will Deacon,
linux-pm, linux-arm-kernel, Shuah Khan, linux-kselftest
On Fri, Sep 25, 2026 at 6:55 PM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
>
> This RFC adds offload-only swap areas for deliberate cold-page offload to
> backends unsuitable for pressure reclaim. ZFS zvol swap has documented
> deadlocks under memory pressure [3]; compressed swap is another target
> because writes may need memory despite free logical slots.
What is the high-level user-visible impact of this series?
I am curious: if we never run out of swap file space on the
non-offload swap area, does that mean we don't need this patch series?
The series discusses proactive reclaim and direct reclaim. The
offload-only swap area is for proactive reclaim. However, direct
reclaim can use both types of swap areas. Wouldn't using the
non-offload case cause memory allocation which you want to avoid?
I was wondering if we should actually flip things around. Mark the
type of swap device that avoids memory allocation during swap out.
Then direct reclaim prioritizes using those.That would make the happy
path allocate less memory during direct reclaim. Right now this series
does not seem to guarantee that direct reclaim stays away from the
area that allocates memory.
Chris
>
> Explicit proactive reclaim can use these areas alongside conventional
> swap, ordered by priority. Ordinary reclaim cannot initiate non-zero
> writes to them, including through retained swap entries. Conventional
> capacity is not reserved for emergencies. Operators choose the policy;
> the kernel does not measure headroom or make an allocating backend safe.
>
> TMO/Senpai [4] and DAMON_RECLAIM [5] are related cold-page reclaim
> approaches. This RFC admits offload-only swap through memory.reclaim and
> per-node reclaim, but not DAMON. Other related work includes per-cgroup
> zswap writeback control [6], Virtual Swap Space v4 [1] and swap tiers v10
> [2].
>
> The util-linux companion [8] proposes swapon --offload-only and an fstab
> option. Older swapon silently ignores the fstab token, so persistent
> activation needs discussion. Apply the separate i915 fix [9] first to avoid
> a pre-existing folio-lock leak when shmem writeback is skipped.
>
> The series fixes zram selftest device tracking and error reporting, adds
> the swap policy with DRM eligibility checks, and tests routing, workingset
> activation and retained-entry write refusal. It is based on mm-new at
> 995829088503.
>
> Feedback is particularly welcome on:
>
> - Activation-time eligibility versus backend placement or migration:
> is retained-entry refusal, with possible reclaim churn or OOM,
> acceptable without migration?
> - Eligible-capacity accounting, including overcommit limits and OOM
> scoring; cache recovery, cluster invalidation and workingset activation.
> - Persistent activation and visibility: current swap listings do not
> expose the policy.
>
> Validation includes x86 builds, affected arm64 MTE objects and builds
> without swap or memory cgroups. QEMU tests covered routing, retained-entry
> recovery, data integrity and cleanup, with negative controls for write
> refusal and teardown failure.
>
> I have also been using this policy on my own machine with experimental
> bcachefs swap support, with no problems observed so far. Deliberate stress
> testing is confined to VMs; I do not deliberately stress-test this machine.
>
> Changes since v1, incorporating Sashiko's public review [7]:
>
> - Improve selftest isolation, retained-swap measurement and memlock skips.
> - Track allocated zram devices, wait for udev probes and report cleanup
> errors without deleting data through a mount that could not be released.
> - Clarify activation, capacity reporting and conventional fallback limits.
>
> [1] https://patchew.org/linux/20260825153238.2695446-1-nphamcs@gmail.com/
> [2] https://patchew.org/linux/20260713025644.170839-1-youngjun.park@lge.com/
> [3] https://openzfs.github.io/openzfs-docs/Project%20and%20Community/FAQ.html#using-a-zvol-for-a-swap-device-on-linux
> [4] https://engineering.fb.com/2022/06/20/data-infrastructure/transparent-memory-offloading-more-memory-at-a-fraction-of-the-cost-and-power/
> [5] https://www.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
> [6] https://www.kernel.org/doc/html/latest/admin-guide/mm/zswap.html
> [7] https://sashiko.dev/#/patchset/20260914104501.3960616-1-matthias.goergens@gmail.com
> [8] https://github.com/util-linux/util-linux/pull/4633
> [9] https://lore.kernel.org/all/20260914105606.3997649-1-matthias.goergens@gmail.com/
>
> v1: https://lore.kernel.org/all/20260914104501.3960616-1-matthias.goergens@gmail.com/
>
> Matthias Goergens (4):
> selftests: zram: track owned devices and report cleanup failures
> mm: restrict offload-only swap to proactive reclaim
> selftests: zram: cover offload-only swap policy
> selftests: zram: cover retained offload-only entries
>
> Documentation/mm/swap.rst | 105 ++++
> drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
> .../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
> drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
> drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
> drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
> drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
> include/linux/swap.h | 29 +-
> include/linux/vm_event_item.h | 1 +
> mm/memcontrol.c | 20 +-
> mm/page_io.c | 24 +-
> mm/swapfile.c | 123 ++++-
> mm/vmscan.c | 31 +-
> mm/vmstat.c | 1 +
> tools/testing/selftests/zram/.gitignore | 2 +
> tools/testing/selftests/zram/Makefile | 4 +-
> tools/testing/selftests/zram/README | 20 +-
> tools/testing/selftests/zram/config | 10 +-
> tools/testing/selftests/zram/settings | 1 +
> tools/testing/selftests/zram/swap_offload.c | 425 +++++++++++++++++
> .../selftests/zram/workingset_offload.c | 207 ++++++++
> tools/testing/selftests/zram/zram.sh | 9 +
> tools/testing/selftests/zram/zram01.sh | 9 +-
> tools/testing/selftests/zram/zram02.sh | 9 +-
> tools/testing/selftests/zram/zram03.sh | 174 +++++++
> tools/testing/selftests/zram/zram04.sh | 164 +++++++
> tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++
> tools/testing/selftests/zram/zram_lib.sh | 189 ++++++--
> 28 files changed, 1926 insertions(+), 94 deletions(-)
> create mode 100644 tools/testing/selftests/zram/settings
> create mode 100644 tools/testing/selftests/zram/swap_offload.c
> create mode 100644 tools/testing/selftests/zram/workingset_offload.c
> create mode 100755 tools/testing/selftests/zram/zram03.sh
> create mode 100755 tools/testing/selftests/zram/zram04.sh
> create mode 100755 tools/testing/selftests/zram/zram05.sh
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
` (3 preceding siblings ...)
2026-09-28 0:19 ` Chris Li
@ 2026-09-28 5:19 ` Christoph Hellwig
2026-09-28 15:19 ` Matthias Goergens
4 siblings, 1 reply; 13+ messages in thread
From: Christoph Hellwig @ 2026-09-28 5:19 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J . Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest
On Sat, Sep 26, 2026 at 12:55:13PM +0800, Matthias Goergens wrote:
> This RFC adds offload-only swap areas for deliberate cold-page offload to
> backends unsuitable for pressure reclaim. ZFS zvol swap has documented
> deadlocks under memory pressure [3]; compressed swap is another target
> because writes may need memory despite free logical slots.
Out of tree code not matter. Out of tree that violates the linux
license terms even less.
Either way block devices must be safe to be called from paging paths,
not just for swap but also for file system based paging. So this is a
bug in the implementation, and not something worked around in the swap
code.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-28 0:19 ` Chris Li
@ 2026-09-28 15:19 ` Matthias Goergens
2026-09-29 0:27 ` Chris Li
0 siblings, 1 reply; 13+ messages in thread
From: Matthias Goergens @ 2026-09-28 15:19 UTC (permalink / raw)
To: Chris Li
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx, dri-devel,
Rafael J . Wysocki, Pavel Machek, Catalin Marinas, Will Deacon,
linux-pm, linux-arm-kernel, Shuah Khan, linux-kselftest
Hi Chris,
Thanks for resolving the conflicts yourself and reviewing it.
> What is the high-level user-visible impact of this series?
Being able to use more interesting swap backends and logic when the
machine is not under memory pressure. It is much easier to write a
swap backend that may occasionally allocate memory than one that is
guaranteed never to, so today such backends are either unsafe as swap
or ruled out (btrfs, for example, refuses swapfiles that are
copy-on-write, checksummed or compressed).
> I am curious: if we never run out of swap file space on the
> non-offload swap area, does that mean we don't need this patch series?
No: running out of conventional swap isn't the point. Without the
series, any active swap area may be written under pressure, so a
backend that may allocate can't be used as swap at all, however much
conventional swap there is next to it.
> However, direct reclaim can use both types of swap areas.
In v2 it can't: only memory.reclaim, per-node reclaim and MGLRU's
debugfs eviction may write new data to an offload-only area; direct
reclaim, kswapd, MADV_PAGEOUT and DAMON reclaim may not. But I think
your suggestion to flip it round is closer to what I want: mark the
areas whose writes may allocate, and keep pressure reclaim away from
those, rather than tying it to what started the reclaim. I'll work
that into v3, together with Kairui's suggestion to build on the swap
tiers work.
Thanks,
Matthias
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-28 5:19 ` Christoph Hellwig
@ 2026-09-28 15:19 ` Matthias Goergens
0 siblings, 0 replies; 13+ messages in thread
From: Matthias Goergens @ 2026-09-28 15:19 UTC (permalink / raw)
To: Christoph Hellwig
Cc: Andrew Morton, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J . Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest
Hi Christoph,
> Either way block devices must be safe to be called from paging paths,
> not just for swap but also for file system based paging. So this is
> a bug in the implementation, and not something worked around in the
> swap code.
Agreed that a backend used for paging has to make progress under
pressure, and the zvol example was a poor lead. The series isn't
meant to excuse a backend that can deadlock; that would still be a bug
to fix in the backend.
What I'm after is different: backends that are correct but need
memory to accept a write, where the better policy is to use them for
cold-page offload when memory isn't tight and keep pressure reclaim on
areas that don't need to allocate. Two in-tree examples: zram
allocates on writes and fails with -ENOMEM, with no fallback, while its
logical size still shows free slots; and filesystem swapfiles must
give up copy-on-write, checksums and compression today, because swap
writes bypass the filesystem. I'll make that the case for v3 instead.
Thanks,
Matthias
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 10:16 ` Kairui Song
@ 2026-09-28 15:19 ` Matthias Goergens
2026-10-04 18:13 ` Youngjun Park
0 siblings, 1 reply; 13+ messages in thread
From: Matthias Goergens @ 2026-09-28 15:19 UTC (permalink / raw)
To: Kairui Song
Cc: Andrew Morton, Chris Li, Youngjun Park, Baoquan He,
Johannes Weiner, David Hildenbrand, Michal Hocko, Shakeel Butt,
Kemeng Shi, Nhat Pham, Yosry Ahmed, Barry Song, linux-mm,
linux-kernel, cgroups, linux-api, Alejandro Colomar, linux-man,
Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx, dri-devel,
Rafael J . Wysocki, Pavel Machek, Catalin Marinas, Will Deacon,
linux-pm, linux-arm-kernel, Shuah Khan, linux-kselftest,
Kairui Song
Hi Kairui,
> Just like we already have a "swappiness=" parameter in the reclaim
> interface, perhaps adding a "swap.tier=" would be better and much
> cleaner to achieve the same goal? Based on Youngjun's work here:
> https://lore.kernel.org/all/20260916183437.2946306-1-youngjun.park@lge.com/
Thanks, I think that's the right base. I tried it: on top of
Youngjun's v11, a "swap.tier=" argument to memory.reclaim is about 40
lines. In a VM with zram at a higher priority than a disk swap
device, and a cgroup whose tier mask allowed only the disk, ordinary
reclaim in that cgroup stayed on the disk, while memory.reclaim with
"swap.tier=" pointing at zram went to zram only.
What it doesn't cover yet is reclaim outside any configured cgroup:
under global pressure, a cgroup nobody configured still reached the
zram tier. So for v3 I'd like to add a system-wide limit on which
tiers pressure reclaim may use, on top of the per-cgroup settings.
I'll wait for Youngjun's v12 and build on that.
Thanks,
Matthias
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-27 17:33 ` Andy Lutomirski
@ 2026-09-28 15:24 ` Matthias Goergens
0 siblings, 0 replies; 13+ messages in thread
From: Matthias Goergens @ 2026-09-28 15:24 UTC (permalink / raw)
To: Andy Lutomirski
Cc: Andrew Morton, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Vivi Rodrigo, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest
Hi Andy,
Thanks for reading it, and for the direct feedback.
> if the system is under overall memory pressure, why do you care what
> triggered the particular swap operation that is being processed?
You're right, and that question gets at what I actually want better
than the series does. The goal is: when there is no memory pressure,
swap-out may do more work, including allocating memory (compression,
copy-on-write, filesystem-backed swap), because moving cold pages out
to make room for page cache is what swap is for most of the time.
Under real pressure, swap-out must not depend on that.
The RFC was the smallest change I could think of in that direction.
It used "who started the reclaim" as a stand-in for "is there
pressure", and that is wrong both ways: memory.reclaim during global
pressure may allocate, while kswapd with plenty of free memory may
not. Building on Kairui's suggestion in this thread (swap tiers), or
on virtual swap, may well be the better route, and I'm looking at both
for v3.
> Why can't the kernel figure this out itself?
It can: the backend knows whether its writes may allocate (zram knows
it compresses, a filesystem knows at swapon whether the swapfile is
copy-on-write or compressed), so it should declare that, rather than
the administrator setting a flag.
On the practical questions: v2 does nothing to balance fullness
between areas or to empty conventional swap, and does not change OOM
selection. Those are fair gaps and I'll address them, or say
explicitly what is out of scope, in v3.
On the recursion question: within the reclaiming task it can't
happen. memory.reclaim and per-node reclaim both run with PF_MEMALLOC
set (memalloc_noreclaim_save() in try_to_free_mem_cgroup_pages() and
__node_reclaim()), so an allocation the backend makes in that task
never enters direct reclaim; it either fails or, unless it passes
__GFP_NOMEMALLOC, dips into the reserves. Two things do need care. A
backend allocating there can drain the emergency reserves unless it
passes __GFP_NOMEMALLOC or fails fast. And work the backend hands to
another thread, such as a filesystem or zvol worker, runs without
PF_MEMALLOC and can enter reclaim itself; that is only safe if that
reclaim never waits for the write the worker is meant to complete.
> The remainder of the writeup is IMO somewhat incoherent.
Agreed. I cut the v1 cover letter down too far and lost the
definitions it depended on. v3's will start from the goal above and
define its terms.
> Adding "eligible" to a bunch of calls does nothing to explain what's
> going on.
It exists because v2 made "how much swap is free" depend on who asks.
Reclaim itself checks free swap before scanning anonymous pages
(can_reclaim_anon_pages(), MGLRU's get_swappiness()) and before
allocating a slot (folio_alloc_swap()), and callers such as the GPU
shrinkers check it before pushing objects towards swap. With an
offload-only area, each of those would otherwise count space that the
calling reclaim may not use. In v3 I'd keep get_nr_swap_pages()
meaning "usable by ordinary reclaim", so those checks stay as they are
and only the offload path asks for a different count.
Thanks,
Matthias
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-28 15:19 ` Matthias Goergens
@ 2026-09-29 0:27 ` Chris Li
0 siblings, 0 replies; 13+ messages in thread
From: Chris Li @ 2026-09-29 0:27 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx, dri-devel,
Rafael J . Wysocki, Pavel Machek, Catalin Marinas, Will Deacon,
linux-pm, linux-arm-kernel, Shuah Khan, linux-kselftest
On Mon, Sep 28, 2026 at 5:19 AM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
>
> Hi Chris,
>
> Thanks for resolving the conflicts yourself and reviewing it.
>
> > What is the high-level user-visible impact of this series?
>
> Being able to use more interesting swap backends and logic when the
> machine is not under memory pressure. It is much easier to write a
> swap backend that may occasionally allocate memory than one that is
> guaranteed never to, so today such backends are either unsafe as swap
I actually don't know of a swap backend that absolutely will not
allocate memory on swap out yet. Some backends allocate more than
others. Because the proactive reclaim can write to the non-offloaded
swap backend. That brings me back to my original question, in what way
this series helps.
So the answer seems to be that previously, some backends were unusable
by swap, due to possible memory allocation. With this series, those
back ends are usable for the proactive reclaim now. It is a 0 to 0.5
improvement because direct reclaim can't use it yet.
> or ruled out (btrfs, for example, refuses swapfiles that are
> copy-on-write, checksummed or compressed).
>
> > I am curious: if we never run out of swap file space on the
> > non-offload swap area, does that mean we don't need this patch series?
>
> No: running out of conventional swap isn't the point. Without the
> series, any active swap area may be written under pressure, so a
> backend that may allocate can't be used as swap at all, however much
> conventional swap there is next to it.
But proactive reclaim can also use the non-offload swap area. So if we
have plenty of non-offload swap area, both proactive reclaim and
direct reclaim can use that non-offload area. The value of this series
isn't apparent. If you don't have a traditional swap-capable backend,
then your non-offload area is zero. That is considered a special case
of running out of non-offload swap area: never having a traditional
swap area in the first place.
>
> > However, direct reclaim can use both types of swap areas.
>
> In v2 it can't: only memory.reclaim, per-node reclaim and MGLRU's
Sorry I meant proactive reclaim can use both types, nothing constrains
proactive reclaim.
> debugfs eviction may write new data to an offload-only area; direct
> reclaim, kswapd, MADV_PAGEOUT and DAMON reclaim may not. But I think
> your suggestion to flip it round is closer to what I want: mark the
Flipping it around might provide additional benefit. I know some users
maintain off tree patches to turn off zswap on the direct reclaim path
exactly because zswap might allocate more memory before it can free
some. If the kernel can be smart about it. It provides additional
value to the status quo.
> areas whose writes may allocate, and keep pressure reclaim away from
> those, rather than tying it to what started the reclaim. I'll work
> that into v3, together with Kairui's suggestion to build on the swap
> tiers work.
Yes, I feel that cluster-level marking might not be needed, you can
remember which si has the offload flag. However the per cpu cache
needs to know which context might require using a different cached
cluster. That is the trickiest part. The rest of the patch is just
enough plumbing to preserve whether the swap-out context is proactive
or not.
Chris
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-28 15:19 ` Matthias Goergens
@ 2026-10-04 18:13 ` Youngjun Park
0 siblings, 0 replies; 13+ messages in thread
From: Youngjun Park @ 2026-10-04 18:13 UTC (permalink / raw)
To: Matthias Goergens
Cc: Kairui Song, Andrew Morton, Chris Li, Baoquan He, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx, dri-devel,
Rafael J . Wysocki, Pavel Machek, Catalin Marinas, Will Deacon,
linux-pm, linux-arm-kernel, Shuah Khan, linux-kselftest,
Kairui Song
On 2026-09-28 23:19, Matthias Goergens wrote:
Hi Matthias !
> > Just like we already have a "swappiness=" parameter in the reclaim
> > interface, perhaps adding a "swap.tier=" would be better and much
> > cleaner to achieve the same goal? Based on Youngjun's work here:
> > https://lore.kernel.org/all/20260916183437.2946306-1-youngjun.park@lge.com/
>
> Thanks, I think that's the right base. I tried it: on top of
> Youngjun's v11, a "swap.tier=" argument to memory.reclaim is about 40
> lines. In a VM with zram at a higher priority than a disk swap
> device, and a cgroup whose tier mask allowed only the disk, ordinary
> reclaim in that cgroup stayed on the disk, while memory.reclaim with
> "swap.tier=" pointing at zram went to zram only.
There is an earlier attempt that may be worth a look.
https://lore.kernel.org/all/20260618044857.69439-1-jiahao.kernel@gmail.com/
If you go this way, the discussion in that thread worth refering too.
> What it doesn't cover yet is reclaim outside any configured cgroup:
> under global pressure, a cgroup nobody configured still reached the
Right, so for now every cgroup has to be set to exclude zram, unless
another idea or more code covers it.
> zram tier. So for v3 I'd like to add a system-wide limit on which
> tiers pressure reclaim may use, on top of the per-cgroup settings.
> I'll wait for Youngjun's v12 and build on that.
I have looked through the whole series, and I will keep this use case
in mind while working on v12. :)
Also, to share what I had in mind, I was planning a sysfs interface
for each swap tier, /sys/kernel/mm/swap/tiers/<tier name> (exact
naming TBD). Maybe the system-wide limit could be handled there as one
of its use cases?
Let's discuss it in more detail after v12.
Thanks,
Youngjun
^ permalink raw reply [flat|nested] 13+ messages in thread
end of thread, other threads:[~2026-10-04 18:14 UTC | newest]
Thread overview: 13+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
2026-09-26 10:16 ` Kairui Song
2026-09-28 15:19 ` Matthias Goergens
2026-10-04 18:13 ` Youngjun Park
2026-09-26 23:32 ` Chris Li
2026-09-27 0:03 ` Chris Li
2026-09-27 17:33 ` Andy Lutomirski
2026-09-28 15:24 ` Matthias Goergens
2026-09-28 0:19 ` Chris Li
2026-09-28 15:19 ` Matthias Goergens
2026-09-29 0:27 ` Chris Li
2026-09-28 5:19 ` Christoph Hellwig
2026-09-28 15:19 ` Matthias Goergens
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox