From: Kairui Song <ryncsn@gmail.com>
To: Youngjun Park <youngjun.park@lge.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,
"Rafael J. Wysocki" <rafael@kernel.org>,
Kairui Song <kasong@tencent.com>, Chris Li <chrisl@kernel.org>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Baoquan He <baoquan.he@linux.dev>,
Barry Song <baohua@kernel.org>, Pavel Machek <pavel@kernel.org>,
Len Brown <lenb@kernel.org>,
linux-mm@kvack.org, linux-pm@vger.kernel.org,
her0gyugyu@gmail.com, taejoon.song@lge.com
Subject: Re: [RFC PATCH 00/10] mm/swap, PM: hibernate: improve image slot allocation and I/O
Date: Wed, 30 Sep 2026 01:25:54 +0800 [thread overview]
Message-ID: <aruGQpdsq7_O-AC8@KASONG-MC4> (raw)
In-Reply-To: <20260915031658.1505680-1-youngjun.park@lge.com>
On Tue, Sep 15, 2026 at 12:16:48PM +0800, Youngjun Park wrote:
> This series improves how hibernation allocates swap slots for its
> image. It starts from a few observations about what happens while
> the image is written. With them the allocator can hand the image
> contiguous runs, the image I/O can be batched per run, and
> hibernation gets faster. Contiguous I/O pattern is friendly to flash device also.
>
> Slot allocation matters for hibernation speed. After commit
> 0ff67f990bd4 ("mm, swap: remove swap slot cache") in v6.15, writing
> the image was about ten times slower on some SSDs [1], until
> commit 396f57b57200 ("mm, swap: speed up hibernation allocation and
> writeout") fixed it in v7.1.
>
> Note
> - I was responsible for the observation and design,
> while receiving substantial support from LLMs throughout the series.
> - Submitting this patch to confirm whether this work is progressing or not.
Hi Youngjun
Thanks for the patch!
Splitting out the hibernation specific part out of swapfile.c
looks a good idea to me. We can then compile that file conditionally
rather than having a huge ifdef block in swapfile.c. Would be nice
to get more feedback from hibnernation side.
>
> Observations
> ============
>
> 1. While the image is written, swap can only hand out slots that are
> already free. Swap cache cannot be reclaimed to make more.
>
> folio_swapcache_freeable() refuses every folio while storage is
> suspended. The image may already hold a folio as clean swap cache.
> If its slot were freed and reused for the image, the resumed kernel
> could later drop that folio and read it back from the slot, now
> with the wrong data.
>
> Storage is suspended before the image is written, so every
> allocation for the image falls in that window.
>
> hibernate()
> freeze_processes() user space is frozen
> hibernation_snapshot()
> freeze_kernel_threads() kswapd is frozen
> hibernate_preallocate_memory() may swap out to shrink memory
> pm_restrict_gfp_mask() storage is suspended from here
> create_image() the snapshot is taken
> swsusp_write() image slots are allocated here
> power_down()
...
> 2. Image slots are not freed and reused while the image is written.
> And whatever the allocator changes during the write is gone at
> resume, because the resumed kernel is the snapshot.
>
> With observation 1, once the image owns a free cluster it can use
> all of it. Nothing has to be recorded per slot, neither in the
> swap table nor as a memcg id.
That's pretty nice indeed. But I'm a bit concerned about having a
specific cluster isolation path for hibernation.
Will it be cleaner to provide a more generic cluster sized
allocation (PMD sized) so common swap can benefit too?
And I'm not sure is IO batching working properly before?
How much performance gain is due to the cluster sized IO?
But if other appraoches won't work or hibernation is really
special, and we can keep all the hibernation tricky clean and
simple in just one place, maybe it's not too bad.
> 3. User space and kswapd do not swap out while the image is written,
> and allocations from the page allocator cannot start swap I/O.
> What is left is rare. DAMON pageout and the memcg high work can
> still reach swap(This is all I found. anything else?),
> and both were seen running in that window.
>
> So the allocator can favor the image then. Other users take slots
> from the nonfull and frag clusters and leave the free clusters to
> the image.
>
> With these the allocator gets simpler and faster, and batching the
> image read and write becomes easy.
>
> What the series does
> ====================
>
> 1-2 fixes, a slot leak after a failed test_resume and a NULL
> dereference for a swapfile with no block device (some bug fix)
> 3 move the hibernation code to mm/swap_hibernate.c (refactor)
> 4 skip swap cache reclaim while storage is suspended (optimization)
> 5-6 hand the image whole free clusters, in disk order (exploit contiguous space)
> 7-8 write and read the image one bio per run (batch I/O)
> 9-10 optional reservation at swapon, hibernate=reserve (assure contiguous space)
>
> Based on mm-new (383fc05d4650) with patches 2 to 4 of [3] under it.
> Patch 1 of [3] is in mm-new as 10d9012e83ef.
>
> Note. [3] gives hibernation slots their own swap table entry, keeps
> readahead off them, and frees them by offset alone. The single slot
> path of patch 5 builds on that.
Nice, maybe that series need a refresh to get merged first.
>
> Results
> =======
>
> Setup
> - qemu, 12G RAM, 4 CPUs, no KVM. Times only compare against each
> other.
> - swap on virtio-blk as a non-rotational device
> - image 5.0G, written with hibernate=nocompress
> - 3 reps of two hibernations each. Times are medians of the 4 to 6
> samples per cell that no host load hit.
> - base is patch 3 and allocates as mm-new does, allocator is
> patch 6, allocator + bio is patch 8
>
> Rows. runs is how many contiguous stretches of the device the image
> ends up in. bios is how many bios the kernel allocates and submits to
> write it. In both cases the image fits in free clusters, so the
> fallback to nonfull and frag clusters is not measured. Percentages
> are against base.
>
> Shuffled free list. A device that has been in use, emptied.
> - 6G swap, 5.4G of 2M tmpfs files swapped out, then all removed in
> random order
> - every cluster is free, the free list is in free order, not in
> disk order
>
> base allocator allocator + bio
> write, s 10.60 9.32 (-12%) 7.94 (-25%)
> read, s 9.15 8.92 (-3%) 8.10 (-11%)
> runs 3560 1 1
> bios 1.30M 1.30M 12.7K
>
> Holes in nonfull clusters. What taking free clusters first buys.
> - 12G swap, one 4G file swapped out, every other 64K of it freed
> - 2G of 64K holes in 2048 clusters, 8G of free clusters
> - mm-new fills the holes first, the series takes the free clusters
>
> base allocator allocator + bio
> write, s 10.69 10.07 (-6%) 8.81 (-18%)
> read, s 12.21 10.89 (-11%) 9.78 (-20%)
> runs 31899 1 1
> bios 1.30M 1.30M 12.8K
>
> Summary against base
> - the allocator cuts write time by 6 to 12%
> - allocator + bio cuts write time by 18 to 25% and read time by
> 11 to 20%
>
> These are VM numbers without compression. Compression, the default,
> and real hardware are still to be checked.
>
> Next steps
> ==========
>
> Things to keep working on after this RFC. Comments are welcome.
>
> 1. Dropping swap cache before hibernation starts, so more slots are
> free. This series does not do that.
>
> 2. Whether the reservation in patches 9 and 10 is worth keeping. It
> makes sure the image gets contiguous slots when swap has room to
> spare.
>
> 3. A block device of its own for hibernation instead of swap. Not
> taken for now. Sharing one device keeps the spare space useful, the
> existing infrastructure stays, and the ideas above give much the same
> effect.
So is the idea for 2 and 3 here to make sure hibernation always success by
avoid the allocator using too much for common swap?
We only want one of them I think, and we need to be careful here to not
make the maintainance messup by adding too many knobs...
>
> 4. Whether the extent tree can go. A normal hibernation never walks
> it, the swap state comes back as it was at the snapshot. It is only
> walked to free the slots after an error or a wake from hybrid sleep.
> With the slots marked in the swap table [3] and taken as whole
> clusters, a free could find them without it.
The extent tree is not a hibernation issue right? At least for block based
swap the extent tree is useless (only one node). I think swap_ops can be
used to make this limited to certain swap_ops (e.g. file swap ops).
And maybe, the swap_ops can provide some interface for hibenation usage to
make things cleaner?
>
> 5. Whether SNAPSHOT_ALLOC_SWAP_PAGE should refuse a request made before
> storage is suspended. Such a request gets single slots from the
> normal allocator today, and s2disk only asks after
> SNAPSHOT_CREATE_IMAGE anyway.
Is that a even a right thing to do during hibernation?
> 6. Two cases are not measured yet. A device with both a shuffled free
> list and partly used clusters. An image bigger than the free
> clusters, so part of it comes from nonfull and frag clusters.
I think that's fine, free cluster shuffle should not effect the
performance much as 2M is a pretty big IO unit. For the
fragmentation batching IO should be very helpful.
next prev parent reply other threads:[~2026-09-29 17:26 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-15 3:16 [RFC PATCH 00/10] mm/swap, PM: hibernate: improve image slot allocation and I/O Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 01/10] PM: hibernate: give the image's swap slots back when test_resume fails Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 02/10] mm, swap: skip swap devices without a block device in hibernation lookups Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 03/10] mm, swap: move hibernation swap code to mm/swap_hibernate.c Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 04/10] mm, swap: skip swap cache reclaim while storage is suspended Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 05/10] mm, swap: hand the hibernation image whole free clusters Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 06/10] mm, swap: hand the image's free clusters out in disk order Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 07/10] PM: hibernate: build one bio per contiguous run of the image Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 08/10] PM: hibernate: read the image back a run at a time Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 09/10] PM: hibernate: tell swap how much space an image needs Youngjun Park
2026-09-15 3:16 ` [RFC PATCH 10/10] mm, swap: hold swap space back for a hibernation image at swapon Youngjun Park
2026-09-29 17:25 ` Kairui Song [this message]
2026-10-04 17:39 ` [RFC PATCH 00/10] mm/swap, PM: hibernate: improve image slot allocation and I/O Youngjun Park
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aruGQpdsq7_O-AC8@KASONG-MC4 \
--to=ryncsn@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=chrisl@kernel.org \
--cc=her0gyugyu@gmail.com \
--cc=kasong@tencent.com \
--cc=lenb@kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-pm@vger.kernel.org \
--cc=nphamcs@gmail.com \
--cc=pavel@kernel.org \
--cc=rafael@kernel.org \
--cc=shikemeng@huaweicloud.com \
--cc=taejoon.song@lge.com \
--cc=youngjun.park@lge.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox