From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: 天狼 <rockeet@gmail.com>
Cc: Jan Kara <jack@suse.cz>,
"David Hildenbrand (Arm)" <david@kernel.org>,
linux-mm@kvack.org, linux-fsdevel@vger.kernel.org,
"Liam R. Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Jann Horn <jannh@google.com>,
Matthew Wilcox <willy@infradead.org>
Subject: Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
Date: Thu, 1 Oct 2026 09:35:25 +0100 [thread overview]
Message-ID: <ar4WSaLYBs1db1-3@gremlin> (raw)
In-Reply-To: <CAAE3jtexTRdqzAodzCkAy5NtH_s3VhtutXbxMjangf8BtB+-dQ@mail.gmail.com>
Hi Peng,
(Point of etiquette - and you are perhaps not aware - in general in the
kernel we reply inline to people rather than sending big lists in
reply.)
>1. With buffered I/O, the data already resides in shmem, and writing
>it to the regular file allocates another set of page-cache pages. When
>both copies are resident, this roughly doubles the memory occupied by
>the data. Traditional LSM MemTables are usually small, so duplicating
>one is relatively inexpensive (they waste so much elsewhere that they
>never even reach the scale where this becomes a concern). Our
>MemTables can be very large, making duplication a substantial waste of
>memory that significantly undermines the advantages of our design.
Then use O_DIRECT :)
>
>2. During the MemTable-to-SST transition, existing readers still
>access the data through the MemTable, while new readers access it
>through the SST. We therefore cannot immediately release the shmem
>copy. Our current ConvertToSST instead renames the existing file and
>appends metadata. The inode remains the same, so the existing MemTable
>mapping and the new SST mapping share the same page-cache pages
>throughout this overlap period.
That's your choice, not a requirement. The SST is the MemTable plus
metadata, so append the metadata in shmem and serve new readers from there
too.
>
>3. AIO or io_uring with direct I/O avoids allocating the destination
>page cache during the write, but it does not turn the source shmem
>pages into page-cache pages of the SST file. Subsequent SST mmap reads
>populate a separate page cache, while existing MemTable readers still
>need the shmem pages. The duplicate memory therefore remains a problem
>during the overlap period.
It doesn't need to. Serve everyone from shmem until existing readers drain,
then drop it and switch to the SST. Only one copy is ever resident.
>
>4. Temporary files are a much broader use case. Periodic automatic
>writeback is unnecessary for temporary data; writeback for memory
>reclaim is what serves a practical purpose. A per-file exemption from
>age-based periodic writeback would benefit these workloads as well.
Temporary data that doesn't need to reach disk belongs in shmem :)
Overall - each objection so far has been to the cost of changing your
design rather than to the userspace approaches not working.
I agree with Jan that the inode hint is viable in shape, but we'd need a
good reason to add it, and avoiding a design change isn't one :)
David, Jan and I have all suggested a shmem-based userspace
approach. Please explore that properly first.
I took the time to write up the approach, in detail, at
https://lore.kernel.org/linux-mm/ar0mdJKLC0cHBoKM@gremlin/ hopefully that's
helpful.
--
Cheers, Lorenzo
next prev parent reply other threads:[~2026-10-01 8:35 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09 ` 天狼
2026-09-28 19:12 ` David Hildenbrand (Arm)
2026-09-29 4:06 ` 天狼
2026-09-29 6:59 ` David Hildenbrand (Arm)
2026-09-29 8:59 ` Lorenzo Stoakes (ARM)
2026-09-30 10:27 ` 天狼
2026-09-30 11:38 ` David Hildenbrand (Arm)
2026-09-30 12:08 ` Lorenzo Stoakes (ARM)
2026-09-30 12:11 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14 ` 天狼
2026-09-30 14:24 ` Lorenzo Stoakes (ARM)
2026-09-30 15:33 ` Lorenzo Stoakes (ARM)
2026-09-30 15:29 ` 天狼
2026-09-30 15:43 ` Lorenzo Stoakes (ARM)
2026-09-30 16:38 ` Jan Kara
2026-10-01 4:28 ` 天狼
2026-10-01 8:35 ` Lorenzo Stoakes (ARM) [this message]
2026-10-01 13:21 ` 天狼
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ar4WSaLYBs1db1-3@gremlin \
--to=ljs@kernel.org \
--cc=david@kernel.org \
--cc=jack@suse.cz \
--cc=jannh@google.com \
--cc=liam@infradead.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=rockeet@gmail.com \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox