Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: 天狼 <rockeet@gmail.com>
Cc: Jan Kara <jack@suse.cz>,
	"David Hildenbrand (Arm)" <david@kernel.org>,
	 linux-mm@kvack.org, linux-fsdevel@vger.kernel.org,
	 "Liam R. Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	 Jann Horn <jannh@google.com>,
	Matthew Wilcox <willy@infradead.org>
Subject: Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
Date: Thu, 1 Oct 2026 09:35:25 +0100	[thread overview]
Message-ID: <ar4WSaLYBs1db1-3@gremlin> (raw)
In-Reply-To: <CAAE3jtexTRdqzAodzCkAy5NtH_s3VhtutXbxMjangf8BtB+-dQ@mail.gmail.com>

Hi Peng,

(Point of etiquette - and you are perhaps not aware - in general in the
 kernel we reply inline to people rather than sending big lists in
 reply.)

>1. With buffered I/O, the data already resides in shmem, and writing
>it to the regular file allocates another set of page-cache pages. When
>both copies are resident, this roughly doubles the memory occupied by
>the data. Traditional LSM MemTables are usually small, so duplicating
>one is relatively inexpensive (they waste so much elsewhere that they
>never even reach the scale where this becomes a concern). Our
>MemTables can be very large, making duplication a substantial waste of
>memory that significantly undermines the advantages of our design.

Then use O_DIRECT :)

>
>2. During the MemTable-to-SST transition, existing readers still
>access the data through the MemTable, while new readers access it
>through the SST. We therefore cannot immediately release the shmem
>copy. Our current ConvertToSST instead renames the existing file and
>appends metadata. The inode remains the same, so the existing MemTable
>mapping and the new SST mapping share the same page-cache pages
>throughout this overlap period.

That's your choice, not a requirement. The SST is the MemTable plus
metadata, so append the metadata in shmem and serve new readers from there
too.

>
>3. AIO or io_uring with direct I/O avoids allocating the destination
>page cache during the write, but it does not turn the source shmem
>pages into page-cache pages of the SST file. Subsequent SST mmap reads
>populate a separate page cache, while existing MemTable readers still
>need the shmem pages. The duplicate memory therefore remains a problem
>during the overlap period.

It doesn't need to. Serve everyone from shmem until existing readers drain,
then drop it and switch to the SST. Only one copy is ever resident.

>
>4. Temporary files are a much broader use case. Periodic automatic
>writeback is unnecessary for temporary data; writeback for memory
>reclaim is what serves a practical purpose. A per-file exemption from
>age-based periodic writeback would benefit these workloads as well.

Temporary data that doesn't need to reach disk belongs in shmem :)

Overall - each objection so far has been to the cost of changing your
design rather than to the userspace approaches not working.

I agree with Jan that the inode hint is viable in shape, but we'd need a
good reason to add it, and avoiding a design change isn't one :)

David, Jan and I have all suggested a shmem-based userspace
approach. Please explore that properly first.

I took the time to write up the approach, in detail, at
https://lore.kernel.org/linux-mm/ar0mdJKLC0cHBoKM@gremlin/ hopefully that's
helpful.

--
Cheers, Lorenzo


  reply	other threads:[~2026-10-01  8:35 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28  7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09   ` 天狼
2026-09-28 19:12     ` David Hildenbrand (Arm)
2026-09-29  4:06       ` 天狼
2026-09-29  6:59         ` David Hildenbrand (Arm)
2026-09-29  8:59           ` Lorenzo Stoakes (ARM)
2026-09-30 10:27             ` 天狼
2026-09-30 11:38               ` David Hildenbrand (Arm)
2026-09-30 12:08               ` Lorenzo Stoakes (ARM)
2026-09-30 12:11                 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14                   ` 天狼
2026-09-30 14:24                     ` Lorenzo Stoakes (ARM)
2026-09-30 15:33                       ` Lorenzo Stoakes (ARM)
2026-09-30 15:29                     ` 天狼
2026-09-30 15:43                       ` Lorenzo Stoakes (ARM)
2026-09-30 16:38                       ` Jan Kara
2026-10-01  4:28                         ` 天狼
2026-10-01  8:35                           ` Lorenzo Stoakes (ARM) [this message]
2026-10-01 13:21                             ` 天狼

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ar4WSaLYBs1db1-3@gremlin \
    --to=ljs@kernel.org \
    --cc=david@kernel.org \
    --cc=jack@suse.cz \
    --cc=jannh@google.com \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=rockeet@gmail.com \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox