Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: 天狼 <rockeet@gmail.com>,
	linux-mm@kvack.org, linux-fsdevel@vger.kernel.org,
	"Liam R. Howlett" <liam@infradead.org>,
	"Vlastimil Babka" <vbabka@kernel.org>,
	"Jann Horn" <jannh@google.com>,
	"Matthew Wilcox" <willy@infradead.org>, "Jan Kara" <jack@suse.cz>
Subject: Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
Date: Tue, 29 Sep 2026 09:59:25 +0100	[thread overview]
Message-ID: <art8EKRKMdwSDCkr@gremlin> (raw)
In-Reply-To: <1ce9cf3a-15dc-4815-ba29-2e22b05206fb@kernel.org>

On Tue, Sep 29, 2026 at 08:59:57AM +0200, David Hildenbrand (Arm) wrote:
> On 9/29/26 06:06, 天狼 wrote:
> > Hi David,
> >
>
> Hi,
>
> >> What is supposed to happen if the process crashes when updating the
> >> MAP_SHARED region halfway through?
> >
> > The data structure survives an interrupted update through its
> > application-level copy-on-write and lock-free concurrency design. A
> > file-backed MAP_SHARED mapping preserves the latest userspace writes after
> > the process dies, even if they have not yet reached disk, provided the
> > kernel remains running.
> >
> > This does not depend on preventing writeback. Deferring writeback is purely
> > a performance optimization, not a correctness requirement.
> >
> >> It would be helpful if the use case + data structure would be explained
> >> in a bit more detail.
> >
> > The data structure originally used uint32_t offsets instead of pointers to
> > reduce pointer overhead. On that basis, we implemented lock-free concurrent
> > reads and writes, obtaining crash safety for free at the same time. We use
> > a file-backed MAP_SHARED mapping to turn that capability into an actual
> > feature that survives process crashes.
>
> Ok, thanks. Just be sure: you really need an actual file on disk. Shared memory
> / shmem / memfd is not sufficient?
>

I also assume MAP_PRIVATE mapping the file so it's anon CoW'd is not sufficient
either?

> >> What exactly is the problem with writeback here? Unnecessary I/O? Is
> >> writeback the problem or actual reclaim after writeback?
> >
> > Unnecessary I/O. Reclaim after writeback is not the problem I am trying to
> > address.
>
> Ok.

I mean if you're mmap'ing it you're literally mapping the page cache folio and
dirtying that folio.

So the kernel really does have to do writeback and it might be quite problematic
trying to prevent the writeback algorithm from doing its work on a specific
range.

I think the shape of a viable solution really is either MAP_PRIVATE-mapping it
or using some anon shmem/memfd as David suggests.

I think it'd be problematic to effectively corrupt the page cache by saying 'hey
this is dirty but just clear dirty state and pretend this is what the disk has'
or something.


>
> >
> > During construction, background writeback can write intermediate contents
> > that the application will soon overwrite. Those repeated writes waste disk
> > bandwidth. Delaying writeback would avoid that waste where possible; it is
> > not required for correctness.
>
> Understood.

Yeah again delaying writeback could interfere with the writeback balancing logic
which tries to keep dirty page levels sane and is fair and balanced so processes
that writeback a lot get delayed in doing so under heavy dirtying.

I think anything like this would interfere with that.

>
> >
> >> In a fuse server you can in theory delay the writeback request.
> >
> > Thanks for the suggestion. That sounds like a possible way to experiment
> > with delayed writeback. The feature I am requesting would make this
> > advisory behavior available to applications using ordinary file-backed
> > shared mappings.
> >
> >> So fadvise would be an option. However, this "defer mode" is really odd.
> >> It sounds more like you would want to have a custom policy there [...]
> >
> > madvise is not the only option, deferring writeback for the entire file
> > would also be acceptable.
> >
> > The application only wants to indicate that it is still modifying the data
> > and would prefer background writeback to happen later. The kernel would
> > retain control over the actual timing. Memory pressure could override the
> > hint, and explicit synchronization would keep its normal semantics.
> >
> > The hint should also expire automatically when the owning mapping or handle
> > is released, including on process exit.
>
> Ok, so while you are updating the large mmap'ed file concurrently, you don't
> want writeback to go crazy, because you know that you will modify the memory
> immediately anyway.

Yeah see above, I really think the only sensible solution is a MAP_PRIVATE CoW'd
mapping or memfd etc.

>
> Let me CC some more people.

Christian and probably Willy also? But I'm not so sure there's anything sensible
to do here other than something-anon.

>
>
> --
> Cheers,
>
> David

--
Cheers, Lorenzo


  reply	other threads:[~2026-09-29  8:59 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28  7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09   ` 天狼
2026-09-28 19:12     ` David Hildenbrand (Arm)
2026-09-29  4:06       ` 天狼
2026-09-29  6:59         ` David Hildenbrand (Arm)
2026-09-29  8:59           ` Lorenzo Stoakes (ARM) [this message]
2026-09-30 10:27             ` 天狼
2026-09-30 11:38               ` David Hildenbrand (Arm)
2026-09-30 12:08               ` Lorenzo Stoakes (ARM)
2026-09-30 12:11                 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14                   ` 天狼
2026-09-30 14:24                     ` Lorenzo Stoakes (ARM)
2026-09-30 15:33                       ` Lorenzo Stoakes (ARM)
2026-09-30 15:29                     ` 天狼
2026-09-30 15:43                       ` Lorenzo Stoakes (ARM)
2026-09-30 16:38                       ` Jan Kara
2026-10-01  4:28                         ` 天狼
2026-10-01  8:35                           ` Lorenzo Stoakes (ARM)
2026-10-01 13:21                             ` 天狼

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=art8EKRKMdwSDCkr@gremlin \
    --to=ljs@kernel.org \
    --cc=david@kernel.org \
    --cc=jack@suse.cz \
    --cc=jannh@google.com \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=rockeet@gmail.com \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox