From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: 天狼 <rockeet@gmail.com>
Cc: "David Hildenbrand (Arm)" <david@kernel.org>,
linux-mm@kvack.org, linux-fsdevel@vger.kernel.org,
"Liam R. Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Jann Horn <jannh@google.com>,
Matthew Wilcox <willy@infradead.org>, Jan Kara <jack@suse.cz>
Subject: Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
Date: Wed, 30 Sep 2026 13:08:33 +0100 [thread overview]
Message-ID: <arz0auULA33fqoWu@gremlin> (raw)
In-Reply-To: <CAAE3jtftxMkCfWJ77W-xvV2x7jdAN8C8kyofYC3fEivM-Cq6UQ@mail.gmail.com>
On Wed, Sep 30, 2026 at 06:27:14PM +0800, 天狼 wrote:
> > You really need an actual file on disk. Shared memory / shmem / memfd is
> > not sufficient?
>
> Yes, an ordinary disk-backed file is the intended output, and subsequent
> consumers access it through mmap.
>
> Shared memory can preserve state across a process crash if the shared
> object remains alive. However, using shmem/memfd would introduce a separate
> staging and transfer step to put the data into the final file. Building
> directly in a file-backed MAP_SHARED mapping lets us preserve the latest
> userspace writes across process crashes and retain the same file-cache
> pages for subsequent consumers.
In general, writing to files through MAP_SHARED mappings is discouraged.
>
> > I also assume MAP_PRIVATE mapping the file so it's anon CoW'd is not
> > sufficient either?
>
> Correct. With MAP_PRIVATE, modifications remain private and are not carried
> through to the underlying file. If the constructing process crashes before
> transferring those modifications to the file, another process reopening the
> file cannot access them. That would lose the process-crash survival
> property we obtain from MAP_SHARED.
But you can solve that with a memfd?
>
> > It'd be problematic to effectively corrupt the page cache by saying "hey
> > this is dirty but just clear dirty state and pretend this is what the
> > disk has".
>
> To clarify, the proposal does not involve clearing dirty state or
> pretending that memory and disk contents match. The pages would remain
> dirty and continue to count toward the existing dirty-page limits.
>
> The requested hint only expresses that the application expects to modify
> the data again soon and would prefer background writeback to happen later.
But that's the same as saying 'don't writeback'?
When there's a lot of dirtied pages on disk writes are delayed in proportion to
how dirtied with various other heuristics taken into account.
I don't see how this requested feature wouldn't just give programs a way to work
around that, meanwhile everything else doing dirty writeback are _impacted by
the dirtying that the requesting process has done_.
That seems really problematic.
>
> > Delaying writeback could interfere with the writeback balancing logic.
>
> I understand that concern. The hint should not exempt the application from
> dirty-page limits or throttling. If writeback is needed for balancing,
> memory pressure, or explicit synchronization, the kernel should override
> the hint.
Hmm.
The delay before background writeback happens is, by default, quite long,
until you have a lot of dirtied pages.
dirty_expire_centisecs defaults to 3000, i.e. 30s before it is considered
for writeback.
And then that only happens every dirty_writeback_centisecs, defaults to
500, i.e. every 5 seconds.
If you're not writing to the mapping within 5 seconds even, let alone 30
seconds then that seems like an issue with your program.
I guess you are re-dirtying again with later changes, but the _triggering_
of writeback is at an inode granularity on dirtying - so you want to delay
writing back some pages because later pages are not meaningful?
All the while you are adding to the total dirty pages but not paying the
price anywhere, nor allowing the dirty page count to reduce.
Unfortunately there's no real way to avoid the per-inode thing, and
delaying valid writeback there because SOME of it isn't 'valid' is not
really sensible.
Either that, or you are actually concerned when there are enough
dirty pages to trigger background writeback i.e. dirty_background_ratio is
exceeded (where the above limits don't apply).
But later you say you don't want the proposed change to impact how that
behaves (the balance_dirty_pages() logic), which contradicts that.
So are you taking ~35s to actually write to this properly?
And is it that you're wanting the _triggering_ of writeback to be
per-folio? That isn't really sensible.
And nor is it really to delay writeback for dirty folios because later ones
are 'invalid'.
>
> The intended benefit is to avoid writing intermediate versions when the
> kernel has room to defer that work. There is no requirement to prevent
> writeback until construction finishes.
>
> madvise is not the only option, deferring writeback for the entire file
> would also be acceptable.
>
> Could such an advisory preference be accommodated within the existing
> writeback balancing policy while retaining its limits and fairness?
I think the fundamental friction here is that you're charging a cost to a
global limit then doing something that prevents that cost being paid (you
can't writeback), and meanwhile the process can wait for arbitrary time
before having to write to it.
And the fact that it is taking you so long to actually write what is
intended there that this matters suggests to me there's something wrong
with what you're doing here.
I'd say in general the shape of the solution would be something
cgroups-ish, but since all of this is global stuff and the knobs are all
global I think you need to think about your design a little more.
Having a process that is kept around maybe with an appropriate OOM score
that issues a memfd would get you the sharing even if process dies stuff,
and then it can periodically write to a file when necessary.
I think that'd work a lot better, and I don't really see a kernel change
here that would make sense.
All in all my intuition is that you're looking for a kernel solution to
something that should be implemented differently in userland.
And in general we've seen a pattern of 'writes to MAP_SHARED, experiences
problems' before.
And the solution is generally - use something like memfd, have a process
that does writebacks manually for you as you need.
So I strongly suggest you do that! :)
--
Cheers, Lorenzo
next prev parent reply other threads:[~2026-09-30 12:08 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09 ` 天狼
2026-09-28 19:12 ` David Hildenbrand (Arm)
2026-09-29 4:06 ` 天狼
2026-09-29 6:59 ` David Hildenbrand (Arm)
2026-09-29 8:59 ` Lorenzo Stoakes (ARM)
2026-09-30 10:27 ` 天狼
2026-09-30 11:38 ` David Hildenbrand (Arm)
2026-09-30 12:08 ` Lorenzo Stoakes (ARM) [this message]
2026-09-30 12:11 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14 ` 天狼
2026-09-30 14:24 ` Lorenzo Stoakes (ARM)
2026-09-30 15:33 ` Lorenzo Stoakes (ARM)
2026-09-30 15:29 ` 天狼
2026-09-30 15:43 ` Lorenzo Stoakes (ARM)
2026-09-30 16:38 ` Jan Kara
2026-10-01 4:28 ` 天狼
2026-10-01 8:35 ` Lorenzo Stoakes (ARM)
2026-10-01 13:21 ` 天狼
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arz0auULA33fqoWu@gremlin \
--to=ljs@kernel.org \
--cc=david@kernel.org \
--cc=jack@suse.cz \
--cc=jannh@google.com \
--cc=liam@infradead.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=rockeet@gmail.com \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).