Linux filesystem development
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: 天狼 <rockeet@gmail.com>, "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
Cc: linux-mm@kvack.org, linux-fsdevel@vger.kernel.org,
	"Liam R. Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>, Jann Horn <jannh@google.com>,
	Matthew Wilcox <willy@infradead.org>, Jan Kara <jack@suse.cz>
Subject: Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
Date: Wed, 30 Sep 2026 13:38:39 +0200	[thread overview]
Message-ID: <de323c38-4719-45fa-b4ac-6dae1b51b086@kernel.org> (raw)
In-Reply-To: <CAAE3jtftxMkCfWJ77W-xvV2x7jdAN8C8kyofYC3fEivM-Cq6UQ@mail.gmail.com>

On 9/30/26 12:27, 天狼 wrote:
> Hi David, Lorenzo,
>

Hi,

>> You really need an actual file on disk. Shared memory / shmem / memfd is
>> not sufficient?
> 
> Yes, an ordinary disk-backed file is the intended output, and subsequent
> consumers access it through mmap.
> 
> Shared memory can preserve state across a process crash if the shared
> object remains alive. However, using shmem/memfd would introduce a separate
> staging and transfer step to put the data into the final file. Building
> directly in a file-backed MAP_SHARED mapping lets us preserve the latest
> userspace writes across process crashes and retain the same file-cache
> pages for subsequent consumers.

Why do you really need the data on a file in disk?

What you could do is, load it once from the file into shmem, then let everybody
work on shmem, and have some background thread that periodically writes the
shmem content out to the real file on disk.

Then, you could completely control when I/O would happen.

[...]

>> It'd be problematic to effectively corrupt the page cache by saying "hey
>> this is dirty but just clear dirty state and pretend this is what the
>> disk has".
> 
> To clarify, the proposal does not involve clearing dirty state or
> pretending that memory and disk contents match. The pages would remain
> dirty and continue to count toward the existing dirty-page limits.
> 

IIUC, vmscan would still write out the folio if dirty, so I'd assume that at
least memory reclaim would not be affected, only background writeback.

> The requested hint only expresses that the application expects to modify
> the data again soon and would prefer background writeback to happen later.
> 
>> Delaying writeback could interfere with the writeback balancing logic.
> 
> I understand that concern. The hint should not exempt the application from
> dirty-page limits or throttling. If writeback is needed for balancing,
> memory pressure, or explicit synchronization, the kernel should override
> the hint.
> 
> The intended benefit is to avoid writing intermediate versions when the
> kernel has room to defer that work. There is no requirement to prevent
> writeback until construction finishes.
> 
> madvise is not the only option, deferring writeback for the entire file
> would also be acceptable.
> 
> Could such an advisory preference be accommodated within the existing
> writeback balancing policy while retaining its limits and fairness?
When using MAP_SHARED with disk-based files for VM memory, people ran into
similar problems: the VM will constantly dirty many folios and background
writeback will just permanently write these out and wear storage. For VMs you
really only want the memory persisted in the file once you e.g., migrate the VM.
Essentially, once the VM is paused.

One of the reasons why we tell people to not use that combination and use
shmem/hugetlb instead. I do wonder whether there is some similarity ... we don't
want brackground writeback as long as there is heavy activity on these files.

-- 
Cheers,

David

  reply	other threads:[~2026-09-30 11:38 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28  7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09   ` 天狼
2026-09-28 19:12     ` David Hildenbrand (Arm)
2026-09-29  4:06       ` 天狼
2026-09-29  6:59         ` David Hildenbrand (Arm)
2026-09-29  8:59           ` Lorenzo Stoakes (ARM)
2026-09-30 10:27             ` 天狼
2026-09-30 11:38               ` David Hildenbrand (Arm) [this message]
2026-09-30 12:08               ` Lorenzo Stoakes (ARM)
2026-09-30 12:11                 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14                   ` 天狼
2026-09-30 14:24                     ` Lorenzo Stoakes (ARM)
2026-09-30 15:33                       ` Lorenzo Stoakes (ARM)
2026-09-30 15:29                     ` 天狼
2026-09-30 15:43                       ` Lorenzo Stoakes (ARM)
2026-09-30 16:38                       ` Jan Kara
2026-10-01  4:28                         ` 天狼
2026-10-01  8:35                           ` Lorenzo Stoakes (ARM)
2026-10-01 13:21                             ` 天狼

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=de323c38-4719-45fa-b4ac-6dae1b51b086@kernel.org \
    --to=david@kernel.org \
    --cc=jack@suse.cz \
    --cc=jannh@google.com \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=rockeet@gmail.com \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox