Linux filesystem development
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: 天狼 <rockeet@gmail.com>
Cc: linux-mm@kvack.org, linux-fsdevel@vger.kernel.org,
	"Lorenzo Stoakes (Arm)" <ljs@kernel.org>,
	"Liam R. Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>, Jann Horn <jannh@google.com>
Subject: Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
Date: Mon, 28 Sep 2026 21:12:30 +0200	[thread overview]
Message-ID: <85aff663-2131-47da-ac04-8f0799a49b91@kernel.org> (raw)
In-Reply-To: <CAAE3jtfeKJLSypScqRAEEgvO8En5kJfpbcFn5X+LQ2YL+xz_vA@mail.gmail.com>

On 9/28/26 15:09, 天狼 wrote:
> Hi David,

Hi,

> 
> The use case is a concurrent, lock-free, relocatable data structure that
> uses offsets rather than absolute pointers. It implements copy-on-write
> at the data-structure level and is designed to remain recoverable after
> a process crash. We therefore build it directly in a file-backed
> MAP_SHARED mapping, so its state is not lost when the constructing
> process exits unexpectedly.

What is supposed to happen if the process crashes when updating the MAP_SHARED
region halfway through?

It would be helpful if the use case + data structure would be explained in a
bit more detail.

> 
> Building it in anonymous memory and calling write() only after
> construction would lose that state if the process crashed before the
> write. It would also require copying the entire structure into the
> file's page cache.

Yes. Unless to would try to bypass the page cache of course (if possible for
your use case).

> 
> During construction, writes are scattered throughout the mapping,
> dirtying pages rapidly and repeatedly redirtying pages that may already
> have been written back. The goal is to reduce writeback of these
> intermediate states while retaining the same page-cache pages for
> subsequent readers.

What exactly is the poblem with writeback here? Unnecessary I/O? Is writeback
the problem or actual reclaim after writeback?

(I recall that in a fuse server you can in theory delay the writeback request.
So maybe you could get something going by serving your file through fuse and
enlightening the fuse server about it. Just a random idea.)

> 
> The hint would remain strictly best-effort: memory pressure and explicit
> synchronization could override it. By "process crash," I mean termination
> of the process while the kernel remains running; recovery after a system
> crash or power loss is a separate durability concern.
> 
> Regarding the interface, madvise was only a suggestion. A file-level
> advisory interface would also work for this use case: deferring background
> writeback for the entire file, rather than individual mapped ranges, is
> acceptable.

So fadvise would be an option. However, this "defer mode" is really odd. It
sounds more like you would want to have a custom policy there, instead of
hardcoding something that really not a lot might want. Hmmm

-- 
Cheers,

David

  reply	other threads:[~2026-09-28 19:12 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28  7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09   ` 天狼
2026-09-28 19:12     ` David Hildenbrand (Arm) [this message]
2026-09-29  4:06       ` 天狼
2026-09-29  6:59         ` David Hildenbrand (Arm)
2026-09-29  8:59           ` Lorenzo Stoakes (ARM)
2026-09-30 10:27             ` 天狼
2026-09-30 11:38               ` David Hildenbrand (Arm)
2026-09-30 12:08               ` Lorenzo Stoakes (ARM)
2026-09-30 12:11                 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14                   ` 天狼
2026-09-30 14:24                     ` Lorenzo Stoakes (ARM)
2026-09-30 15:33                       ` Lorenzo Stoakes (ARM)
2026-09-30 15:29                     ` 天狼
2026-09-30 15:43                       ` Lorenzo Stoakes (ARM)
2026-09-30 16:38                       ` Jan Kara
2026-10-01  4:28                         ` 天狼
2026-10-01  8:35                           ` Lorenzo Stoakes (ARM)
2026-10-01 13:21                             ` 天狼

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=85aff663-2131-47da-ac04-8f0799a49b91@kernel.org \
    --to=david@kernel.org \
    --cc=jannh@google.com \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=rockeet@gmail.com \
    --cc=vbabka@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox