Linux filesystem development
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: 天狼 <rockeet@gmail.com>
Cc: "David Hildenbrand (Arm)" <david@kernel.org>,
	linux-mm@kvack.org,  linux-fsdevel@vger.kernel.org,
	"Liam R. Howlett" <liam@infradead.org>,
	 Vlastimil Babka <vbabka@kernel.org>,
	Jann Horn <jannh@google.com>,
	 Matthew Wilcox <willy@infradead.org>, Jan Kara <jack@suse.cz>
Subject: Re: [RFC] madvise: best-effort deferred writeback for shared file mappings
Date: Wed, 30 Sep 2026 16:33:11 +0100	[thread overview]
Message-ID: <ar0mdJKLC0cHBoKM@gremlin> (raw)
In-Reply-To: <ar0a9y2E0nDTIC8V@gremlin>

On Wed, Sep 30, 2026 at 03:24:53PM +0100, Lorenzo Stoakes (ARM) wrote:
> I strongly recommend you look at a shmem solution.

To expand upon that:

The general structure would be something like a minimal process, with an
oom_score_adj set such that it can't be killed (or is at least very unlikely).

Have it do something like allocate a memfd.

Then have that be as minimal as possible - it sets things up, it hands off
memfds over a UNIX socket for instance, and is set up to periodically write
back to disk.

That way you have something that should never crash/exit which gives you the
guarantees you need.

You can then also control how to synchronise things between processes, when to
write back and how much, etc.

You can keep it in a cgroup as well to isolate it from the rest of the system
too, potentially.

Then have anything that does anything non-trivial be clients of that.

Everything works as before, you can ensure that intermediate state doesn't get
written back through whatever mechanism you want to employ, and you avoid all of
the pitfalls of writing to a MAP_SHARED file-backed mapping.

If the memfd isn't enough for you, you could also have it establish a tmpfs file
which is then shared between all of the processes. That will survive even the
establishing process dying.

The main win here is that you can choose when and how to writeback using any
mechanism you like.

For instance, you could have a dirty bitmap be part of the shared memory and
atomically update the relevant bit when it's ready to be written back.

That kind of approach gives you total control over when and how things are
written back.

And avoids the known issues with writing to a MAP_SHARED file:

* Terrible error handling - random SIGBUS's, network file systems in particular
  are very problematic, writeback errors are hard to obtain (fsync(), msync()
  needed to even get them) - note the tmpfs compromise mentioned above has this
  issue too.

* No control over writeback - Exactly your issue. This is what you're trying to
  work around.

* Writeback stalls - as above. The dirty writeback balancing bites there.

* Reclaim - dirty pages can't be reclaimed until written back.

* Fault overhead - every fresh dirtying write is a page fault (that's how dirty
  tracking works) and you have that overhead throughout. Once written back it's
  cleaned and causes the same cost again next time.

The TL;DR is you're trying to manipulate logic whose whole job it is to
writeback to not do that.

So stop doing that :) and life is much easier.

--
Cheers, Lorenzo

  reply	other threads:[~2026-09-30 15:33 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28  7:07 [RFC] madvise: best-effort deferred writeback for shared file mappings 天狼
2026-09-28 11:29 ` David Hildenbrand (Arm)
2026-09-28 13:09   ` 天狼
2026-09-28 19:12     ` David Hildenbrand (Arm)
2026-09-29  4:06       ` 天狼
2026-09-29  6:59         ` David Hildenbrand (Arm)
2026-09-29  8:59           ` Lorenzo Stoakes (ARM)
2026-09-30 10:27             ` 天狼
2026-09-30 11:38               ` David Hildenbrand (Arm)
2026-09-30 12:08               ` Lorenzo Stoakes (ARM)
2026-09-30 12:11                 ` Lorenzo Stoakes (ARM)
2026-09-30 14:14                   ` 天狼
2026-09-30 14:24                     ` Lorenzo Stoakes (ARM)
2026-09-30 15:33                       ` Lorenzo Stoakes (ARM) [this message]
2026-09-30 15:29                     ` 天狼
2026-09-30 15:43                       ` Lorenzo Stoakes (ARM)
2026-09-30 16:38                       ` Jan Kara
2026-10-01  4:28                         ` 天狼
2026-10-01  8:35                           ` Lorenzo Stoakes (ARM)
2026-10-01 13:21                             ` 天狼

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ar0mdJKLC0cHBoKM@gremlin \
    --to=ljs@kernel.org \
    --cc=david@kernel.org \
    --cc=jack@suse.cz \
    --cc=jannh@google.com \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=rockeet@gmail.com \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox