Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Matthew Wilcox <willy@infradead.org>
To: Jann Horn <jannh@google.com>
Cc: Pedro Falcato <pfalcato@suse.de>, Christoph Hellwig <hch@lst.de>,
	David Howells <dhowells@redhat.com>,
	John Hubbard <jhubbard@nvidia.com>, Jan Kara <jack@suse.cz>,
	Rik van Riel <riel@surriel.com>, Qu Wenruo <wqu@suse.com>,
	"Darrick J. Wong" <djwong@kernel.org>,
	linux-btrfs@vger.kernel.org, linux-fsdevel@vger.kernel.org,
	linux-mm@kvack.org, linux-xfs@vger.kernel.org
Subject: Re: Removing ->dirty_folio
Date: Tue, 25 Aug 2026 20:54:25 +0100	[thread overview]
Message-ID: <ao3y8TtEGjniZ_4f@casper.infradead.org> (raw)
In-Reply-To: <CAG48ez0vXVUu9j4rYUeCqgwyHKAiN_6fL--D90WvX9q7B41V-A@mail.gmail.com>

On Tue, Aug 25, 2026 at 09:35:57PM +0200, Jann Horn wrote:
> > The problem is GUP.  We have no way to force the GUP caller to go
> > through page_mkwrite again.  So instead we make the GUP caller call
> > folio_mark_dirty_lock() which many just don't, and generally we get away
> > with it.  But it's a bug, and a bad interface.
> 
> We currently have get_user_pages*() and pin_user_pages*(), where only
> the pin_*() version is properly usable for write access, right? And
> dropping such pins should always go through unpin_*() helpers?

Right.  I'm stuffing cheese into my ears and pretending that people
aren't calling get_user_pages() to do write accesses.  We should be
able to use this work to flush out the last remaining ones -- we
can put in various assertions that folios should still be dirty where
we currently have folio_mark_dirty() calls.

> Could we strictly ban using get_user_pages*() for write access, and
> use the unpin_*() helpers to somehow enforce that folios are always
> dirtied on unpin? I guess the problem with that is that we have no
> state that tracks whether the pin was read-only, and finding free bits
> in struct page to keep track of this is hard?

We'd need a count, not just a bit or two.  A pincount for all folios
is on its way (eventually) but I wans't planning on tracking writable vs
read-only pins.

> > My proposal is this:
> >
> >  - Fileystems take note of folio_maybe_dma_pinned() during writeback.
> >    If it's true, do the writeback, but retain/recreate whatever data
> >    structures you need in order to write the folio again; behave as if
> >    ->page_mkdirty() had been called again for each page in the folio is
> >    marked as dirty.
> 
> I don't understand this part of the MM/VFS machinery well - would this
> mean that a long-term pin of a dirty pagecache folio could cause an
> unbounded number of disk writes in regular intervals, even if nothing
> actually writes into the folio? I'm guessing that could be bad for
> cheap flash storage, but maybe this is in the category of "yes that
> would be bad but it would be userspace's fault".

Yes.  I think the only alternative would be storing a checksum of the
contents of the folio and seeing if it changed since the last write.
As you say, this is userspace doing something incredibly odd (I really
don't think people make a habit of mmap()ing files shared writable and
then giving RDMA longterm write accesses to them.  Not on machines with
poor quality flash storage anyway).

Or we could say "these pages only get written back on requested fsync()
rather than periodically".


  reply	other threads:[~2026-08-25 19:54 UTC|newest]

Thread overview: 28+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-24 19:08 Removing ->dirty_folio Matthew Wilcox
2026-08-24 19:25 ` John Hubbard
2026-08-24 19:43   ` Matthew Wilcox
2026-08-24 19:51     ` John Hubbard
2026-08-24 21:05       ` Matthew Wilcox
2026-08-25  7:59     ` Pedro Falcato
2026-08-24 19:33 ` Rik van Riel
2026-08-25  8:21   ` David Hildenbrand (Arm)
2026-08-24 21:27 ` Boris Burkov
2026-08-25 19:47   ` Matthew Wilcox
2026-08-24 22:38 ` Qu Wenruo
2026-08-25 19:26   ` Matthew Wilcox
2026-08-25 22:33     ` Qu Wenruo
2026-08-26  4:53       ` Christoph Hellwig
2026-08-26  8:01     ` Christoph Hellwig
2026-08-25  6:48 ` David Howells
2026-08-25 19:16   ` Matthew Wilcox
2026-08-25  7:39 ` Christoph Hellwig
2026-08-25 19:14   ` Matthew Wilcox
2026-08-25  8:25 ` Pedro Falcato
2026-08-25 18:35   ` Matthew Wilcox
2026-08-26  7:54     ` Christoph Hellwig
2026-08-25 19:35 ` Jann Horn
2026-08-25 19:54   ` Matthew Wilcox [this message]
2026-08-25 20:06     ` Jann Horn
2026-08-26  0:02       ` Matthew Wilcox
2026-08-26  5:07     ` Christoph Hellwig
2026-08-26  7:42       ` Christoph Hellwig

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ao3y8TtEGjniZ_4f@casper.infradead.org \
    --to=willy@infradead.org \
    --cc=dhowells@redhat.com \
    --cc=djwong@kernel.org \
    --cc=hch@lst.de \
    --cc=jack@suse.cz \
    --cc=jannh@google.com \
    --cc=jhubbard@nvidia.com \
    --cc=linux-btrfs@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-xfs@vger.kernel.org \
    --cc=pfalcato@suse.de \
    --cc=riel@surriel.com \
    --cc=wqu@suse.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox