All of lore.kernel.org
 help / color / mirror / Atom feed
From: Kiryl Shutsemau <kas@kernel.org>
To: akpm@linux-foundation.org,
	 "Matthew Wilcox (Oracle)" <willy@infradead.org>,
	David Hildenbrand <david@kernel.org>,
	 Boris Burkov <boris@bur.io>
Cc: Lorenzo Stoakes <ljs@kernel.org>,
	 "Liam R. Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	 Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	 Michal Hocko <mhocko@suse.com>, Jan Kara <jack@suse.cz>,
	Rik van Riel <riel@surriel.com>,  Harry Yoo <harry@kernel.org>,
	Lance Yang <lance.yang@linux.dev>, Jann Horn <jannh@google.com>,
	 Alexander Viro <viro@zeniv.linux.org.uk>,
	Christian Brauner <brauner@kernel.org>,
	 "Darrick J. Wong" <djwong@kernel.org>,
	Carlos Maiolino <cem@kernel.org>,
	 Usama Arif <usama.arif@linux.dev>,
	Pedro Falcato <pfalcato@suse.de>,
	linux-mm@kvack.org,  linux-fsdevel@vger.kernel.org,
	linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org,
	 kernel-team@meta.com
Subject: Re: [RFC PATCH 0/5] mm: sub-folio dirty tracking for PTE-mapped mmap writes
Date: Mon, 7 Sep 2026 11:15:15 +0100	[thread overview]
Message-ID: <ap6LFvuOuWf_BkGi@thinkstation> (raw)
In-Reply-To: <20260903182943.662461-1-kirill@shutemov.name>

On Thu, Sep 03, 2026 at 07:29:38PM +0100, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> 
> A store through a shared file mapping dirties the whole folio. With large
> page cache folios that turns a 4K store into 2M of writeback: one dirty
> bit per folio, and writeback has no way to know which part changed.
> 
> XFS already knows better. iomap tracks dirty state per block and
> iomap_writeback_folio() submits only the dirty ranges, and the buffered
> write path sets just the range it copied. Only the mmap path throws that
> away, because iomap_dirty_folio() covers the whole folio.
> 
> Narrowing the dirtying at page_mkwrite() time does not work on its own:
> set_pte_range() batch-maps a whole folio writable on the first shared
> write fault, so the stores that follow never fault and never reach the
> filesystem.
> 
> So harvest the hardware instead. folio_clear_dirty_for_io() already calls
> folio_mkclean(), whose rmap walk reads pte_dirty() for every entry of the
> folio and throws it away. Those bits are the only record of which parts
> of a large folio were written through a mapping. Collect them there and
> hand the filesystem the runs that were dirty, through a new
> a_ops->dirty_folio_range().

Boris pointed me to Matthew's proposal to remove ->dirty_folio:

https://lore.kernel.org/all/aoyWln-Gt-yvZQkE@casper.infradead.org

I agree that the current ->dirty_folio() makes little sense and that
dirtying the folio can be bundled into ->page_mkwrite(), as they are
matched 1-to-1.

My proposal makes the distinction between making the folio writable and
making it dirty meaningful. ->page_mkwrite() allocates whatever is needed
on the filesystem side to track dirty state and drive writeback for the
*folio*, while ->dirty_folio_range() marks part of the folio dirty.

We can still drop ->dirty_folio(). A filesystem can provide
->dirty_folio_range() if it wants fine-grained (sub-folio) dirty
tracking.

A separate question is whether we want to avoid installing a writable PMD
entry for filesystems that want fine-grained dirty tracking. I have a
patch for this, but it deserves a separate discussion once we agree that we
want this for PTE-mapped folios first.

Any feedback?

> All of this is about PTE-mapped folios. A PMD-mapped folio has a single
> dirty bit for the 2M it maps, so there is nothing finer to harvest, and
> it keeps writing back whole. Keeping shared write faults off PMDs is a
> separate patch and not part of this posting.
> 
> On a 512M file in 2M folios on XFS, storing one byte per folio and
> calling msync() wrote 512M before and writes 1M after, with identical
> minor fault counts.
> 
> Not addressed here:
> 
>  - Dirty accounting stays folio-granular. A 4K store still counts as 2M
>    against dirty_ratio and balance_dirty_pages().
>  - iomap_page_mkwrite() still allocates blocks for the whole folio.
>  - Filesystems without per-block dirty state see no change.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov

  parent reply	other threads:[~2026-09-07 10:15 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03 18:29 [RFC PATCH 0/5] mm: sub-folio dirty tracking for PTE-mapped mmap writes Kiryl Shutsemau
2026-09-03 18:29 ` [RFC PATCH 1/5] mm: let folio_mkclean() report which pages had dirty PTEs Kiryl Shutsemau
2026-09-09 10:12   ` Usama Arif
2026-09-10 13:36     ` Kiryl Shutsemau
2026-09-03 18:29 ` [RFC PATCH 2/5] mm: add a_ops->dirty_folio_range() and use the mkclean dirty harvest Kiryl Shutsemau
2026-09-09 10:51   ` Usama Arif
2026-09-10 14:54     ` Kiryl Shutsemau
2026-09-03 18:29 ` [RFC PATCH 3/5] mm: keep the mmap dirty range down to the faulting page Kiryl Shutsemau
2026-09-03 18:29 ` [RFC PATCH 4/5] iomap: narrow page_mkwrite() dirtying " Kiryl Shutsemau
2026-09-03 18:29 ` [RFC PATCH 5/5] xfs: track mmap dirty state per block Kiryl Shutsemau
2026-09-03 19:55 ` [RFC PATCH 0/5] mm: sub-folio dirty tracking for PTE-mapped mmap writes Pedro Falcato
2026-09-03 21:18   ` Kiryl Shutsemau
2026-09-07 10:15 ` Kiryl Shutsemau [this message]
2026-09-09 10:02 ` Usama Arif
2026-09-09 10:15   ` Kiryl Shutsemau

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ap6LFvuOuWf_BkGi@thinkstation \
    --to=kas@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=boris@bur.io \
    --cc=brauner@kernel.org \
    --cc=cem@kernel.org \
    --cc=david@kernel.org \
    --cc=djwong@kernel.org \
    --cc=harry@kernel.org \
    --cc=jack@suse.cz \
    --cc=jannh@google.com \
    --cc=kernel-team@meta.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-xfs@vger.kernel.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=pfalcato@suse.de \
    --cc=riel@surriel.com \
    --cc=rppt@kernel.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.