Linux filesystem development
 help / color / mirror / Atom feed
From: David Howells <dhowells@redhat.com>
To: Paulo Alcantara <pc@manguebit.org>
Cc: David Howells <dhowells@redhat.com>,
	Christian Brauner <christian@brauner.io>,
	Matthew Wilcox <willy@infradead.org>,
	Christoph Hellwig <hch@infradead.org>,
	Jens Axboe <axboe@kernel.dk>, Leon Romanovsky <leon@kernel.org>,
	Namjae Jeon <linkinjeon@kernel.org>,
	ChenXiaoSong <chenxiaosong@chenxiaosong.com>,
	Marc Dionne <marc.dionne@auristor.com>,
	Stefan Metzmacher <metze@samba.org>,
	Eric Van Hensbergen <ericvh@kernel.org>,
	Dominique Martinet <asmadeus@codewreck.org>,
	Ilya Dryomov <idryomov@gmail.com>,
	netfs@lists.linux.dev, linux-afs@lists.infradead.org,
	linux-cifs@vger.kernel.org, linux-nfs@vger.kernel.org,
	ceph-devel@vger.kernel.org, v9fs@lists.linux.dev,
	linux-erofs@lists.ozlabs.org, linux-fsdevel@vger.kernel.org,
	linux-kernel@vger.kernel.org
Subject: [PATCH v11 28/36] netfs: Simplify writeback cleanup
Date: Wed,  2 Sep 2026 18:33:40 +0100	[thread overview]
Message-ID: <20260902173350.3468672-29-dhowells@redhat.com> (raw)
In-Reply-To: <20260902173350.3468672-1-dhowells@redhat.com>

Currently, the netfslib buffered writeback algorithm walks the list of
folios, using that to determine the folios that need to be unlocked.  This
is tricky, however, as different streams really want different folios or
different parts of folios (e.g. data that's read from the server will be
written to the cache, but not written back to the server, and a small
region that can be written to the server may need to be rounded out for DIO
write to the cache).

Further, the collector thread may be walking the folio list at the same
time that the application thread is filling it - and at the same time as
things are doing I/O to or from it.

Finally, it requires careful cleanup during collection, such that there's
always at least one link remaining in the bvecq chain so that the consumer
never gets disconnected from the producer.

Instead, use the end-writeback iteration tool to simplify the cleanup of
writeback by walking the inode's pagecache xarray to find the folios that
will be unmarked rather than looking in the list of bio_vecs that refer to
those folios.

This makes use of the list of file sections under writeback to keep track
of what needs to be dealt with.

Using the end-writeback iterator affords the possibility of improving
efficiency of the process by doing bulk access and bulk modification of the
xarray and stats.

Signed-off-by: David Howells <dhowells@redhat.com>
cc: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Christoph Hellwig <hch@infradead.org>
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
 fs/netfs/internal.h          |   2 +-
 fs/netfs/write_collect.c     | 170 +++++++++++++----------------------
 fs/netfs/write_issue.c       |   6 +-
 include/trace/events/netfs.h |  30 +++++++
 4 files changed, 97 insertions(+), 111 deletions(-)

diff --git a/fs/netfs/internal.h b/fs/netfs/internal.h
index e6e768ad9648..b62c5b7e43d9 100644
--- a/fs/netfs/internal.h
+++ b/fs/netfs/internal.h
@@ -222,7 +222,7 @@ static inline void netfs_stat_d(atomic_t *stat)
 /*
  * write_collect.c
  */
-int netfs_folio_written_back(struct folio *folio);
+void netfs_folio_written_back(struct netfs_io_request *wreq, struct folio *folio);
 bool netfs_write_collection(struct netfs_io_request *wreq);
 void netfs_write_collection_worker(struct work_struct *work);
 
diff --git a/fs/netfs/write_collect.c b/fs/netfs/write_collect.c
index 40f8b1485992..2510d3dd2d58 100644
--- a/fs/netfs/write_collect.c
+++ b/fs/netfs/write_collect.c
@@ -21,57 +21,6 @@
 #define NEED_RETRY		0x10	/* A front op requests retrying */
 #define SAW_FAILURE		0x20	/* One stream or hit a permanent failure */
 
-struct end_writeback_ctrl {
-	struct xa_state xas;
-	uoff_t fend;
-};
-
-/**
- * end_writeback_iter - End writeback for the folios within the range
- * @mapping: The pagecache to modify
- * @from: Pointer to the starting position
- * @to: The end position (exclusive)
- * @ctrl: Iterator state
- * @folio: The last return from this function or NULL if first call
- *
- * Remove the writeback mark on folios that are entirely within in the given
- * range, where @from is included in the range, but @to is excluded from the
- * range.
- *
- * Return: The folio to end writeback upon or NULL if the next folio isn't
- * wholly within the range.  Note that this does not necessarily imply that the
- * ending is complete.  @ctrl->fend is updated to point directly beyond the
- * folio that was considered and may be negative if cast to loff_t.
- */
-static
-struct folio *end_writeback_iter(struct address_space *mapping,
-				 uoff_t from, uoff_t to,
-				 struct end_writeback_ctrl *ctrl,
-				 struct folio *folio)
-{
-	if (!folio) {
-		lockdep_assert_in_rcu_read_lock();
-		ctrl->xas = (struct xa_state)
-			__XA_STATE(&mapping->i_pages, from / PAGE_SIZE, 0, 0);
-		folio = xas_find(&ctrl->xas, (to - 1) / PAGE_SIZE);
-	} else {
-		folio_end_writeback(folio);
-	retry:
-		folio = xas_next_entry(&ctrl->xas, (to - 1) / PAGE_SIZE);
-	}
-
-	if (xas_retry(&ctrl->xas, folio))
-		goto retry;
-
-	if (!folio)
-		return NULL;
-
-	ctrl->fend = folio_next_pos(folio);
-	if (ctrl->fend > to)
-		folio = NULL;
-	return folio;
-}
-
 static void netfs_dump_request(const struct netfs_io_request *rreq)
 {
 	pr_err("Request R=%08x r=%d fl=%lx or=%x e=%ld\n",
@@ -105,14 +54,22 @@ static void netfs_dump_request(const struct netfs_io_request *rreq)
  * that we are not allowed to lock the folio here on pain of deadlocking with
  * truncate.
  */
-int netfs_folio_written_back(struct folio *folio)
+void netfs_folio_written_back(struct netfs_io_request *wreq, struct folio *folio)
 {
 	enum netfs_folio_trace why = netfs_folio_trace_endwb;
 	struct inode *inode = folio_inode(folio);
 	struct netfs_inode *ictx = netfs_inode(inode);
 	struct netfs_folio *finfo;
 	struct netfs_group *group = NULL;
-	int gcount = 0;
+
+	if (WARN_ONCE(!folio_test_writeback(folio),
+		      "R=%08x: folio %lx is not under writeback\n",
+		      wreq->debug_id, folio->index)) {
+		trace_netfs_folio(folio, netfs_folio_trace_not_under_wback);
+		netfs_dump_request(wreq);
+	}
+
+	trace_netfs_collect_folio(wreq, folio);
 
 	if ((finfo = netfs_folio_info(folio))) {
 		/* Streaming writes cannot be redirtied whilst under writeback,
@@ -128,7 +85,7 @@ int netfs_folio_written_back(struct folio *folio)
 
 		folio_detach_private(folio);
 		group = finfo->netfs_group;
-		gcount++;
+		wreq->nr_group_rel++;
 		kfree(finfo);
 		why = netfs_folio_trace_endwb_s;
 		goto end_wb;
@@ -148,15 +105,34 @@ int netfs_folio_written_back(struct folio *folio)
 		why = netfs_folio_trace_redirtied;
 		if (!folio_test_dirty(folio)) {
 			folio_detach_private(folio);
-			gcount++;
+			wreq->nr_group_rel++;
 			why = netfs_folio_trace_endwb_g;
 		}
 	}
 
 end_wb:
 	trace_netfs_folio(folio, why);
-	folio_end_writeback(folio);
-	return gcount;
+}
+
+static bool netfs_writeback_unlock_range(struct netfs_io_request *wreq, uoff_t stop_at)
+{
+	struct end_writeback_ctrl ctrl = {};
+	struct folio *folio = NULL;
+	bool progress = false;
+
+	rcu_read_lock();
+	for (;;) {
+		folio = end_writeback_iter(wreq->mapping, &ctrl,
+					   wreq->cleaned_to, stop_at, folio);
+		if (!folio)
+			break;
+		netfs_folio_written_back(wreq, folio);
+		wreq->cleaned_to = ctrl.fend;
+		progress = true;
+	}
+	rcu_read_unlock();
+
+	return progress;
 }
 
 /*
@@ -165,15 +141,7 @@ int netfs_folio_written_back(struct folio *folio)
 static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
 					  unsigned int *notes)
 {
-	struct bvecq *bvecq = wreq->collect_cursor.bvecq;
-	unsigned int slot = wreq->collect_cursor.slot;
-	uoff_t collected_to = wreq->collected_to;
-
-	if (WARN_ON_ONCE(!bvecq)) {
-		pr_err("[!] Writeback unlock found empty buffer!\n");
-		netfs_dump_request(wreq);
-		return;
-	}
+	struct netfs_writeback *wback, *next;
 
 	if (wreq->origin == NETFS_PGPRIV2_COPY_TO_CACHE) {
 		if (netfs_pgpriv2_unlock_copied_folios(wreq))
@@ -181,57 +149,45 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
 		return;
 	}
 
+	wback = wreq->writebacks;
+
 	for (;;) {
-		struct folio *folio;
-		struct netfs_folio *finfo;
-		uoff_t fpos, fend;
-		size_t fsize, flen;
-
-		/* Try to clean up the head of the queue if it appears to be
-		 * used up, but we need to be very careful - the cleanup can
-		 * catch the dispatcher, which could lead to us having nothing
-		 * left in the queue, causing the front and back pointers to
-		 * end up on different tracks.  To avoid this, we must always
-		 * keep at least one segment in the queue.
-		 */
-		if (!bvecq_acquire_slot(bvecq, slot)) {
-			wreq->collect_cursor.slot = slot;
-			if (!bvecq_delete_spent(&wreq->collect_cursor))
-				return;
-			bvecq = wreq->collect_cursor.bvecq;
-			slot  = wreq->collect_cursor.slot;
-		}
+		uoff_t stop_at;
+		size_t len;
 
-		folio = page_folio(bvecq->bv[slot].bv_page);
-		if (WARN_ONCE(!folio_test_writeback(folio),
-			      "R=%08x: folio %lx is not under writeback\n",
-			      wreq->debug_id, folio->index))
-			trace_netfs_folio(folio, netfs_folio_trace_not_under_wback);
+		/* Jump over discontiguities. */
+		if (wreq->cleaned_to < wback->start)
+			wreq->cleaned_to = wback->start;
 
-		fpos = folio_pos(folio);
-		fsize = folio_size(folio);
-		finfo = netfs_folio_info(folio);
-		flen = finfo ? finfo->dirty_offset + finfo->dirty_len : fsize;
+		if (wreq->collected_to <= wreq->cleaned_to)
+			break;
 
-		fend = min_t(uoff_t, fpos + flen, wreq->i_size);
+		/* Order read of region length before reading folios. */
+		len = smp_load_acquire(&wback->len);
 
-		trace_netfs_collect_folio(wreq, folio);
+		if (wreq->cleaned_to >= wback->start + len) {
+			/* Order read of next before recheck length. */
+			next = smp_load_acquire(&wback->next);
+			if (!next)
+				break; /* We don't remove the tail writeback. */
 
-		/* Unlock any folio we've transferred all of. */
-		if (collected_to < fend)
-			break;
+			/* Order read of region length before reading folios. */
+			if (len != smp_load_acquire(&wback->len))
+				continue; /* len/next update race. */
 
-		wreq->nr_group_rel += netfs_folio_written_back(folio);
-		wreq->cleaned_to = fpos + fsize;
-		*notes |= MADE_PROGRESS;
+			mempool_free(wback, &netfs_bvecq_pool);
+			wreq->writebacks = next;
+			wback = next;
+			continue;
+		}
+
+		stop_at = min(wreq->collected_to, wback->start + len);
 
-		bvecq->bv[slot].bv_page = NULL;
-		slot++;
-		if (fpos + fsize >= collected_to)
+		trace_netfs_collect_folios(wreq, wback->start, len);
+		if (!netfs_writeback_unlock_range(wreq, stop_at))
 			break;
+		*notes |= MADE_PROGRESS;
 	}
-
-	wreq->collect_cursor.slot = slot;
 }
 
 /*
diff --git a/fs/netfs/write_issue.c b/fs/netfs/write_issue.c
index ace790b127ca..086869bd0db1 100644
--- a/fs/netfs/write_issue.c
+++ b/fs/netfs/write_issue.c
@@ -369,7 +369,7 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
 		_debug("beyond eof");
 		folio_start_writeback(folio);
 		folio_unlock(folio);
-		wreq->nr_group_rel += netfs_folio_written_back(folio);
+		netfs_folio_written_back(wreq, folio);
 		netfs_put_group_many(wreq->group, wreq->nr_group_rel);
 		wreq->nr_group_rel = 0;
 		return 0;
@@ -457,13 +457,13 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
 		if (!cache->avail) {
 			trace_netfs_folio(folio, netfs_folio_trace_cancel_copy);
 			netfs_issue_write(wreq, upload);
-			netfs_folio_written_back(folio);
+			netfs_folio_written_back(wreq, folio);
 			return 0;
 		}
 		trace_netfs_folio(folio, netfs_folio_trace_store_copy);
 	} else if (!upload->avail && !cache->avail) {
 		trace_netfs_folio(folio, netfs_folio_trace_cancel_store);
-		netfs_folio_written_back(folio);
+		netfs_folio_written_back(wreq, folio);
 		return 0;
 	} else if (!upload->construct) {
 		trace_netfs_folio(folio, netfs_folio_trace_store);
diff --git a/include/trace/events/netfs.h b/include/trace/events/netfs.h
index 444ae7d3a01a..006bd2c73670 100644
--- a/include/trace/events/netfs.h
+++ b/include/trace/events/netfs.h
@@ -773,6 +773,36 @@ TRACE_EVENT(netfs_collect_stream,
 		      __entry->collected_to, __entry->issued_to)
 	    );
 
+TRACE_EVENT(netfs_collect_folios,
+	    TP_PROTO(const struct netfs_io_request *wreq,
+		     uoff_t range_start, size_t range_len),
+
+	    TP_ARGS(wreq, range_start, range_len),
+
+	    TP_STRUCT__entry(
+		    __field(unsigned int,	wreq)
+		    __field(size_t,		range_len)
+		    __field(uoff_t,		range_start)
+		    __field(uoff_t,		cleaned_to)
+		    __field(uoff_t,		collected_to)
+			     ),
+
+	    TP_fast_assign(
+		    __entry->wreq		= wreq->debug_id;
+		    __entry->range_len		= range_len;
+		    __entry->range_start	= range_start;
+		    __entry->cleaned_to		= wreq->cleaned_to;
+		    __entry->collected_to	= wreq->collected_to;
+			   ),
+
+	    TP_printk("R=%08x r=%llx-%llx cln=%llx col=%llx",
+		      __entry->wreq,
+		      __entry->range_start,
+		      __entry->range_start + __entry->range_len,
+		      __entry->cleaned_to,
+		      __entry->collected_to)
+	    );
+
 TRACE_EVENT(netfs_bvecq,
 	    TP_PROTO(const struct bvecq *bq,
 		     enum netfs_bvecq_trace trace),


  parent reply	other threads:[~2026-09-02 17:37 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02 17:33 [PATCH v11 00/36] netfs: Keep track of folios in a segmented bio_vec[] chain David Howells
2026-09-02 17:33 ` [PATCH v11 01/36] block: Fix start and length check added to iov_iter_extract_bvecs() David Howells
2026-09-02 17:33 ` [PATCH v11 02/36] mm: Make readahead store folio count in readahead_control David Howells
2026-09-02 17:33 ` [PATCH v11 03/36] mm: Add a bulk end-writeback tool David Howells
2026-09-02 17:33 ` [PATCH v11 04/36] netfs: Use uoff_t instead of unsigned long long and loff_t David Howells
2026-09-02 17:33 ` [PATCH v11 05/36] Add a function to kmap one page of a multipage bio_vec David Howells
2026-09-02 17:33 ` [PATCH v11 06/36] iov_iter: Make iov_iter_get_pages*() wrap iov_iter_extract_pages() David Howells
2026-09-02 17:33 ` [PATCH v11 07/36] iov_iter: Add a segmented queue of bio_vec[] David Howells
2026-09-02 17:33 ` [PATCH v11 08/36] netfs: Add some tools for managing bvecq chains David Howells
2026-09-02 17:33 ` [PATCH v11 09/36] netfs: Make mempool available for bvecq David Howells
2026-09-02 17:33 ` [PATCH v11 10/36] netfs: Add a function to extract from an iter into a bvecq David Howells
2026-09-02 17:33 ` [PATCH v11 11/36] afs: Use a bvecq to hold dir content rather than folioq David Howells
2026-09-02 17:33 ` [PATCH v11 12/36] cifs: Use a bvecq for buffering instead of a folioq David Howells
2026-09-02 17:33 ` [PATCH v11 13/36] smbdirect: Support ITER_BVECQ in smbdirect_map_sges_from_iter() David Howells
2026-09-02 17:33 ` [PATCH v11 14/36] netfs: Remove the writethrough code David Howells
2026-09-02 17:33 ` [PATCH v11 15/36] netfs: trace: Change the "clear" folio traces to "endwb" David Howells
2026-09-02 17:33 ` [PATCH v11 16/36] netfs: trace: Rejig a couple of the tracepoints David Howells
2026-09-02 17:33 ` [PATCH v11 17/36] netfs: Add some functions to wrap the all-queued handling David Howells
2026-09-02 17:33 ` [PATCH v11 18/36] netfs: Make deprecated PG_private_2 support optional David Howells
2026-09-02 17:33 ` [PATCH v11 19/36] cachefiles: Don't rely on backing fs storage map for most use cases David Howells
2026-09-02 17:33 ` [PATCH v11 20/36] netfs: Add the cache object ID to netfs_read/write tracepoints David Howells
2026-09-02 17:33 ` [PATCH v11 21/36] netfs: Switch to using bvecq rather than folio_queue and rolling_buffer David Howells
2026-09-02 17:33 ` [PATCH v11 22/36] smbdirect: Remove support for ITER_FOLIOQ from smbdirect_map_sges_from_iter() David Howells
2026-09-02 17:33 ` [PATCH v11 23/36] netfs: Remove netfs_alloc/free_folioq_buffer() David Howells
2026-09-02 17:33 ` [PATCH v11 24/36] netfs: Remove netfs_extract_user_iter() David Howells
2026-09-02 17:33 ` [PATCH v11 25/36] iov_iter: Remove ITER_FOLIOQ David Howells
2026-09-02 17:33 ` [PATCH v11 26/36] netfs: Remove folio_queue and rolling_buffer David Howells
2026-09-02 17:33 ` [PATCH v11 27/36] netfs: Build a list of regions undergoing writeback David Howells
2026-09-02 17:33 ` David Howells [this message]
2026-09-02 17:33 ` [PATCH v11 29/36] netfs: Simplify read abandonment David Howells
2026-09-02 17:33 ` [PATCH v11 30/36] netfs: Check for too much data being read David Howells
2026-09-02 17:33 ` [PATCH v11 31/36] netfs: Add a method to get an estimate of the amount that can be written David Howells
2026-09-02 17:33 ` [PATCH v11 32/36] netfs: Rework writeback to estimate David Howells
2026-09-02 17:33 ` [PATCH v11 33/36] netfs: Set subrequest->source at alloc before trace emission David Howells
2026-09-02 17:33 ` [PATCH v11 34/36] netfs: Combine prepare and issue ops David Howells
2026-09-02 17:33 ` [PATCH v11 35/36] netfs: Clean up now-unused code David Howells
2026-09-02 17:33 ` [PATCH v11 36/36] cachefiles: Preset the state xattr when creating a new file David Howells
2026-09-03  6:06 ` [PATCH v11 00/36] netfs: Keep track of folios in a segmented bio_vec[] chain Christoph Hellwig

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260902173350.3468672-29-dhowells@redhat.com \
    --to=dhowells@redhat.com \
    --cc=asmadeus@codewreck.org \
    --cc=axboe@kernel.dk \
    --cc=ceph-devel@vger.kernel.org \
    --cc=chenxiaosong@chenxiaosong.com \
    --cc=christian@brauner.io \
    --cc=ericvh@kernel.org \
    --cc=hch@infradead.org \
    --cc=idryomov@gmail.com \
    --cc=leon@kernel.org \
    --cc=linkinjeon@kernel.org \
    --cc=linux-afs@lists.infradead.org \
    --cc=linux-cifs@vger.kernel.org \
    --cc=linux-erofs@lists.ozlabs.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    --cc=marc.dionne@auristor.com \
    --cc=metze@samba.org \
    --cc=netfs@lists.linux.dev \
    --cc=pc@manguebit.org \
    --cc=v9fs@lists.linux.dev \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox