All of lore.kernel.org
 help / color / mirror / Atom feed
From: Zi Yan <ziy@nvidia.com>
To: David Hildenbrand <david@kernel.org>,
	 "Matthew Wilcox (Oracle)" <willy@infradead.org>,
	 Andrew Morton <akpm@linux-foundation.org>,
	 Muchun Song <muchun.song@linux.dev>,
	Lorenzo Stoakes <ljs@kernel.org>,
	 "Liam R. Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	 Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	 Michal Hocko <mhocko@suse.com>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	 Nico Pache <nico.pache@linux.dev>,
	Ryan Roberts <ryan.roberts@arm.com>,  Dev Jain <dev.jain@arm.com>,
	Barry Song <baohua@kernel.org>,
	 Lance Yang <lance.yang@linux.dev>,
	Usama Arif <usama.arif@linux.dev>,
	 Gregory Price <gourry@gourry.net>,
	 Ying Huang <ying.huang@linux.alibaba.com>,
	 Alistair Popple <apopple@nvidia.com>,
	Johannes Weiner <hannes@cmpxchg.org>,
	 Qi Zheng <qi.zheng@linux.dev>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	 Kairui Song <kasong@tencent.com>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	 Zi Yan <ziy@nvidia.com>, Gao Xiang <xiang@kernel.org>,
	 Chao Yu <chao@kernel.org>, Jan Kara <jack@suse.cz>,
	 Yue Hu <zbestahu@gmail.com>,
	Jeffle Xu <jefflexu@linux.alibaba.com>,
	 Sandeep Dhavale <dhavale@google.com>,
	Hongbo Li <hongbohbli@tencent.com>,
	 Chunhai Guo <guochunhai@vivo.com>,
	linux-erofs@lists.ozlabs.org,  linux-fsdevel@vger.kernel.org
Subject: [PATCH v2 07/14] erofs: mm/pagemap: add readahead_folio_last() to avoid folio->private
Date: Mon, 31 Aug 2026 15:25:30 -0400	[thread overview]
Message-ID: <20260831-remove-pg_private-v2-7-3668159cd9e8@nvidia.com> (raw)
In-Reply-To: <20260831-remove-pg_private-v2-0-3668159cd9e8@nvidia.com>

erofs needs to traverse readahead folios in reverse order to achieve
maximum performance by
1. reading all folios from readahead_folio();
2. storing the prior folio pointer in folio->private;
3. traverse from the last folio to the first one.

Add readahead_folio_last() to achieve the same function without using
folio->private. __readahead_advance() helper shares readahead_control
adjustment code among __readahead_folio(), readahead_folio_last(), and
__readahead_batch() by checking new private member, _forward, of
readahead_control.

It prepares for a future commit that replaces PG_private checks with
!folio->private checks. After switching the checks, erofs's use of
folio->private without bumping folio refcount can cause unexpected
outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
reachable.

No functional change intended.

Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Gao Xiang <xiang@kernel.org>
To: Chao Yu <chao@kernel.org>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
Cc: Yue Hu <zbestahu@gmail.com>
Cc: Jeffle Xu <jefflexu@linux.alibaba.com>
Cc: Sandeep Dhavale <dhavale@google.com>
Cc: Hongbo Li <hongbohbli@tencent.com>
Cc: Chunhai Guo <guochunhai@vivo.com>
Cc: linux-erofs@lists.ozlabs.org
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
---
 fs/erofs/zdata.c        | 13 +++---------
 include/linux/pagemap.h | 56 ++++++++++++++++++++++++++++++++++++++++++-------
 2 files changed, 51 insertions(+), 18 deletions(-)

diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
index e1e25ca0d1904..78fd7d980e957 100644
--- a/fs/erofs/zdata.c
+++ b/fs/erofs/zdata.c
@@ -1898,21 +1898,14 @@ static void z_erofs_readahead(struct readahead_control *rac)
 	struct inode *realinode = erofs_real_inode(sharedinode, &need_iput);
 	Z_EROFS_DEFINE_FRONTEND(f, realinode, sharedinode, readahead_pos(rac));
 	unsigned int nrpages = readahead_count(rac);
-	struct folio *head = NULL, *folio;
+	struct folio *folio;
 	int err;
 
 	trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
 	z_erofs_pcluster_readmore(&f, rac, true);
-	while ((folio = readahead_folio(rac))) {
-		folio->private = head;
-		head = folio;
-	}
-
-	/* traverse in reverse order for best metadata I/O performance */
-	while (head) {
-		folio = head;
-		head = folio_get_private(folio);
 
+	/* traverse from last to first for best metadata I/O performance */
+	while ((folio = readahead_folio_last(rac))) {
 		err = z_erofs_scan_folio(&f, folio, true);
 		if (err && err != -EINTR)
 			erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 0adfa6605653d..546e987d5cda7 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -1416,6 +1416,7 @@ struct readahead_control {
 	bool dropbehind;
 	bool _workingset;
 	unsigned long _pflags;
+	bool _forward;
 };
 
 #define DEFINE_READAHEAD(ractl, f, r, m, i)				\
@@ -1480,18 +1481,25 @@ void page_cache_async_readahead(struct address_space *mapping,
 	page_cache_async_ra(&ractl, folio, req_count);
 }
 
+static inline void __readahead_advance(struct readahead_control *rac)
+{
+	if (rac->_forward)
+		rac->_index += rac->_batch_count;
+
+	rac->_nr_pages -= rac->_batch_count;
+	rac->_batch_count = 0;
+}
+
 static inline struct folio *__readahead_folio(struct readahead_control *ractl)
 {
 	struct folio *folio;
 
 	BUG_ON(ractl->_batch_count > ractl->_nr_pages);
-	ractl->_nr_pages -= ractl->_batch_count;
-	ractl->_index += ractl->_batch_count;
+	__readahead_advance(ractl);
+	ractl->_forward = true;
 
-	if (!ractl->_nr_pages) {
-		ractl->_batch_count = 0;
+	if (!ractl->_nr_pages)
 		return NULL;
-	}
 
 	folio = xa_load(&ractl->mapping->i_pages, ractl->_index);
 	VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
@@ -1517,6 +1525,39 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
 	return folio;
 }
 
+/**
+ * readahead_folio_last - Get the next folio to read, from the tail.
+ * @ractl: The current readahead request.
+ *
+ * Like readahead_folio(), but walks the range back-to-front. The folio is
+ * returned locked with its refcount dropped; the caller unlocks it once I/O
+ * completes. Compound folios are returned once, at their head index.
+ *
+ * Context: The folio is locked.
+ * Return: A pointer to the next folio, or %NULL when done.
+ */
+static inline struct folio *readahead_folio_last(struct readahead_control *ractl)
+{
+	struct folio *folio;
+
+	/* Drop the previously returned batch from the remaining range. */
+	__readahead_advance(ractl);
+	ractl->_forward = false;
+
+	if (!ractl->_nr_pages)
+		return NULL;
+
+	/* xa_load() follows sibling entries, so a tail index returns the head */
+	folio = xa_load(&ractl->mapping->i_pages,
+			ractl->_index + ractl->_nr_pages - 1);
+	VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
+
+	ractl->_batch_count = folio_nr_pages(folio);
+
+	folio_put(folio);
+	return folio;
+}
+
 static inline unsigned int __readahead_batch(struct readahead_control *rac,
 		struct page **array, unsigned int array_sz)
 {
@@ -1525,9 +1566,8 @@ static inline unsigned int __readahead_batch(struct readahead_control *rac,
 	struct folio *folio;
 
 	BUG_ON(rac->_batch_count > rac->_nr_pages);
-	rac->_nr_pages -= rac->_batch_count;
-	rac->_index += rac->_batch_count;
-	rac->_batch_count = 0;
+	__readahead_advance(rac);
+	rac->_forward = true;
 
 	xas_set(&xas, rac->_index);
 	rcu_read_lock();

-- 
2.53.0


  parent reply	other threads:[~2026-08-31 19:26 UTC|newest]

Thread overview: 35+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31 19:25 [PATCH v2 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
2026-08-31 19:25 ` [f2fs-dev] " Zi Yan via Linux-f2fs-devel
2026-08-31 19:25 ` Zi Yan
2026-08-31 19:25 ` [PATCH v2 01/14] mm/zsmalloc: replace PG_private with pointer comparison Zi Yan
2026-08-31 19:25 ` [PATCH v2 02/14] perf/ring_buffer: stop using PG_private as AUX page high-order marker Zi Yan
2026-08-31 21:35   ` sashiko-bot
2026-08-31 19:25 ` [PATCH v2 03/14] xen/grant-table: stop setting PG_private on pages for grant mapping Zi Yan
2026-08-31 19:25 ` [PATCH v2 04/14] fscrypt: stop setting PG_private on bounce page Zi Yan
2026-08-31 19:25 ` [PATCH v2 05/14] mm/hugetlb: use direct assignment instead of folio_change_private() Zi Yan
2026-09-03  0:05   ` Gregory Price
2026-08-31 19:25 ` [PATCH v2 06/14] f2fs: stop using PG_private Zi Yan
2026-08-31 19:25   ` [f2fs-dev] " Zi Yan via Linux-f2fs-devel
2026-09-03 15:37   ` Jaegeuk Kim
2026-09-03 15:37     ` [f2fs-dev] " Jaegeuk Kim via Linux-f2fs-devel
2026-08-31 19:25 ` Zi Yan [this message]
2026-08-31 19:25 ` [PATCH v2 08/14] erofs: use folio_attach/detach_private() instead of direct assignment Zi Yan
2026-08-31 19:25 ` [PATCH v2 09/14] mm/page-flags: check page/folio->private instead of PG_private Zi Yan
2026-08-31 23:21   ` sashiko-bot
2026-09-01  2:11   ` Zi Yan
2026-08-31 19:25 ` [PATCH v2 10/14] mm/page-flags: introduce folio_test_fs_private() Zi Yan
2026-08-31 19:25 ` [PATCH v2 11/14] treewide: remove folio_set/clear_private() Zi Yan
2026-08-31 19:25 ` [PATCH v2 12/14] treewide: replace PagePrivate() with page_private() Zi Yan
2026-08-31 19:25 ` [PATCH v2 13/14] treewide: adjust comments on PagePrivate and PG_private Zi Yan
2026-08-31 19:25   ` Zi Yan
2026-08-31 19:25 ` [PATCH v2 14/14] mm/page-flags: remove PG_private Zi Yan
2026-09-01  0:16   ` sashiko-bot
2026-09-01  2:17   ` Zi Yan
2026-09-01 15:55   ` Steven Rostedt
2026-09-01 16:01     ` Zi Yan
2026-09-01 17:50       ` Steven Rostedt
2026-09-02 17:09   ` Usama Arif
2026-09-02 17:57     ` Zi Yan
2026-09-03 16:10 ` [f2fs-dev] [PATCH v2 00/14] Remove PG_private by using page/folio->private checks instead patchwork-bot+f2fs
2026-09-03 16:10   ` patchwork-bot+f2fs
2026-09-03 16:10   ` patchwork-bot+f2fs--- via Linux-f2fs-devel

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260831-remove-pg_private-v2-7-3668159cd9e8@nvidia.com \
    --to=ziy@nvidia.com \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=chao@kernel.org \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=dhavale@google.com \
    --cc=gourry@gourry.net \
    --cc=guochunhai@vivo.com \
    --cc=hannes@cmpxchg.org \
    --cc=hongbohbli@tencent.com \
    --cc=jack@suse.cz \
    --cc=jefflexu@linux.alibaba.com \
    --cc=kasong@tencent.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-erofs@lists.ozlabs.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=nico.pache@linux.dev \
    --cc=qi.zheng@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shakeel.butt@linux.dev \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    --cc=xiang@kernel.org \
    --cc=ying.huang@linux.alibaba.com \
    --cc=zbestahu@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.