From: "Zi Yan" <ziy@nvidia.com>
To: "Jan Kara" <jack@suse.cz>
Cc: "David Hildenbrand" <david@kernel.org>,
"Matthew Wilcox (Oracle)" <willy@infradead.org>,
"Andrew Morton" <akpm@linux-foundation.org>,
"Muchun Song" <muchun.song@linux.dev>,
"Lorenzo Stoakes" <ljs@kernel.org>,
"Liam R. Howlett" <liam@infradead.org>,
"Vlastimil Babka" <vbabka@kernel.org>,
"Mike Rapoport" <rppt@kernel.org>,
"Suren Baghdasaryan" <surenb@google.com>,
"Michal Hocko" <mhocko@suse.com>,
"Baolin Wang" <baolin.wang@linux.alibaba.com>,
"Nico Pache" <nico.pache@linux.dev>,
"Ryan Roberts" <ryan.roberts@arm.com>,
"Dev Jain" <dev.jain@arm.com>, "Barry Song" <baohua@kernel.org>,
"Lance Yang" <lance.yang@linux.dev>,
"Usama Arif" <usama.arif@linux.dev>,
"Gregory Price" <gourry@gourry.net>,
"Ying Huang" <ying.huang@linux.alibaba.com>,
"Alistair Popple" <apopple@nvidia.com>,
"Johannes Weiner" <hannes@cmpxchg.org>,
"Qi Zheng" <qi.zheng@linux.dev>,
"Shakeel Butt" <shakeel.butt@linux.dev>,
"Kairui Song" <kasong@tencent.com>, <linux-mm@kvack.org>,
<linux-kernel@vger.kernel.org>, "Gao Xiang" <xiang@kernel.org>,
"Chao Yu" <chao@kernel.org>, "Yue Hu" <zbestahu@gmail.com>,
"Jeffle Xu" <jefflexu@linux.alibaba.com>,
"Sandeep Dhavale" <dhavale@google.com>,
"Hongbo Li" <hongbohbli@tencent.com>,
"Chunhai Guo" <guochunhai@vivo.com>,
<linux-erofs@lists.ozlabs.org>, <linux-fsdevel@vger.kernel.org>
Subject: Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
Date: Mon, 03 Aug 2026 12:56:36 -0400 [thread overview]
Message-ID: <DKFGTVA0XD01.OZYNYUOU9FTX@nvidia.com> (raw)
In-Reply-To: <332rknj4vo3cfhvfhhlf6pvg37s3lbrnzbbnv4swa6gctsiu6a@ndotgokvnglc>
On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
> On Fri 31-07-26 22:13:30, Zi Yan wrote:
>> erofs needs to traverse readahead folios in reverse order to achieve
>> maximum performance by
>> 1. reading all folios from readahead_folio();
>> 2. storing the prior folio pointer in folio->private;
>> 3. traverse from the last folio to the first one.
>>
>> Add readahead_folio_reverse() to achieve the same function without using
>> folio->private.
>>
>> It prepares for a future commit that replaces PG_private checks with
>> !folio->private checks. After switching the checks, erofs's use of
>> folio->private without bumping folio refcount can cause unexpected
>> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
>> reachable.
>>
>> No funtional change intended.
>>
>> Assisted-by: Claude:claude-opus-4-8
>> Assisted-by: Codex:gpt-5
>> Signed-off-by: Zi Yan <ziy@nvidia.com>
>> To: Gao Xiang <xiang@kernel.org>
>> To: Chao Yu <chao@kernel.org>
>> To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
>> To: Jan Kara <jack@suse.cz>
>> Cc: Yue Hu <zbestahu@gmail.com>
>> Cc: Jeffle Xu <jefflexu@linux.alibaba.com>
>> Cc: Sandeep Dhavale <dhavale@google.com>
>> Cc: Hongbo Li <hongbohbli@tencent.com>
>> Cc: Chunhai Guo <guochunhai@vivo.com>
>> Cc: linux-erofs@lists.ozlabs.org
>> Cc: linux-kernel@vger.kernel.org
>> Cc: linux-fsdevel@vger.kernel.org
>> Cc: linux-mm@kvack.org
>
> One comment regarding the generic infrastructure below.
>
>> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
>> index 4e8b2b29f6d3e..90904a4d173b7 100644
>> --- a/include/linux/pagemap.h
>> +++ b/include/linux/pagemap.h
>> @@ -1549,6 +1549,37 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
>> return folio;
>> }
>>
>> +/**
>> + * readahead_folio_reverse - Get the next folio to read, from the tail.
>> + * @ractl: The current readahead request.
>> + *
>> + * Like readahead_folio(), but walks the range back-to-front. The folio is
>> + * returned locked with its refcount dropped; the caller unlocks it once I/O
>> + * completes. Compound folios are returned once, at their head index.
>> + *
>> + * Context: The folio is locked.
>> + * Return: A pointer to the next folio, or %NULL when done.
>> + */
>> +static inline struct folio *readahead_folio_reverse(struct readahead_control *ractl)
>> +{
>> + struct folio *folio;
>> +
>> + if (!ractl->_nr_pages)
>> + return NULL;
>> +
>> + /* xa_load() follows sibling entries, so a tail index returns the head */
>> + folio = xa_load(&ractl->mapping->i_pages,
>> + ractl->_index + ractl->_nr_pages - 1);
>> + VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
>> +
>> + /* Shrink the window from the tail down to this folio's head index */
>> + ractl->_nr_pages = folio->index - ractl->_index;
>> + ractl->_batch_count = 0;
>
> Thanks for the patch! Currently there's the invariant that the returned
> folio is still inside the _index .. _index+_nr_pages range. I think when we
> are providing a generic helper, we should keep that to make code more
> robust for the future when more people start using it.
Definitely.
>
> What I'd suggest doing is add bool in struct readahead_control telling
> whether the last folio (batch) was taken from the head or tail of the
> range, advance _nr_pages and _index accordingly in the functions returning
> folios (probably hide this in a helper function __readahead_advance()
> because it will be used in 3 places) and maybe call this new function
> readahead_folio_last() instead of _reverse() (but I have only a slight
> preference here so .._reverse() is ok with me if other people prefer it).
The below is what I come up with. I did not add a bool to
readahead_control, since I think that is the decision of caller of
__readahead_advance(). But let me know if you disagree.
diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
index b59f2745a8e72..23f423c22ac8c 100644
--- a/fs/erofs/zdata.c
+++ b/fs/erofs/zdata.c
@@ -1908,8 +1908,8 @@ static void z_erofs_readahead(struct readahead_control *rac)
trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
z_erofs_pcluster_readmore(&f, rac, true);
- /* traverse in reverse order for best metadata I/O performance */
- while ((folio = readahead_folio_reverse(rac))) {
+ /* traverse from last to first for best metadata I/O performance */
+ while ((folio = readahead_folio_last(rac))) {
err = z_erofs_scan_folio(&f, folio, true);
if (err && err != -EINTR)
erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 5ca5aa365f319..2cc3de5594518 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -1510,13 +1510,21 @@ void page_cache_async_readahead(struct address_space *mapping,
page_cache_async_ra(&ractl, folio, req_count);
}
+static inline void __readahead_advance(struct readahead_control *rac,
+ bool read_from_head)
+{
+ if (read_from_head)
+ rac->_index += rac->_batch_count;
+
+ rac->_nr_pages -= rac->_batch_count;
+}
+
static inline struct folio *__readahead_folio(struct readahead_control *ractl)
{
- struct folio *folio;
+ struct folio *folio = NULL;
BUG_ON(ractl->_batch_count > ractl->_nr_pages);
- ractl->_nr_pages -= ractl->_batch_count;
- ractl->_index += ractl->_batch_count;
+ __readahead_advance(ractl, /* read_from_head= */ true);
if (!ractl->_nr_pages) {
ractl->_batch_count = 0;
@@ -1548,7 +1556,7 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
}
/**
- * readahead_folio_reverse - Get the next folio to read, from the tail.
+ * readahead_folio_last - Get the next folio to read, from the tail.
* @ractl: The current readahead request.
*
* Like readahead_folio(), but walks the range back-to-front. The folio is
@@ -1558,21 +1566,24 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
* Context: The folio is locked.
* Return: A pointer to the next folio, or %NULL when done.
*/
-static inline struct folio *readahead_folio_reverse(struct readahead_control *ractl)
+static inline struct folio *readahead_folio_last(struct readahead_control *ractl)
{
struct folio *folio;
- if (!ractl->_nr_pages)
+ /* Shrink the window from the tail down to this folio's head index */
+ __readahead_advance(ractl, /* read_from_head= */ false);
+
+ if (!ractl->_nr_pages) {
+ ractl->_batch_count = 0;
return NULL;
+ }
/* xa_load() follows sibling entries, so a tail index returns the head */
folio = xa_load(&ractl->mapping->i_pages,
ractl->_index + ractl->_nr_pages - 1);
VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
- /* Shrink the window from the tail down to this folio's head index */
- ractl->_nr_pages = folio->index - ractl->_index;
- ractl->_batch_count = 0;
+ ractl->_batch_count = folio_nr_pages(folio);
folio_put(folio);
return folio;
@@ -1583,11 +1594,10 @@ static inline unsigned int __readahead_batch(struct readahead_control *rac,
{
unsigned int i = 0;
XA_STATE(xas, &rac->mapping->i_pages, 0);
- struct folio *folio;
+ struct folio *folio = NULL;
BUG_ON(rac->_batch_count > rac->_nr_pages);
- rac->_nr_pages -= rac->_batch_count;
- rac->_index += rac->_batch_count;
+ __readahead_advance(rac, /* read_from_head= */ true);
rac->_batch_count = 0;
xas_set(&xas, rac->_index);
--
Best Regards,
Yan, Zi
next prev parent reply other threads:[~2026-08-03 16:56 UTC|newest]
Thread overview: 62+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-01 2:13 [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
2026-08-01 2:13 ` [f2fs-dev] " Zi Yan via Linux-f2fs-devel
2026-08-01 2:13 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 01/14] mm/zsmalloc: replace PG_private with pointer comparison Zi Yan
2026-08-01 14:12 ` Usama Arif
2026-08-01 23:49 ` Zi Yan
2026-08-02 12:05 ` Usama Arif
2026-08-02 18:38 ` Zi Yan
2026-08-03 15:04 ` Johannes Weiner
2026-08-03 15:34 ` Zi Yan
2026-08-03 16:53 ` Johannes Weiner
2026-08-03 21:05 ` Zi Yan
2026-08-04 0:51 ` Johannes Weiner
2026-08-04 6:37 ` Sergey Senozhatsky
2026-08-04 15:28 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 02/14] perf/ring_buffer: stop using PG_private as AUX page high-order marker Zi Yan
2026-08-01 14:32 ` Usama Arif
2026-08-02 1:20 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 03/14] xen/grant-table: stop setting PG_private on pages for grant mapping Zi Yan
2026-08-01 14:42 ` Usama Arif
2026-08-02 1:24 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 04/14] fs/crypto: stop setting PG_private on bounce page Zi Yan
2026-08-01 14:52 ` Usama Arif
2026-08-02 1:29 ` Zi Yan
2026-08-03 18:40 ` Eric Biggers
2026-08-01 2:13 ` [PATCH RFC 05/14] mm/hugetlb: use direct assignment instead of folio_change_private() Zi Yan
2026-08-02 12:16 ` Usama Arif
2026-08-02 18:38 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 06/14] fs/f2fs: stop using PG_private Zi Yan
2026-08-01 2:13 ` [f2fs-dev] " Zi Yan via Linux-f2fs-devel
2026-08-03 11:17 ` Chao Yu via Linux-f2fs-devel
2026-08-03 11:17 ` Chao Yu
2026-08-03 15:48 ` [f2fs-dev] " Usama Arif
2026-08-03 15:48 ` Usama Arif
2026-08-01 2:13 ` [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private Zi Yan
2026-08-03 9:54 ` Jan Kara
2026-08-03 16:56 ` Zi Yan [this message]
2026-08-04 9:32 ` Jan Kara
2026-08-04 15:54 ` Zi Yan
2026-08-04 17:04 ` Jan Kara
2026-08-04 17:09 ` Zi Yan
2026-08-05 2:37 ` Zi Yan
2026-08-05 9:25 ` Jan Kara
2026-08-05 11:42 ` Zi Yan
2026-08-05 13:51 ` Zi Yan
2026-08-03 23:55 ` Gao Xiang
2026-08-01 2:13 ` [PATCH RFC 08/14] fs/erofs: use folio_attach/detach_private() instead of direct assignment Zi Yan
2026-08-03 23:40 ` Gao Xiang
2026-08-05 2:41 ` Zi Yan
2026-08-05 4:17 ` Gao Xiang
2026-08-01 2:13 ` [PATCH RFC 09/14] mm/page-flags: check page/folio->private instead of PG_private Zi Yan
2026-08-01 2:13 ` [PATCH RFC 10/14] mm/page-flags: introduce folio_test_fs_private() Zi Yan
2026-08-01 2:13 ` [PATCH RFC 11/14] treewide: remove folio_set/clear_private() Zi Yan
2026-08-01 2:13 ` [PATCH RFC 12/14] treewide: replace PagePrivate() with page_private() Zi Yan
2026-08-01 2:13 ` [PATCH RFC 13/14] treewide: adjust comments on PagePrivate and PG_private Zi Yan
2026-08-01 2:13 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 14/14] mm/page-flags: remove PG_private Zi Yan
2026-08-03 9:07 ` [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Jürgen Groß
2026-08-03 9:07 ` Jürgen Groß
2026-08-03 18:13 ` Zi Yan
2026-08-03 18:13 ` [f2fs-dev] " Zi Yan via Linux-f2fs-devel
2026-08-03 18:13 ` Zi Yan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DKFGTVA0XD01.OZYNYUOU9FTX@nvidia.com \
--to=ziy@nvidia.com \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=chao@kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=dhavale@google.com \
--cc=gourry@gourry.net \
--cc=guochunhai@vivo.com \
--cc=hannes@cmpxchg.org \
--cc=hongbohbli@tencent.com \
--cc=jack@suse.cz \
--cc=jefflexu@linux.alibaba.com \
--cc=kasong@tencent.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-erofs@lists.ozlabs.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=nico.pache@linux.dev \
--cc=qi.zheng@linux.dev \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=xiang@kernel.org \
--cc=ying.huang@linux.alibaba.com \
--cc=zbestahu@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.