From: "Zi Yan" <ziy@nvidia.com>
To: "Jan Kara" <jack@suse.cz>
Cc: "David Hildenbrand" <david@kernel.org>,
"Matthew Wilcox (Oracle)" <willy@infradead.org>,
"Andrew Morton" <akpm@linux-foundation.org>,
"Muchun Song" <muchun.song@linux.dev>,
"Lorenzo Stoakes" <ljs@kernel.org>,
"Liam R. Howlett" <liam@infradead.org>,
"Vlastimil Babka" <vbabka@kernel.org>,
"Mike Rapoport" <rppt@kernel.org>,
"Suren Baghdasaryan" <surenb@google.com>,
"Michal Hocko" <mhocko@suse.com>,
"Baolin Wang" <baolin.wang@linux.alibaba.com>,
"Nico Pache" <nico.pache@linux.dev>,
"Ryan Roberts" <ryan.roberts@arm.com>,
"Dev Jain" <dev.jain@arm.com>, "Barry Song" <baohua@kernel.org>,
"Lance Yang" <lance.yang@linux.dev>,
"Usama Arif" <usama.arif@linux.dev>,
"Gregory Price" <gourry@gourry.net>,
"Ying Huang" <ying.huang@linux.alibaba.com>,
"Alistair Popple" <apopple@nvidia.com>,
"Johannes Weiner" <hannes@cmpxchg.org>,
"Qi Zheng" <qi.zheng@linux.dev>,
"Shakeel Butt" <shakeel.butt@linux.dev>,
"Kairui Song" <kasong@tencent.com>, <linux-mm@kvack.org>,
<linux-kernel@vger.kernel.org>, "Gao Xiang" <xiang@kernel.org>,
"Chao Yu" <chao@kernel.org>, "Yue Hu" <zbestahu@gmail.com>,
"Jeffle Xu" <jefflexu@linux.alibaba.com>,
"Sandeep Dhavale" <dhavale@google.com>,
"Hongbo Li" <hongbohbli@tencent.com>,
"Chunhai Guo" <guochunhai@vivo.com>,
<linux-erofs@lists.ozlabs.org>, <linux-fsdevel@vger.kernel.org>
Subject: Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
Date: Wed, 05 Aug 2026 07:42:37 -0400 [thread overview]
Message-ID: <DKGZEK0AQP6R.1RTQEY44OQD65@nvidia.com> (raw)
In-Reply-To: <ivg5x7hjo4tmpd34i6wrd5cp2v56c3vorwru7h6kqxnqlcpte7@lvd7xa4ncwxu>
On Wed Aug 5, 2026 at 5:25 AM EDT, Jan Kara wrote:
> On Tue 04-08-26 13:09:46, Zi Yan wrote:
>> On Tue Aug 4, 2026 at 1:04 PM EDT, Jan Kara wrote:
>> > On Tue 04-08-26 11:54:41, Zi Yan wrote:
>> >> On Tue Aug 4, 2026 at 5:32 AM EDT, Jan Kara wrote:
>> >> > On Mon 03-08-26 12:56:36, Zi Yan wrote:
>> >> >> On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
>> >> >> > On Fri 31-07-26 22:13:30, Zi Yan wrote:
>> >> >> >> erofs needs to traverse readahead folios in reverse order to achieve
>> >> >> >> maximum performance by
>> >> >> >> 1. reading all folios from readahead_folio();
>> >> >> >> 2. storing the prior folio pointer in folio->private;
>> >> >> >> 3. traverse from the last folio to the first one.
>> >> >> >>
>> >> >> >> Add readahead_folio_reverse() to achieve the same function without using
>> >> >> >> folio->private.
>> >> >> >>
>> >> >> >> It prepares for a future commit that replaces PG_private checks with
>> >> >> >> !folio->private checks. After switching the checks, erofs's use of
>> >> >> >> folio->private without bumping folio refcount can cause unexpected
>> >> >> >> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
>> >> >> >> reachable.
>> >>
>> >> <snip>
>> >>
>> >> >>
>> >> >> The below is what I come up with. I did not add a bool to
>> >> >> readahead_control, since I think that is the decision of caller of
>> >> >> __readahead_advance(). But let me know if you disagree.
>> >> >
>> >> > The reason why I wanted bool in readahead_control is that if some code
>> >> > ends up mixing readahead_folio() with readahead_folio_last() things will
>> >> > get confused (because __readahead_advance() really wants to skip the batch
>> >> > returned from the *previous* call to readahead_folio[_last]()). With the
>> >> > bool in rac, even mixed use will properly advance the state of the
>> >> > readahead_control. I don't think mixed use is very realistic (at this
>> >> > point at least) so I'm ok with leaving that for later if you don't like it.
>> >>
>> >> Got it. I am trying to figure out your mental model of how the mix of
>> >> readahead_folio() and readahead_folio_last() works with the bool inside
>> >> ractl. By looking at readahead_folio_last() code, it is almost the same
>> >> as readahead_folio() with __readahead_folio() inlined
>> >> (__readahead_folio() is only used by readahead_folio(), so the inline
>> >> can happen without any issue). As a result, we can get rid of
>> >> readahead_folio_last(), add set_readahead_direction() to set the
>> >> embedded bool read_from_head, and use readahead_folio() only. This
>> >> removes redundant code in readahead_folio_last(). One thing I am not
>> >> certain is whether we want to
>> >>
>> >> 1. use set_readahead_direction() explicit and warn readahead_folio() if
>> >> read_from_head is not initialized, or
>> >>
>> >> 2. set read_from_head to true by default, so that only erofs needs to
>> >> call set_readahead_direction() to change read_from_head.
>> >>
>> >> The former is less confusing but changes how readahead_folio() works;
>> >> the latter is simpler but implicit read_from_head state might confuse
>> >> people at some point.
>> >
>> > My idea was: readahead_folio() will call __readahead_advance() and then set
>> > rac->forward = true. readahead_folio_last() will call __readahead_advance()
>> > and set rac->forward = false. __readahead_advance() advances from beginning
>> > / end based on rac->_forward value.
>>
>> Got it. I can do that. Just to be clear, it should be that
>> readahead_folio() first sets rac->forward = true, then calls
>> __readahead_advance(), since __readahead_advance() advances based on
>> rac->forward, right? readahead_folio_last() as well.
>
> No. I wrote "and then set" which means after and that is what I really
> wanted to say. You still don't seem to be understanding the logic of handling
> the _batch_count. _batch_count is the length of the returned batch.
> __readahead_advance() updates _index and _nr_pages to remove the folios
> returned in the last batch from the range. So _forward needs to contain
> whether the last returned batch was taken from the beginning or the end of
> the range and __readahead_advance() uses it to update current range
> accordingly (before we go and return the next batch). We cannot clobber
> _forward before calling __readahead_advance(). I hope things are clearer
> now.
Got it. Sorry I made some assumption instead of asking my question, so I
misinterpret your words. My question is who sets the initial value of
_forward? So that __readahead_advance() can update _index and _nr_pages
correctly at the first time __readahead_folio() is called?
__readahead_folio() does:
1. update _nr_pages and _index,
2. return NULL if _nr_pages is 0 and set _batch_count to 0,
3. return folio using xa_load and set _batch_count to folio_nr_pages().
after the change:
1. call __readahead_advance() to update _index, _nr_pages, and
_batch_count based on _forward,
2. update _forward to true, since it is __readahead_folio()
3. return NULL or folio based on _nr_pages.
Then the first time __readahead_folio() is called, who sets _forward to
make 1 work correctly?
Thanks.
--
Best Regards,
Yan, Zi
next prev parent reply other threads:[~2026-08-05 11:42 UTC|newest]
Thread overview: 63+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-01 2:13 [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
2026-08-01 2:13 ` [f2fs-dev] " Zi Yan via Linux-f2fs-devel
2026-08-01 2:13 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 01/14] mm/zsmalloc: replace PG_private with pointer comparison Zi Yan
2026-08-01 14:12 ` Usama Arif
2026-08-01 23:49 ` Zi Yan
2026-08-02 12:05 ` Usama Arif
2026-08-02 18:38 ` Zi Yan
2026-08-03 15:04 ` Johannes Weiner
2026-08-03 15:34 ` Zi Yan
2026-08-03 16:53 ` Johannes Weiner
2026-08-03 21:05 ` Zi Yan
2026-08-04 0:51 ` Johannes Weiner
2026-08-04 6:37 ` Sergey Senozhatsky
2026-08-04 15:28 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 02/14] perf/ring_buffer: stop using PG_private as AUX page high-order marker Zi Yan
2026-08-01 14:32 ` Usama Arif
2026-08-02 1:20 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 03/14] xen/grant-table: stop setting PG_private on pages for grant mapping Zi Yan
2026-08-01 14:42 ` Usama Arif
2026-08-02 1:24 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 04/14] fs/crypto: stop setting PG_private on bounce page Zi Yan
2026-08-01 14:52 ` Usama Arif
2026-08-02 1:29 ` Zi Yan
2026-08-03 18:40 ` Eric Biggers
2026-08-01 2:13 ` [PATCH RFC 05/14] mm/hugetlb: use direct assignment instead of folio_change_private() Zi Yan
2026-08-02 12:16 ` Usama Arif
2026-08-02 18:38 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 06/14] fs/f2fs: stop using PG_private Zi Yan
2026-08-01 2:13 ` [f2fs-dev] " Zi Yan via Linux-f2fs-devel
2026-08-03 11:17 ` Chao Yu via Linux-f2fs-devel
2026-08-03 11:17 ` Chao Yu
2026-08-03 15:48 ` [f2fs-dev] " Usama Arif
2026-08-03 15:48 ` Usama Arif
2026-08-01 2:13 ` [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private Zi Yan
2026-08-03 9:54 ` Jan Kara
2026-08-03 16:56 ` Zi Yan
2026-08-04 9:32 ` Jan Kara
2026-08-04 15:54 ` Zi Yan
2026-08-04 17:04 ` Jan Kara
2026-08-04 17:09 ` Zi Yan
2026-08-05 2:37 ` Zi Yan
2026-08-05 9:25 ` Jan Kara
2026-08-05 11:42 ` Zi Yan [this message]
2026-08-05 13:51 ` Zi Yan
2026-08-05 16:10 ` Jan Kara
2026-08-03 23:55 ` Gao Xiang
2026-08-01 2:13 ` [PATCH RFC 08/14] fs/erofs: use folio_attach/detach_private() instead of direct assignment Zi Yan
2026-08-03 23:40 ` Gao Xiang
2026-08-05 2:41 ` Zi Yan
2026-08-05 4:17 ` Gao Xiang
2026-08-01 2:13 ` [PATCH RFC 09/14] mm/page-flags: check page/folio->private instead of PG_private Zi Yan
2026-08-01 2:13 ` [PATCH RFC 10/14] mm/page-flags: introduce folio_test_fs_private() Zi Yan
2026-08-01 2:13 ` [PATCH RFC 11/14] treewide: remove folio_set/clear_private() Zi Yan
2026-08-01 2:13 ` [PATCH RFC 12/14] treewide: replace PagePrivate() with page_private() Zi Yan
2026-08-01 2:13 ` [PATCH RFC 13/14] treewide: adjust comments on PagePrivate and PG_private Zi Yan
2026-08-01 2:13 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 14/14] mm/page-flags: remove PG_private Zi Yan
2026-08-03 9:07 ` [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Jürgen Groß
2026-08-03 9:07 ` Jürgen Groß
2026-08-03 18:13 ` Zi Yan
2026-08-03 18:13 ` [f2fs-dev] " Zi Yan via Linux-f2fs-devel
2026-08-03 18:13 ` Zi Yan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DKGZEK0AQP6R.1RTQEY44OQD65@nvidia.com \
--to=ziy@nvidia.com \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=chao@kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=dhavale@google.com \
--cc=gourry@gourry.net \
--cc=guochunhai@vivo.com \
--cc=hannes@cmpxchg.org \
--cc=hongbohbli@tencent.com \
--cc=jack@suse.cz \
--cc=jefflexu@linux.alibaba.com \
--cc=kasong@tencent.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-erofs@lists.ozlabs.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=nico.pache@linux.dev \
--cc=qi.zheng@linux.dev \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=xiang@kernel.org \
--cc=ying.huang@linux.alibaba.com \
--cc=zbestahu@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.