* [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
2026-08-01 2:13 [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
@ 2026-08-01 2:13 ` Zi Yan
2026-08-03 9:54 ` Jan Kara
2026-08-03 23:55 ` Gao Xiang
2026-08-01 2:13 ` [PATCH RFC 09/14] mm/page-flags: check page/folio->private instead of PG_private Zi Yan
` (4 subsequent siblings)
5 siblings, 2 replies; 15+ messages in thread
From: Zi Yan @ 2026-08-01 2:13 UTC (permalink / raw)
To: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song
Cc: linux-mm, linux-kernel, Zi Yan, Gao Xiang, Chao Yu, Jan Kara,
Yue Hu, Jeffle Xu, Sandeep Dhavale, Hongbo Li, Chunhai Guo,
linux-erofs, linux-fsdevel
erofs needs to traverse readahead folios in reverse order to achieve
maximum performance by
1. reading all folios from readahead_folio();
2. storing the prior folio pointer in folio->private;
3. traverse from the last folio to the first one.
Add readahead_folio_reverse() to achieve the same function without using
folio->private.
It prepares for a future commit that replaces PG_private checks with
!folio->private checks. After switching the checks, erofs's use of
folio->private without bumping folio refcount can cause unexpected
outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
reachable.
No funtional change intended.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Gao Xiang <xiang@kernel.org>
To: Chao Yu <chao@kernel.org>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
Cc: Yue Hu <zbestahu@gmail.com>
Cc: Jeffle Xu <jefflexu@linux.alibaba.com>
Cc: Sandeep Dhavale <dhavale@google.com>
Cc: Hongbo Li <hongbohbli@tencent.com>
Cc: Chunhai Guo <guochunhai@vivo.com>
Cc: linux-erofs@lists.ozlabs.org
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
---
fs/erofs/zdata.c | 11 ++---------
include/linux/pagemap.h | 31 +++++++++++++++++++++++++++++++
2 files changed, 33 insertions(+), 9 deletions(-)
diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
index 74520e9102596..b59f2745a8e72 100644
--- a/fs/erofs/zdata.c
+++ b/fs/erofs/zdata.c
@@ -1902,21 +1902,14 @@ static void z_erofs_readahead(struct readahead_control *rac)
struct inode *realinode = erofs_real_inode(sharedinode, &need_iput);
Z_EROFS_DEFINE_FRONTEND(f, realinode, sharedinode, readahead_pos(rac));
unsigned int nrpages = readahead_count(rac);
- struct folio *head = NULL, *folio;
+ struct folio *folio;
int err;
trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
z_erofs_pcluster_readmore(&f, rac, true);
- while ((folio = readahead_folio(rac))) {
- folio->private = head;
- head = folio;
- }
/* traverse in reverse order for best metadata I/O performance */
- while (head) {
- folio = head;
- head = folio_get_private(folio);
-
+ while ((folio = readahead_folio_reverse(rac))) {
err = z_erofs_scan_folio(&f, folio, true);
if (err && err != -EINTR)
erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 4e8b2b29f6d3e..90904a4d173b7 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -1549,6 +1549,37 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
return folio;
}
+/**
+ * readahead_folio_reverse - Get the next folio to read, from the tail.
+ * @ractl: The current readahead request.
+ *
+ * Like readahead_folio(), but walks the range back-to-front. The folio is
+ * returned locked with its refcount dropped; the caller unlocks it once I/O
+ * completes. Compound folios are returned once, at their head index.
+ *
+ * Context: The folio is locked.
+ * Return: A pointer to the next folio, or %NULL when done.
+ */
+static inline struct folio *readahead_folio_reverse(struct readahead_control *ractl)
+{
+ struct folio *folio;
+
+ if (!ractl->_nr_pages)
+ return NULL;
+
+ /* xa_load() follows sibling entries, so a tail index returns the head */
+ folio = xa_load(&ractl->mapping->i_pages,
+ ractl->_index + ractl->_nr_pages - 1);
+ VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
+
+ /* Shrink the window from the tail down to this folio's head index */
+ ractl->_nr_pages = folio->index - ractl->_index;
+ ractl->_batch_count = 0;
+
+ folio_put(folio);
+ return folio;
+}
+
static inline unsigned int __readahead_batch(struct readahead_control *rac,
struct page **array, unsigned int array_sz)
{
--
2.53.0
^ permalink raw reply related [flat|nested] 15+ messages in thread* Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
2026-08-01 2:13 ` [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private Zi Yan
@ 2026-08-03 9:54 ` Jan Kara
2026-08-03 16:56 ` Zi Yan
2026-08-03 23:55 ` Gao Xiang
1 sibling, 1 reply; 15+ messages in thread
From: Jan Kara @ 2026-08-03 9:54 UTC (permalink / raw)
To: Zi Yan
Cc: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song, linux-mm,
linux-kernel, Gao Xiang, Chao Yu, Jan Kara, Yue Hu, Jeffle Xu,
Sandeep Dhavale, Hongbo Li, Chunhai Guo, linux-erofs,
linux-fsdevel
On Fri 31-07-26 22:13:30, Zi Yan wrote:
> erofs needs to traverse readahead folios in reverse order to achieve
> maximum performance by
> 1. reading all folios from readahead_folio();
> 2. storing the prior folio pointer in folio->private;
> 3. traverse from the last folio to the first one.
>
> Add readahead_folio_reverse() to achieve the same function without using
> folio->private.
>
> It prepares for a future commit that replaces PG_private checks with
> !folio->private checks. After switching the checks, erofs's use of
> folio->private without bumping folio refcount can cause unexpected
> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
> reachable.
>
> No funtional change intended.
>
> Assisted-by: Claude:claude-opus-4-8
> Assisted-by: Codex:gpt-5
> Signed-off-by: Zi Yan <ziy@nvidia.com>
> To: Gao Xiang <xiang@kernel.org>
> To: Chao Yu <chao@kernel.org>
> To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
> To: Jan Kara <jack@suse.cz>
> Cc: Yue Hu <zbestahu@gmail.com>
> Cc: Jeffle Xu <jefflexu@linux.alibaba.com>
> Cc: Sandeep Dhavale <dhavale@google.com>
> Cc: Hongbo Li <hongbohbli@tencent.com>
> Cc: Chunhai Guo <guochunhai@vivo.com>
> Cc: linux-erofs@lists.ozlabs.org
> Cc: linux-kernel@vger.kernel.org
> Cc: linux-fsdevel@vger.kernel.org
> Cc: linux-mm@kvack.org
One comment regarding the generic infrastructure below.
> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
> index 4e8b2b29f6d3e..90904a4d173b7 100644
> --- a/include/linux/pagemap.h
> +++ b/include/linux/pagemap.h
> @@ -1549,6 +1549,37 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
> return folio;
> }
>
> +/**
> + * readahead_folio_reverse - Get the next folio to read, from the tail.
> + * @ractl: The current readahead request.
> + *
> + * Like readahead_folio(), but walks the range back-to-front. The folio is
> + * returned locked with its refcount dropped; the caller unlocks it once I/O
> + * completes. Compound folios are returned once, at their head index.
> + *
> + * Context: The folio is locked.
> + * Return: A pointer to the next folio, or %NULL when done.
> + */
> +static inline struct folio *readahead_folio_reverse(struct readahead_control *ractl)
> +{
> + struct folio *folio;
> +
> + if (!ractl->_nr_pages)
> + return NULL;
> +
> + /* xa_load() follows sibling entries, so a tail index returns the head */
> + folio = xa_load(&ractl->mapping->i_pages,
> + ractl->_index + ractl->_nr_pages - 1);
> + VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
> +
> + /* Shrink the window from the tail down to this folio's head index */
> + ractl->_nr_pages = folio->index - ractl->_index;
> + ractl->_batch_count = 0;
Thanks for the patch! Currently there's the invariant that the returned
folio is still inside the _index .. _index+_nr_pages range. I think when we
are providing a generic helper, we should keep that to make code more
robust for the future when more people start using it.
What I'd suggest doing is add bool in struct readahead_control telling
whether the last folio (batch) was taken from the head or tail of the
range, advance _nr_pages and _index accordingly in the functions returning
folios (probably hide this in a helper function __readahead_advance()
because it will be used in 3 places) and maybe call this new function
readahead_folio_last() instead of _reverse() (but I have only a slight
preference here so .._reverse() is ok with me if other people prefer it).
Honza
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
2026-08-03 9:54 ` Jan Kara
@ 2026-08-03 16:56 ` Zi Yan
2026-08-04 9:32 ` Jan Kara
0 siblings, 1 reply; 15+ messages in thread
From: Zi Yan @ 2026-08-03 16:56 UTC (permalink / raw)
To: Jan Kara
Cc: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song, linux-mm,
linux-kernel, Gao Xiang, Chao Yu, Yue Hu, Jeffle Xu,
Sandeep Dhavale, Hongbo Li, Chunhai Guo, linux-erofs,
linux-fsdevel
On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
> On Fri 31-07-26 22:13:30, Zi Yan wrote:
>> erofs needs to traverse readahead folios in reverse order to achieve
>> maximum performance by
>> 1. reading all folios from readahead_folio();
>> 2. storing the prior folio pointer in folio->private;
>> 3. traverse from the last folio to the first one.
>>
>> Add readahead_folio_reverse() to achieve the same function without using
>> folio->private.
>>
>> It prepares for a future commit that replaces PG_private checks with
>> !folio->private checks. After switching the checks, erofs's use of
>> folio->private without bumping folio refcount can cause unexpected
>> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
>> reachable.
>>
>> No funtional change intended.
>>
>> Assisted-by: Claude:claude-opus-4-8
>> Assisted-by: Codex:gpt-5
>> Signed-off-by: Zi Yan <ziy@nvidia.com>
>> To: Gao Xiang <xiang@kernel.org>
>> To: Chao Yu <chao@kernel.org>
>> To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
>> To: Jan Kara <jack@suse.cz>
>> Cc: Yue Hu <zbestahu@gmail.com>
>> Cc: Jeffle Xu <jefflexu@linux.alibaba.com>
>> Cc: Sandeep Dhavale <dhavale@google.com>
>> Cc: Hongbo Li <hongbohbli@tencent.com>
>> Cc: Chunhai Guo <guochunhai@vivo.com>
>> Cc: linux-erofs@lists.ozlabs.org
>> Cc: linux-kernel@vger.kernel.org
>> Cc: linux-fsdevel@vger.kernel.org
>> Cc: linux-mm@kvack.org
>
> One comment regarding the generic infrastructure below.
>
>> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
>> index 4e8b2b29f6d3e..90904a4d173b7 100644
>> --- a/include/linux/pagemap.h
>> +++ b/include/linux/pagemap.h
>> @@ -1549,6 +1549,37 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
>> return folio;
>> }
>>
>> +/**
>> + * readahead_folio_reverse - Get the next folio to read, from the tail.
>> + * @ractl: The current readahead request.
>> + *
>> + * Like readahead_folio(), but walks the range back-to-front. The folio is
>> + * returned locked with its refcount dropped; the caller unlocks it once I/O
>> + * completes. Compound folios are returned once, at their head index.
>> + *
>> + * Context: The folio is locked.
>> + * Return: A pointer to the next folio, or %NULL when done.
>> + */
>> +static inline struct folio *readahead_folio_reverse(struct readahead_control *ractl)
>> +{
>> + struct folio *folio;
>> +
>> + if (!ractl->_nr_pages)
>> + return NULL;
>> +
>> + /* xa_load() follows sibling entries, so a tail index returns the head */
>> + folio = xa_load(&ractl->mapping->i_pages,
>> + ractl->_index + ractl->_nr_pages - 1);
>> + VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
>> +
>> + /* Shrink the window from the tail down to this folio's head index */
>> + ractl->_nr_pages = folio->index - ractl->_index;
>> + ractl->_batch_count = 0;
>
> Thanks for the patch! Currently there's the invariant that the returned
> folio is still inside the _index .. _index+_nr_pages range. I think when we
> are providing a generic helper, we should keep that to make code more
> robust for the future when more people start using it.
Definitely.
>
> What I'd suggest doing is add bool in struct readahead_control telling
> whether the last folio (batch) was taken from the head or tail of the
> range, advance _nr_pages and _index accordingly in the functions returning
> folios (probably hide this in a helper function __readahead_advance()
> because it will be used in 3 places) and maybe call this new function
> readahead_folio_last() instead of _reverse() (but I have only a slight
> preference here so .._reverse() is ok with me if other people prefer it).
The below is what I come up with. I did not add a bool to
readahead_control, since I think that is the decision of caller of
__readahead_advance(). But let me know if you disagree.
diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
index b59f2745a8e72..23f423c22ac8c 100644
--- a/fs/erofs/zdata.c
+++ b/fs/erofs/zdata.c
@@ -1908,8 +1908,8 @@ static void z_erofs_readahead(struct readahead_control *rac)
trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
z_erofs_pcluster_readmore(&f, rac, true);
- /* traverse in reverse order for best metadata I/O performance */
- while ((folio = readahead_folio_reverse(rac))) {
+ /* traverse from last to first for best metadata I/O performance */
+ while ((folio = readahead_folio_last(rac))) {
err = z_erofs_scan_folio(&f, folio, true);
if (err && err != -EINTR)
erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 5ca5aa365f319..2cc3de5594518 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -1510,13 +1510,21 @@ void page_cache_async_readahead(struct address_space *mapping,
page_cache_async_ra(&ractl, folio, req_count);
}
+static inline void __readahead_advance(struct readahead_control *rac,
+ bool read_from_head)
+{
+ if (read_from_head)
+ rac->_index += rac->_batch_count;
+
+ rac->_nr_pages -= rac->_batch_count;
+}
+
static inline struct folio *__readahead_folio(struct readahead_control *ractl)
{
- struct folio *folio;
+ struct folio *folio = NULL;
BUG_ON(ractl->_batch_count > ractl->_nr_pages);
- ractl->_nr_pages -= ractl->_batch_count;
- ractl->_index += ractl->_batch_count;
+ __readahead_advance(ractl, /* read_from_head= */ true);
if (!ractl->_nr_pages) {
ractl->_batch_count = 0;
@@ -1548,7 +1556,7 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
}
/**
- * readahead_folio_reverse - Get the next folio to read, from the tail.
+ * readahead_folio_last - Get the next folio to read, from the tail.
* @ractl: The current readahead request.
*
* Like readahead_folio(), but walks the range back-to-front. The folio is
@@ -1558,21 +1566,24 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
* Context: The folio is locked.
* Return: A pointer to the next folio, or %NULL when done.
*/
-static inline struct folio *readahead_folio_reverse(struct readahead_control *ractl)
+static inline struct folio *readahead_folio_last(struct readahead_control *ractl)
{
struct folio *folio;
- if (!ractl->_nr_pages)
+ /* Shrink the window from the tail down to this folio's head index */
+ __readahead_advance(ractl, /* read_from_head= */ false);
+
+ if (!ractl->_nr_pages) {
+ ractl->_batch_count = 0;
return NULL;
+ }
/* xa_load() follows sibling entries, so a tail index returns the head */
folio = xa_load(&ractl->mapping->i_pages,
ractl->_index + ractl->_nr_pages - 1);
VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
- /* Shrink the window from the tail down to this folio's head index */
- ractl->_nr_pages = folio->index - ractl->_index;
- ractl->_batch_count = 0;
+ ractl->_batch_count = folio_nr_pages(folio);
folio_put(folio);
return folio;
@@ -1583,11 +1594,10 @@ static inline unsigned int __readahead_batch(struct readahead_control *rac,
{
unsigned int i = 0;
XA_STATE(xas, &rac->mapping->i_pages, 0);
- struct folio *folio;
+ struct folio *folio = NULL;
BUG_ON(rac->_batch_count > rac->_nr_pages);
- rac->_nr_pages -= rac->_batch_count;
- rac->_index += rac->_batch_count;
+ __readahead_advance(rac, /* read_from_head= */ true);
rac->_batch_count = 0;
xas_set(&xas, rac->_index);
--
Best Regards,
Yan, Zi
^ permalink raw reply related [flat|nested] 15+ messages in thread* Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
2026-08-03 16:56 ` Zi Yan
@ 2026-08-04 9:32 ` Jan Kara
2026-08-04 15:54 ` Zi Yan
0 siblings, 1 reply; 15+ messages in thread
From: Jan Kara @ 2026-08-04 9:32 UTC (permalink / raw)
To: Zi Yan
Cc: Jan Kara, David Hildenbrand, Matthew Wilcox (Oracle),
Andrew Morton, Muchun Song, Lorenzo Stoakes, Liam R. Howlett,
Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
Baolin Wang, Nico Pache, Ryan Roberts, Dev Jain, Barry Song,
Lance Yang, Usama Arif, Gregory Price, Ying Huang,
Alistair Popple, Johannes Weiner, Qi Zheng, Shakeel Butt,
Kairui Song, linux-mm, linux-kernel, Gao Xiang, Chao Yu, Yue Hu,
Jeffle Xu, Sandeep Dhavale, Hongbo Li, Chunhai Guo, linux-erofs,
linux-fsdevel
On Mon 03-08-26 12:56:36, Zi Yan wrote:
> On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
> > On Fri 31-07-26 22:13:30, Zi Yan wrote:
> >> erofs needs to traverse readahead folios in reverse order to achieve
> >> maximum performance by
> >> 1. reading all folios from readahead_folio();
> >> 2. storing the prior folio pointer in folio->private;
> >> 3. traverse from the last folio to the first one.
> >>
> >> Add readahead_folio_reverse() to achieve the same function without using
> >> folio->private.
> >>
> >> It prepares for a future commit that replaces PG_private checks with
> >> !folio->private checks. After switching the checks, erofs's use of
> >> folio->private without bumping folio refcount can cause unexpected
> >> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
> >> reachable.
> >>
> >> No funtional change intended.
> >>
> >> Assisted-by: Claude:claude-opus-4-8
> >> Assisted-by: Codex:gpt-5
> >> Signed-off-by: Zi Yan <ziy@nvidia.com>
> >> To: Gao Xiang <xiang@kernel.org>
> >> To: Chao Yu <chao@kernel.org>
> >> To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
> >> To: Jan Kara <jack@suse.cz>
> >> Cc: Yue Hu <zbestahu@gmail.com>
> >> Cc: Jeffle Xu <jefflexu@linux.alibaba.com>
> >> Cc: Sandeep Dhavale <dhavale@google.com>
> >> Cc: Hongbo Li <hongbohbli@tencent.com>
> >> Cc: Chunhai Guo <guochunhai@vivo.com>
> >> Cc: linux-erofs@lists.ozlabs.org
> >> Cc: linux-kernel@vger.kernel.org
> >> Cc: linux-fsdevel@vger.kernel.org
> >> Cc: linux-mm@kvack.org
> >
> > One comment regarding the generic infrastructure below.
> >
> >> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
> >> index 4e8b2b29f6d3e..90904a4d173b7 100644
> >> --- a/include/linux/pagemap.h
> >> +++ b/include/linux/pagemap.h
> >> @@ -1549,6 +1549,37 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
> >> return folio;
> >> }
> >>
> >> +/**
> >> + * readahead_folio_reverse - Get the next folio to read, from the tail.
> >> + * @ractl: The current readahead request.
> >> + *
> >> + * Like readahead_folio(), but walks the range back-to-front. The folio is
> >> + * returned locked with its refcount dropped; the caller unlocks it once I/O
> >> + * completes. Compound folios are returned once, at their head index.
> >> + *
> >> + * Context: The folio is locked.
> >> + * Return: A pointer to the next folio, or %NULL when done.
> >> + */
> >> +static inline struct folio *readahead_folio_reverse(struct readahead_control *ractl)
> >> +{
> >> + struct folio *folio;
> >> +
> >> + if (!ractl->_nr_pages)
> >> + return NULL;
> >> +
> >> + /* xa_load() follows sibling entries, so a tail index returns the head */
> >> + folio = xa_load(&ractl->mapping->i_pages,
> >> + ractl->_index + ractl->_nr_pages - 1);
> >> + VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
> >> +
> >> + /* Shrink the window from the tail down to this folio's head index */
> >> + ractl->_nr_pages = folio->index - ractl->_index;
> >> + ractl->_batch_count = 0;
> >
> > Thanks for the patch! Currently there's the invariant that the returned
> > folio is still inside the _index .. _index+_nr_pages range. I think when we
> > are providing a generic helper, we should keep that to make code more
> > robust for the future when more people start using it.
>
> Definitely.
>
> >
> > What I'd suggest doing is add bool in struct readahead_control telling
> > whether the last folio (batch) was taken from the head or tail of the
> > range, advance _nr_pages and _index accordingly in the functions returning
> > folios (probably hide this in a helper function __readahead_advance()
> > because it will be used in 3 places) and maybe call this new function
> > readahead_folio_last() instead of _reverse() (but I have only a slight
> > preference here so .._reverse() is ok with me if other people prefer it).
>
> The below is what I come up with. I did not add a bool to
> readahead_control, since I think that is the decision of caller of
> __readahead_advance(). But let me know if you disagree.
The reason why I wanted bool in readahead_control is that if some code
ends up mixing readahead_folio() with readahead_folio_last() things will
get confused (because __readahead_advance() really wants to skip the batch
returned from the *previous* call to readahead_folio[_last]()). With the
bool in rac, even mixed use will properly advance the state of the
readahead_control. I don't think mixed use is very realistic (at this
point at least) so I'm ok with leaving that for later if you don't like it.
Also I have some minor comments below.
> diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
> index b59f2745a8e72..23f423c22ac8c 100644
> --- a/fs/erofs/zdata.c
> +++ b/fs/erofs/zdata.c
> @@ -1908,8 +1908,8 @@ static void z_erofs_readahead(struct readahead_control *rac)
> trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
> z_erofs_pcluster_readmore(&f, rac, true);
>
> - /* traverse in reverse order for best metadata I/O performance */
> - while ((folio = readahead_folio_reverse(rac))) {
> + /* traverse from last to first for best metadata I/O performance */
> + while ((folio = readahead_folio_last(rac))) {
> err = z_erofs_scan_folio(&f, folio, true);
> if (err && err != -EINTR)
> erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
> index 5ca5aa365f319..2cc3de5594518 100644
> --- a/include/linux/pagemap.h
> +++ b/include/linux/pagemap.h
> @@ -1510,13 +1510,21 @@ void page_cache_async_readahead(struct address_space *mapping,
> page_cache_async_ra(&ractl, folio, req_count);
> }
>
> +static inline void __readahead_advance(struct readahead_control *rac,
> + bool read_from_head)
> +{
> + if (read_from_head)
> + rac->_index += rac->_batch_count;
> +
> + rac->_nr_pages -= rac->_batch_count;
> +}
Maybe we can add:
rac->_batch_count = 0;
as well since after the advance the _batch_count isn't valid anymore? We
can then also remove it from the callers.
> +
> static inline struct folio *__readahead_folio(struct readahead_control *ractl)
> {
> - struct folio *folio;
> + struct folio *folio = NULL;
Not sure why this initialization got here...
>
> BUG_ON(ractl->_batch_count > ractl->_nr_pages);
> - ractl->_nr_pages -= ractl->_batch_count;
> - ractl->_index += ractl->_batch_count;
> + __readahead_advance(ractl, /* read_from_head= */ true);
^^^^
This is not really a kernel style :), please delete this comment. I'm ok with
pure false/true here - it is an internal helper used in few places. If
things get wider use, we tend to switch to 'unsigned flags' with explicit
flag names to ease code reading. But that's not the case here.
> @@ -1548,7 +1556,7 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
> }
>
> /**
> - * readahead_folio_reverse - Get the next folio to read, from the tail.
> + * readahead_folio_last - Get the next folio to read, from the tail.
> * @ractl: The current readahead request.
> *
> * Like readahead_folio(), but walks the range back-to-front. The folio is
> @@ -1558,21 +1566,24 @@ static inline struct folio *readahead_folio(struct readahead_control *ractl)
> * Context: The folio is locked.
> * Return: A pointer to the next folio, or %NULL when done.
> */
> -static inline struct folio *readahead_folio_reverse(struct readahead_control *ractl)
> +static inline struct folio *readahead_folio_last(struct readahead_control *ractl)
> {
> struct folio *folio;
>
> - if (!ractl->_nr_pages)
> + /* Shrink the window from the tail down to this folio's head index */
> + __readahead_advance(ractl, /* read_from_head= */ false);
> +
> + if (!ractl->_nr_pages) {
> + ractl->_batch_count = 0;
> return NULL;
> + }
>
> /* xa_load() follows sibling entries, so a tail index returns the head */
> folio = xa_load(&ractl->mapping->i_pages,
> ractl->_index + ractl->_nr_pages - 1);
> VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
>
> - /* Shrink the window from the tail down to this folio's head index */
> - ractl->_nr_pages = folio->index - ractl->_index;
> - ractl->_batch_count = 0;
> + ractl->_batch_count = folio_nr_pages(folio);
>
> folio_put(folio);
> return folio;
> @@ -1583,11 +1594,10 @@ static inline unsigned int __readahead_batch(struct readahead_control *rac,
> {
> unsigned int i = 0;
> XA_STATE(xas, &rac->mapping->i_pages, 0);
> - struct folio *folio;
> + struct folio *folio = NULL;
Again not sure why this initialization got here...
> BUG_ON(rac->_batch_count > rac->_nr_pages);
> - rac->_nr_pages -= rac->_batch_count;
> - rac->_index += rac->_batch_count;
> + __readahead_advance(rac, /* read_from_head= */ true);
> rac->_batch_count = 0;
>
> xas_set(&xas, rac->_index);
>
>
>
>
> --
> Best Regards,
> Yan, Zi
>
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
2026-08-04 9:32 ` Jan Kara
@ 2026-08-04 15:54 ` Zi Yan
2026-08-04 17:04 ` Jan Kara
0 siblings, 1 reply; 15+ messages in thread
From: Zi Yan @ 2026-08-04 15:54 UTC (permalink / raw)
To: Jan Kara
Cc: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song, linux-mm,
linux-kernel, Gao Xiang, Chao Yu, Yue Hu, Jeffle Xu,
Sandeep Dhavale, Hongbo Li, Chunhai Guo, linux-erofs,
linux-fsdevel
On Tue Aug 4, 2026 at 5:32 AM EDT, Jan Kara wrote:
> On Mon 03-08-26 12:56:36, Zi Yan wrote:
>> On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
>> > On Fri 31-07-26 22:13:30, Zi Yan wrote:
>> >> erofs needs to traverse readahead folios in reverse order to achieve
>> >> maximum performance by
>> >> 1. reading all folios from readahead_folio();
>> >> 2. storing the prior folio pointer in folio->private;
>> >> 3. traverse from the last folio to the first one.
>> >>
>> >> Add readahead_folio_reverse() to achieve the same function without using
>> >> folio->private.
>> >>
>> >> It prepares for a future commit that replaces PG_private checks with
>> >> !folio->private checks. After switching the checks, erofs's use of
>> >> folio->private without bumping folio refcount can cause unexpected
>> >> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
>> >> reachable.
<snip>
>>
>> The below is what I come up with. I did not add a bool to
>> readahead_control, since I think that is the decision of caller of
>> __readahead_advance(). But let me know if you disagree.
>
> The reason why I wanted bool in readahead_control is that if some code
> ends up mixing readahead_folio() with readahead_folio_last() things will
> get confused (because __readahead_advance() really wants to skip the batch
> returned from the *previous* call to readahead_folio[_last]()). With the
> bool in rac, even mixed use will properly advance the state of the
> readahead_control. I don't think mixed use is very realistic (at this
> point at least) so I'm ok with leaving that for later if you don't like it.
Got it. I am trying to figure out your mental model of how the mix of
readahead_folio() and readahead_folio_last() works with the bool inside
ractl. By looking at readahead_folio_last() code, it is almost the same
as readahead_folio() with __readahead_folio() inlined
(__readahead_folio() is only used by readahead_folio(), so the inline
can happen without any issue). As a result, we can get rid of
readahead_folio_last(), add set_readahead_direction() to set the
embedded bool read_from_head, and use readahead_folio() only. This
removes redundant code in readahead_folio_last(). One thing I am not
certain is whether we want to
1. use set_readahead_direction() explicit and warn readahead_folio() if
read_from_head is not initialized, or
2. set read_from_head to true by default, so that only erofs needs to
call set_readahead_direction() to change read_from_head.
The former is less confusing but changes how readahead_folio() works;
the latter is simpler but implicit read_from_head state might confuse
people at some point.
Let me know your thoughts. Thanks.
>
> Also I have some minor comments below.
>
>> diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
>> index b59f2745a8e72..23f423c22ac8c 100644
>> --- a/fs/erofs/zdata.c
>> +++ b/fs/erofs/zdata.c
>> @@ -1908,8 +1908,8 @@ static void z_erofs_readahead(struct readahead_control *rac)
>> trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
>> z_erofs_pcluster_readmore(&f, rac, true);
>>
>> - /* traverse in reverse order for best metadata I/O performance */
>> - while ((folio = readahead_folio_reverse(rac))) {
>> + /* traverse from last to first for best metadata I/O performance */
>> + while ((folio = readahead_folio_last(rac))) {
>> err = z_erofs_scan_folio(&f, folio, true);
>> if (err && err != -EINTR)
>> erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
>> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
>> index 5ca5aa365f319..2cc3de5594518 100644
>> --- a/include/linux/pagemap.h
>> +++ b/include/linux/pagemap.h
>> @@ -1510,13 +1510,21 @@ void page_cache_async_readahead(struct address_space *mapping,
>> page_cache_async_ra(&ractl, folio, req_count);
>> }
>>
>> +static inline void __readahead_advance(struct readahead_control *rac,
>> + bool read_from_head)
>> +{
>> + if (read_from_head)
>> + rac->_index += rac->_batch_count;
>> +
>> + rac->_nr_pages -= rac->_batch_count;
>> +}
>
> Maybe we can add:
>
> rac->_batch_count = 0;
>
> as well since after the advance the _batch_count isn't valid anymore? We
> can then also remove it from the callers.
Yes, for __readahead_batch(), _batch_count is zeroed right after. For
__readahead_folio() and readahead_folio_last(), _batch_count is either
zeroed or overwritten by folio_nr_pages() before any use.
>
>> +
>> static inline struct folio *__readahead_folio(struct readahead_control *ractl)
>> {
>> - struct folio *folio;
>> + struct folio *folio = NULL;
>
> Not sure why this initialization got here...
Will remove it. It is some leftover during my development. Thank you for
pointing it out.
>
>>
>> BUG_ON(ractl->_batch_count > ractl->_nr_pages);
>> - ractl->_nr_pages -= ractl->_batch_count;
>> - ractl->_index += ractl->_batch_count;
>> + __readahead_advance(ractl, /* read_from_head= */ true);
> ^^^^
> This is not really a kernel style :), please delete this comment. I'm ok with
> pure false/true here - it is an internal helper used in few places. If
> things get wider use, we tend to switch to 'unsigned flags' with explicit
> flag names to ease code reading. But that's not the case here.
We did this in some MM code. I will delete the comment like you
suggested.
<snip>
>> @@ -1583,11 +1594,10 @@ static inline unsigned int __readahead_batch(struct readahead_control *rac,
>> {
>> unsigned int i = 0;
>> XA_STATE(xas, &rac->mapping->i_pages, 0);
>> - struct folio *folio;
>> + struct folio *folio = NULL;
>
> Again not sure why this initialization got here...
Will remove.
--
Best Regards,
Yan, Zi
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
2026-08-04 15:54 ` Zi Yan
@ 2026-08-04 17:04 ` Jan Kara
2026-08-04 17:09 ` Zi Yan
0 siblings, 1 reply; 15+ messages in thread
From: Jan Kara @ 2026-08-04 17:04 UTC (permalink / raw)
To: Zi Yan
Cc: Jan Kara, David Hildenbrand, Matthew Wilcox (Oracle),
Andrew Morton, Muchun Song, Lorenzo Stoakes, Liam R. Howlett,
Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
Baolin Wang, Nico Pache, Ryan Roberts, Dev Jain, Barry Song,
Lance Yang, Usama Arif, Gregory Price, Ying Huang,
Alistair Popple, Johannes Weiner, Qi Zheng, Shakeel Butt,
Kairui Song, linux-mm, linux-kernel, Gao Xiang, Chao Yu, Yue Hu,
Jeffle Xu, Sandeep Dhavale, Hongbo Li, Chunhai Guo, linux-erofs,
linux-fsdevel
On Tue 04-08-26 11:54:41, Zi Yan wrote:
> On Tue Aug 4, 2026 at 5:32 AM EDT, Jan Kara wrote:
> > On Mon 03-08-26 12:56:36, Zi Yan wrote:
> >> On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
> >> > On Fri 31-07-26 22:13:30, Zi Yan wrote:
> >> >> erofs needs to traverse readahead folios in reverse order to achieve
> >> >> maximum performance by
> >> >> 1. reading all folios from readahead_folio();
> >> >> 2. storing the prior folio pointer in folio->private;
> >> >> 3. traverse from the last folio to the first one.
> >> >>
> >> >> Add readahead_folio_reverse() to achieve the same function without using
> >> >> folio->private.
> >> >>
> >> >> It prepares for a future commit that replaces PG_private checks with
> >> >> !folio->private checks. After switching the checks, erofs's use of
> >> >> folio->private without bumping folio refcount can cause unexpected
> >> >> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
> >> >> reachable.
>
> <snip>
>
> >>
> >> The below is what I come up with. I did not add a bool to
> >> readahead_control, since I think that is the decision of caller of
> >> __readahead_advance(). But let me know if you disagree.
> >
> > The reason why I wanted bool in readahead_control is that if some code
> > ends up mixing readahead_folio() with readahead_folio_last() things will
> > get confused (because __readahead_advance() really wants to skip the batch
> > returned from the *previous* call to readahead_folio[_last]()). With the
> > bool in rac, even mixed use will properly advance the state of the
> > readahead_control. I don't think mixed use is very realistic (at this
> > point at least) so I'm ok with leaving that for later if you don't like it.
>
> Got it. I am trying to figure out your mental model of how the mix of
> readahead_folio() and readahead_folio_last() works with the bool inside
> ractl. By looking at readahead_folio_last() code, it is almost the same
> as readahead_folio() with __readahead_folio() inlined
> (__readahead_folio() is only used by readahead_folio(), so the inline
> can happen without any issue). As a result, we can get rid of
> readahead_folio_last(), add set_readahead_direction() to set the
> embedded bool read_from_head, and use readahead_folio() only. This
> removes redundant code in readahead_folio_last(). One thing I am not
> certain is whether we want to
>
> 1. use set_readahead_direction() explicit and warn readahead_folio() if
> read_from_head is not initialized, or
>
> 2. set read_from_head to true by default, so that only erofs needs to
> call set_readahead_direction() to change read_from_head.
>
> The former is less confusing but changes how readahead_folio() works;
> the latter is simpler but implicit read_from_head state might confuse
> people at some point.
My idea was: readahead_folio() will call __readahead_advance() and then set
rac->forward = true. readahead_folio_last() will call __readahead_advance()
and set rac->forward = false. __readahead_advance() advances from beginning
/ end based on rac->_forward value.
Honza
> > Also I have some minor comments below.
> >
> >> diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
> >> index b59f2745a8e72..23f423c22ac8c 100644
> >> --- a/fs/erofs/zdata.c
> >> +++ b/fs/erofs/zdata.c
> >> @@ -1908,8 +1908,8 @@ static void z_erofs_readahead(struct readahead_control *rac)
> >> trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
> >> z_erofs_pcluster_readmore(&f, rac, true);
> >>
> >> - /* traverse in reverse order for best metadata I/O performance */
> >> - while ((folio = readahead_folio_reverse(rac))) {
> >> + /* traverse from last to first for best metadata I/O performance */
> >> + while ((folio = readahead_folio_last(rac))) {
> >> err = z_erofs_scan_folio(&f, folio, true);
> >> if (err && err != -EINTR)
> >> erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
> >> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
> >> index 5ca5aa365f319..2cc3de5594518 100644
> >> --- a/include/linux/pagemap.h
> >> +++ b/include/linux/pagemap.h
> >> @@ -1510,13 +1510,21 @@ void page_cache_async_readahead(struct address_space *mapping,
> >> page_cache_async_ra(&ractl, folio, req_count);
> >> }
> >>
> >> +static inline void __readahead_advance(struct readahead_control *rac,
> >> + bool read_from_head)
> >> +{
> >> + if (read_from_head)
> >> + rac->_index += rac->_batch_count;
> >> +
> >> + rac->_nr_pages -= rac->_batch_count;
> >> +}
> >
> > Maybe we can add:
> >
> > rac->_batch_count = 0;
> >
> > as well since after the advance the _batch_count isn't valid anymore? We
> > can then also remove it from the callers.
>
> Yes, for __readahead_batch(), _batch_count is zeroed right after. For
> __readahead_folio() and readahead_folio_last(), _batch_count is either
> zeroed or overwritten by folio_nr_pages() before any use.
>
> >
> >> +
> >> static inline struct folio *__readahead_folio(struct readahead_control *ractl)
> >> {
> >> - struct folio *folio;
> >> + struct folio *folio = NULL;
> >
> > Not sure why this initialization got here...
>
> Will remove it. It is some leftover during my development. Thank you for
> pointing it out.
>
> >
> >>
> >> BUG_ON(ractl->_batch_count > ractl->_nr_pages);
> >> - ractl->_nr_pages -= ractl->_batch_count;
> >> - ractl->_index += ractl->_batch_count;
> >> + __readahead_advance(ractl, /* read_from_head= */ true);
> > ^^^^
> > This is not really a kernel style :), please delete this comment. I'm ok with
> > pure false/true here - it is an internal helper used in few places. If
> > things get wider use, we tend to switch to 'unsigned flags' with explicit
> > flag names to ease code reading. But that's not the case here.
>
> We did this in some MM code. I will delete the comment like you
> suggested.
>
> <snip>
>
> >> @@ -1583,11 +1594,10 @@ static inline unsigned int __readahead_batch(struct readahead_control *rac,
> >> {
> >> unsigned int i = 0;
> >> XA_STATE(xas, &rac->mapping->i_pages, 0);
> >> - struct folio *folio;
> >> + struct folio *folio = NULL;
> >
> > Again not sure why this initialization got here...
>
> Will remove.
>
> --
> Best Regards,
> Yan, Zi
>
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
2026-08-04 17:04 ` Jan Kara
@ 2026-08-04 17:09 ` Zi Yan
0 siblings, 0 replies; 15+ messages in thread
From: Zi Yan @ 2026-08-04 17:09 UTC (permalink / raw)
To: Jan Kara
Cc: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song, linux-mm,
linux-kernel, Gao Xiang, Chao Yu, Yue Hu, Jeffle Xu,
Sandeep Dhavale, Hongbo Li, Chunhai Guo, linux-erofs,
linux-fsdevel
On Tue Aug 4, 2026 at 1:04 PM EDT, Jan Kara wrote:
> On Tue 04-08-26 11:54:41, Zi Yan wrote:
>> On Tue Aug 4, 2026 at 5:32 AM EDT, Jan Kara wrote:
>> > On Mon 03-08-26 12:56:36, Zi Yan wrote:
>> >> On Mon Aug 3, 2026 at 5:54 AM EDT, Jan Kara wrote:
>> >> > On Fri 31-07-26 22:13:30, Zi Yan wrote:
>> >> >> erofs needs to traverse readahead folios in reverse order to achieve
>> >> >> maximum performance by
>> >> >> 1. reading all folios from readahead_folio();
>> >> >> 2. storing the prior folio pointer in folio->private;
>> >> >> 3. traverse from the last folio to the first one.
>> >> >>
>> >> >> Add readahead_folio_reverse() to achieve the same function without using
>> >> >> folio->private.
>> >> >>
>> >> >> It prepares for a future commit that replaces PG_private checks with
>> >> >> !folio->private checks. After switching the checks, erofs's use of
>> >> >> folio->private without bumping folio refcount can cause unexpected
>> >> >> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
>> >> >> reachable.
>>
>> <snip>
>>
>> >>
>> >> The below is what I come up with. I did not add a bool to
>> >> readahead_control, since I think that is the decision of caller of
>> >> __readahead_advance(). But let me know if you disagree.
>> >
>> > The reason why I wanted bool in readahead_control is that if some code
>> > ends up mixing readahead_folio() with readahead_folio_last() things will
>> > get confused (because __readahead_advance() really wants to skip the batch
>> > returned from the *previous* call to readahead_folio[_last]()). With the
>> > bool in rac, even mixed use will properly advance the state of the
>> > readahead_control. I don't think mixed use is very realistic (at this
>> > point at least) so I'm ok with leaving that for later if you don't like it.
>>
>> Got it. I am trying to figure out your mental model of how the mix of
>> readahead_folio() and readahead_folio_last() works with the bool inside
>> ractl. By looking at readahead_folio_last() code, it is almost the same
>> as readahead_folio() with __readahead_folio() inlined
>> (__readahead_folio() is only used by readahead_folio(), so the inline
>> can happen without any issue). As a result, we can get rid of
>> readahead_folio_last(), add set_readahead_direction() to set the
>> embedded bool read_from_head, and use readahead_folio() only. This
>> removes redundant code in readahead_folio_last(). One thing I am not
>> certain is whether we want to
>>
>> 1. use set_readahead_direction() explicit and warn readahead_folio() if
>> read_from_head is not initialized, or
>>
>> 2. set read_from_head to true by default, so that only erofs needs to
>> call set_readahead_direction() to change read_from_head.
>>
>> The former is less confusing but changes how readahead_folio() works;
>> the latter is simpler but implicit read_from_head state might confuse
>> people at some point.
>
> My idea was: readahead_folio() will call __readahead_advance() and then set
> rac->forward = true. readahead_folio_last() will call __readahead_advance()
> and set rac->forward = false. __readahead_advance() advances from beginning
> / end based on rac->_forward value.
Got it. I can do that. Just to be clear, it should be that
readahead_folio() first sets rac->forward = true, then calls
__readahead_advance(), since __readahead_advance() advances based on
rac->forward, right? readahead_folio_last() as well.
--
Best Regards,
Yan, Zi
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private
2026-08-01 2:13 ` [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private Zi Yan
2026-08-03 9:54 ` Jan Kara
@ 2026-08-03 23:55 ` Gao Xiang
1 sibling, 0 replies; 15+ messages in thread
From: Gao Xiang @ 2026-08-03 23:55 UTC (permalink / raw)
To: Zi Yan
Cc: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song, linux-mm,
linux-kernel, Gao Xiang, Chao Yu, Jan Kara, Yue Hu, Jeffle Xu,
Sandeep Dhavale, Hongbo Li, Chunhai Guo, linux-erofs,
linux-fsdevel
On Fri, Jul 31, 2026 at 10:13:30PM -0400, Zi Yan wrote:
> erofs needs to traverse readahead folios in reverse order to achieve
> maximum performance by
> 1. reading all folios from readahead_folio();
> 2. storing the prior folio pointer in folio->private;
> 3. traverse from the last folio to the first one.
>
> Add readahead_folio_reverse() to achieve the same function without using
> folio->private.
>
> It prepares for a future commit that replaces PG_private checks with
> !folio->private checks. After switching the checks, erofs's use of
> folio->private without bumping folio refcount can cause unexpected
> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
> reachable.
>
> No funtional change intended.
>
> Assisted-by: Claude:claude-opus-4-8
> Assisted-by: Codex:gpt-5
> Signed-off-by: Zi Yan <ziy@nvidia.com>
> To: Gao Xiang <xiang@kernel.org>
> To: Chao Yu <chao@kernel.org>
> To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
> To: Jan Kara <jack@suse.cz>
> Cc: Yue Hu <zbestahu@gmail.com>
> Cc: Jeffle Xu <jefflexu@linux.alibaba.com>
> Cc: Sandeep Dhavale <dhavale@google.com>
> Cc: Hongbo Li <hongbohbli@tencent.com>
> Cc: Chunhai Guo <guochunhai@vivo.com>
> Cc: linux-erofs@lists.ozlabs.org
> Cc: linux-kernel@vger.kernel.org
> Cc: linux-fsdevel@vger.kernel.org
> Cc: linux-mm@kvack.org
> ---
> fs/erofs/zdata.c | 11 ++---------
> include/linux/pagemap.h | 31 +++++++++++++++++++++++++++++++
> 2 files changed, 33 insertions(+), 9 deletions(-)
>
> diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
> index 74520e9102596..b59f2745a8e72 100644
> --- a/fs/erofs/zdata.c
> +++ b/fs/erofs/zdata.c
> @@ -1902,21 +1902,14 @@ static void z_erofs_readahead(struct readahead_control *rac)
> struct inode *realinode = erofs_real_inode(sharedinode, &need_iput);
> Z_EROFS_DEFINE_FRONTEND(f, realinode, sharedinode, readahead_pos(rac));
> unsigned int nrpages = readahead_count(rac);
> - struct folio *head = NULL, *folio;
> + struct folio *folio;
> int err;
>
> trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
> z_erofs_pcluster_readmore(&f, rac, true);
> - while ((folio = readahead_folio(rac))) {
> - folio->private = head;
> - head = folio;
> - }
>
> /* traverse in reverse order for best metadata I/O performance */
> - while (head) {
> - folio = head;
> - head = folio_get_private(folio);
> -
> + while ((folio = readahead_folio_reverse(rac))) {
Yes, it's needed due to EROFS compression metadata design and on-demand
partial decompression, the last extent in the readahead request can be
parsed as a partial extent (means from the starting logical offset of
extents to the necessary offset.). Since there may be many extents
in a single readahead request, so it needs to iterate backwards here;
but the actual compressed data I/Os will be issued forwards.
Previously I tend to avoid touching core-mm so it uses folio->private
but if MM folks can provide a new helper, that would be very helpful
(one more words: all folios are locked in the forward order previously,
so it won't have any deadlock risk).
Thanks,
Gao Xiang
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH RFC 09/14] mm/page-flags: check page/folio->private instead of PG_private
2026-08-01 2:13 [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
2026-08-01 2:13 ` [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private Zi Yan
@ 2026-08-01 2:13 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 10/14] mm/page-flags: introduce folio_test_fs_private() Zi Yan
` (3 subsequent siblings)
5 siblings, 0 replies; 15+ messages in thread
From: Zi Yan @ 2026-08-01 2:13 UTC (permalink / raw)
To: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song
Cc: linux-mm, linux-kernel, Zi Yan, Steven Rostedt, Masami Hiramatsu,
Jan Kara, Mathieu Desnoyers, Matthew Brost, Joshua Hahn,
Rakie Kim, Byungchul Park, Axel Rasmussen, Yuanchu Xie, Wei Xu,
linux-fsdevel, linux-trace-kernel
After the changes of the prior commits, page/folio->private != NULL is now
equivalent to checking PG_private.
Stop checking PG_private on pages and folios and use page/folio->private
instead, except swapcache and hugetlb folios, because the former uses a
field (swp_entry_t swap) overlapping with ->private and the latter sets its
flags in ->private. Exclude swapcache and hugetlb when the code is meant to
check PG_private only.
folio_set/clear_private() and Set/ClearPagePrivate() become no-ops.
PG_private is no longer checked at page free time.
KPF_PRIVATE exposes PG_private to userspace. Change its code logic to check
folio->private != NULL and exclude non-pagecache, swapcache, hugetlb, and
anon folios. One minor semantic change, for orphaned pagecache folios
(mapping == NULL) with fs-private data will no longer have KPF_PRIVATE.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Andrew Morton <akpm@linux-foundation.org>
To: David Hildenbrand <david@kernel.org>
To: Steven Rostedt <rostedt@goodmis.org>
To: Masami Hiramatsu <mhiramat@kernel.org>
To: Lorenzo Stoakes <ljs@kernel.org>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
To: Johannes Weiner <hannes@cmpxchg.org>
Cc: "Liam R. Howlett" <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Rakie Kim <rakie.kim@sk.com>
Cc: Byungchul Park <byungchul@sk.com>
Cc: Gregory Price <gourry@gourry.net>
Cc: Ying Huang <ying.huang@linux.alibaba.com>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: Qi Zheng <qi.zheng@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Kairui Song <kasong@tencent.com>
Cc: Axel Rasmussen <axelrasmussen@google.com>
Cc: Yuanchu Xie <yuanchu@google.com>
Cc: Wei Xu <weixugc@google.com>
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: linux-trace-kernel@vger.kernel.org
---
fs/proc/page.c | 6 +++++-
include/linux/mm.h | 11 ++++++-----
include/linux/page-flags.h | 26 +++++++++++++++++++++-----
include/trace/events/pagemap.h | 4 +++-
mm/huge_memory.c | 4 +++-
mm/migrate.c | 3 ++-
mm/page-writeback.c | 5 ++++-
mm/vmscan.c | 3 ++-
8 files changed, 46 insertions(+), 16 deletions(-)
diff --git a/fs/proc/page.c b/fs/proc/page.c
index 260772b20bd99..abfa6f7d890cc 100644
--- a/fs/proc/page.c
+++ b/fs/proc/page.c
@@ -232,7 +232,11 @@ u64 stable_page_flags(const struct page *page)
u |= kpf_copy_bit(k, KPF_RESERVED, PG_reserved);
u |= kpf_copy_bit(k, KPF_OWNER_2, PG_owner_2);
- u |= kpf_copy_bit(k, KPF_PRIVATE, PG_private);
+ /* preserve the original KPF_PRIVATE semantics by excluding non pagecache folios */
+ if (folio->mapping && !folio_test_anon(folio) &&
+ (folio_get_private(folio) && !folio_test_swapcache(folio) &&
+ !folio_test_hugetlb(folio)))
+ u |= BIT_ULL(KPF_PRIVATE);
u |= kpf_copy_bit(k, KPF_PRIVATE_2, PG_private_2);
u |= kpf_copy_bit(k, KPF_OWNER_PRIVATE, PG_owner_priv_1);
u |= kpf_copy_bit(k, KPF_ARCH, PG_arch_1);
diff --git a/include/linux/mm.h b/include/linux/mm.h
index 7fabe6c66b4b7..ebc035ac26ccc 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -2961,9 +2961,9 @@ static inline bool folio_maybe_mapped_shared(struct folio *folio)
* @folio: the folio
*
* Calculate the expected folio refcount, taking references from the pagecache,
- * swapcache, PG_private and page table mappings into account. Useful in
- * combination with folio_ref_count() to detect unexpected references (e.g.,
- * GUP or other temporary references).
+ * swapcache, private data (folio->private != NULL) and page table mappings into
+ * account. Useful in combination with folio_ref_count() to detect unexpected
+ * references (e.g., GUP or other temporary references).
*
* Does currently not consider references from the LRU cache. If the folio
* was isolated from the LRU (which is the case during migration or split),
@@ -3003,8 +3003,9 @@ static inline int folio_expected_ref_count(const struct folio *folio)
if (!folio_test_anon(folio)) {
/* One reference per page from the pagecache. */
ref_count += !!folio->mapping << order;
- /* One reference from PG_private. */
- ref_count += folio_test_private(folio);
+ /* One reference from filesystem private data. */
+ ref_count += !!folio->private && !folio_test_hugetlb(folio) &&
+ !folio_test_swapcache(folio);
}
/* One reference per page table mapping. */
diff --git a/include/linux/page-flags.h b/include/linux/page-flags.h
index 7a863572adce7..8efc61f967302 100644
--- a/include/linux/page-flags.h
+++ b/include/linux/page-flags.h
@@ -578,7 +578,23 @@ FOLIO_FLAG(swapbacked, FOLIO_HEAD_PAGE)
* for its own purposes.
* - PG_private and PG_private_2 cause release_folio() and co to be invoked
*/
-PAGEFLAG(Private, private, PF_ANY)
+
+static __always_inline bool folio_test_private(const struct folio *folio)
+{
+ return folio->private;
+}
+
+static __always_inline int PagePrivate(const struct page *page)
+{
+ return !!page_private(page);
+}
+
+/* no-ops during transition */
+static __always_inline void folio_set_private(struct folio *folio) { }
+static __always_inline void folio_clear_private(struct folio *folio) { }
+static __always_inline void SetPagePrivate(struct page *page) { }
+static __always_inline void ClearPagePrivate(struct page *page) { }
+
FOLIO_FLAG(private_2, FOLIO_HEAD_PAGE)
/* owner_2 can be set on tail pages for anon memory */
@@ -1170,7 +1186,7 @@ static __always_inline void __ClearPageAnonExclusive(struct page *page)
*/
#define PAGE_FLAGS_CHECK_AT_FREE \
(1UL << PG_lru | 1UL << PG_locked | \
- 1UL << PG_private | 1UL << PG_private_2 | \
+ 1UL << PG_private_2 | \
1UL << PG_writeback | 1UL << PG_reserved | \
1UL << PG_active | \
1UL << PG_unevictable | __PG_MLOCKED | LRU_GEN_MASK)
@@ -1194,8 +1210,6 @@ static __always_inline void __ClearPageAnonExclusive(struct page *page)
(0xffUL /* order */ | 1UL << PG_has_hwpoisoned | \
1UL << PG_large_rmappable | 1UL << PG_partially_mapped)
-#define PAGE_FLAGS_PRIVATE \
- (1UL << PG_private | 1UL << PG_private_2)
/**
* folio_has_private - Determine if folio has private stuff
* @folio: The folio to be checked
@@ -1205,7 +1219,9 @@ static __always_inline void __ClearPageAnonExclusive(struct page *page)
*/
static inline int folio_has_private(const struct folio *folio)
{
- return !!(folio->flags.f & PAGE_FLAGS_PRIVATE);
+ return (!!folio->private && !folio_test_swapcache(folio) &&
+ !folio_test_hugetlb(folio)) ||
+ folio_test_private_2(folio);
}
#undef PF_ANY
diff --git a/include/trace/events/pagemap.h b/include/trace/events/pagemap.h
index 36c3a90f0acca..fb9abec40ec79 100644
--- a/include/trace/events/pagemap.h
+++ b/include/trace/events/pagemap.h
@@ -22,7 +22,9 @@
(folio_test_swapcache(folio) ? PAGEMAP_SWAPCACHE : 0) | \
(folio_test_swapbacked(folio) ? PAGEMAP_SWAPBACKED : 0) | \
(folio_test_mappedtodisk(folio) ? PAGEMAP_MAPPEDDISK : 0) | \
- (folio_test_private(folio) ? PAGEMAP_BUFFERS : 0) \
+ (folio_test_private(folio) && \
+ !folio_test_swapcache(folio) && \
+ !folio_test_hugetlb(folio) ? PAGEMAP_BUFFERS : 0) \
)
TRACE_EVENT(mm_lru_insertion,
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 21c92ee48e469..d21a9b8f40d15 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -4799,7 +4799,9 @@ static int split_huge_pages_pid(int pid, unsigned long vaddr_start,
* will try to drop it before split and then check if the folio
* can be split or not. So skip the check here.
*/
- if (!folio_test_private(folio) &&
+ if (!(folio_test_private(folio) &&
+ !folio_test_swapcache(folio) &&
+ !folio_test_hugetlb(folio)) &&
folio_expected_ref_count(folio) != folio_ref_count(folio))
goto next;
diff --git a/mm/migrate.c b/mm/migrate.c
index b937cbd764808..ad5daef8721cf 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1330,7 +1330,8 @@ static int migrate_folio_unmap(new_folio_t get_new_folio,
* free the metadata, so the page can be freed.
*/
if (!src->mapping) {
- if (folio_test_private(src)) {
+ if (folio_test_private(src) && !folio_test_swapcache(src) &&
+ !folio_test_hugetlb(src)) {
try_to_free_buffers(src);
goto out;
}
diff --git a/mm/page-writeback.c b/mm/page-writeback.c
index 47495be68598f..b44468383f900 100644
--- a/mm/page-writeback.c
+++ b/mm/page-writeback.c
@@ -2706,7 +2706,10 @@ bool filemap_dirty_folio(struct address_space *mapping, struct folio *folio)
if (folio_test_set_dirty(folio))
return false;
- __folio_mark_dirty(folio, mapping, !folio_test_private(folio));
+ __folio_mark_dirty(folio, mapping,
+ !(folio_test_private(folio) &&
+ !folio_test_swapcache(folio) &&
+ !folio_test_hugetlb(folio)));
if (mapping->host) {
/* !PageAnon && !swapper_space */
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 17d2b793cbfc4..ab059a2c6ea35 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -954,7 +954,8 @@ static void folio_check_dirty_writeback(struct folio *folio,
*writeback = folio_test_writeback(folio);
/* Verify dirty/writeback state if the filesystem supports it */
- if (!folio_test_private(folio))
+ if (!(folio_test_private(folio) && !folio_test_swapcache(folio) &&
+ !folio_test_hugetlb(folio)))
return;
mapping = folio_mapping(folio);
--
2.53.0
^ permalink raw reply related [flat|nested] 15+ messages in thread* [PATCH RFC 10/14] mm/page-flags: introduce folio_test_fs_private()
2026-08-01 2:13 [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
2026-08-01 2:13 ` [PATCH RFC 07/14] fs/erofs: mm/pagemap: add readahead_folio_reverse() to avoid folio->private Zi Yan
2026-08-01 2:13 ` [PATCH RFC 09/14] mm/page-flags: check page/folio->private instead of PG_private Zi Yan
@ 2026-08-01 2:13 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 11/14] treewide: remove folio_set/clear_private() Zi Yan
` (2 subsequent siblings)
5 siblings, 0 replies; 15+ messages in thread
From: Zi Yan @ 2026-08-01 2:13 UTC (permalink / raw)
To: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song
Cc: linux-mm, linux-kernel, Zi Yan, Steven Rostedt, Masami Hiramatsu,
Mathieu Desnoyers, Matthew Brost, Joshua Hahn, Rakie Kim,
Byungchul Park, linux-fsdevel, linux-trace-kernel
folio_test_fs_private() wraps folio->private != NULL check and excludes
swapcache and hugetlb folios, since swapcache uses swp_entry_t overlapping
with folio->private and hugetlb sets its own flags in folio->private.
Replace open code with the helper, since core MM does this check
frequently.
No functional change intended.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Andrew Morton <akpm@linux-foundation.org>
To: David Hildenbrand <david@kernel.org>
To: Steven Rostedt <rostedt@goodmis.org>
To: Masami Hiramatsu <mhiramat@kernel.org>
To: Lorenzo Stoakes <ljs@kernel.org>
Cc: "Liam R. Howlett" <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Rakie Kim <rakie.kim@sk.com>
Cc: Byungchul Park <byungchul@sk.com>
Cc: Gregory Price <gourry@gourry.net>
Cc: Ying Huang <ying.huang@linux.alibaba.com>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: linux-trace-kernel@vger.kernel.org
---
fs/proc/page.c | 3 +--
include/linux/mm.h | 3 +--
include/linux/page-flags.h | 21 ++++++++++++++++++---
include/trace/events/pagemap.h | 4 +---
mm/huge_memory.c | 4 +---
mm/migrate.c | 3 +--
6 files changed, 23 insertions(+), 15 deletions(-)
diff --git a/fs/proc/page.c b/fs/proc/page.c
index abfa6f7d890cc..aab0e20b8f9fa 100644
--- a/fs/proc/page.c
+++ b/fs/proc/page.c
@@ -234,8 +234,7 @@ u64 stable_page_flags(const struct page *page)
u |= kpf_copy_bit(k, KPF_OWNER_2, PG_owner_2);
/* preserve the original KPF_PRIVATE semantics by excluding non pagecache folios */
if (folio->mapping && !folio_test_anon(folio) &&
- (folio_get_private(folio) && !folio_test_swapcache(folio) &&
- !folio_test_hugetlb(folio)))
+ folio_test_fs_private(folio))
u |= BIT_ULL(KPF_PRIVATE);
u |= kpf_copy_bit(k, KPF_PRIVATE_2, PG_private_2);
u |= kpf_copy_bit(k, KPF_OWNER_PRIVATE, PG_owner_priv_1);
diff --git a/include/linux/mm.h b/include/linux/mm.h
index ebc035ac26ccc..a93eaf7aae545 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -3004,8 +3004,7 @@ static inline int folio_expected_ref_count(const struct folio *folio)
/* One reference per page from the pagecache. */
ref_count += !!folio->mapping << order;
/* One reference from filesystem private data. */
- ref_count += !!folio->private && !folio_test_hugetlb(folio) &&
- !folio_test_swapcache(folio);
+ ref_count += folio_test_fs_private(folio);
}
/* One reference per page table mapping. */
diff --git a/include/linux/page-flags.h b/include/linux/page-flags.h
index 8efc61f967302..0e3628ea080c4 100644
--- a/include/linux/page-flags.h
+++ b/include/linux/page-flags.h
@@ -1210,6 +1210,23 @@ static __always_inline void __ClearPageAnonExclusive(struct page *page)
(0xffUL /* order */ | 1UL << PG_has_hwpoisoned | \
1UL << PG_large_rmappable | 1UL << PG_partially_mapped)
+/**
+ * folio_test_fs_private - check if the folio has filesystem private data
+ * @folio: The folio to check.
+ *
+ * Use this in code that may encounter swapcache or hugetlb folios but only
+ * wants to detect filesystem private data. Swapcache stores swp_entry_t in
+ * folio->swap, a union with folio->private, and hugetlb stores its own flags
+ * in folio->private; both are excluded.
+ *
+ * Return: true if folio->private is set and the folio is neither swapcache
+ * nor hugetlb.
+ */
+static inline bool folio_test_fs_private(const struct folio *folio)
+{
+ return folio_test_private(folio) && !folio_test_swapcache(folio) &&
+ !folio_test_hugetlb(folio);
+}
/**
* folio_has_private - Determine if folio has private stuff
* @folio: The folio to be checked
@@ -1219,9 +1236,7 @@ static __always_inline void __ClearPageAnonExclusive(struct page *page)
*/
static inline int folio_has_private(const struct folio *folio)
{
- return (!!folio->private && !folio_test_swapcache(folio) &&
- !folio_test_hugetlb(folio)) ||
- folio_test_private_2(folio);
+ return folio_test_fs_private(folio) || folio_test_private_2(folio);
}
#undef PF_ANY
diff --git a/include/trace/events/pagemap.h b/include/trace/events/pagemap.h
index fb9abec40ec79..5425ef7bbae6e 100644
--- a/include/trace/events/pagemap.h
+++ b/include/trace/events/pagemap.h
@@ -22,9 +22,7 @@
(folio_test_swapcache(folio) ? PAGEMAP_SWAPCACHE : 0) | \
(folio_test_swapbacked(folio) ? PAGEMAP_SWAPBACKED : 0) | \
(folio_test_mappedtodisk(folio) ? PAGEMAP_MAPPEDDISK : 0) | \
- (folio_test_private(folio) && \
- !folio_test_swapcache(folio) && \
- !folio_test_hugetlb(folio) ? PAGEMAP_BUFFERS : 0) \
+ (folio_test_fs_private(folio) ? PAGEMAP_BUFFERS : 0) \
)
TRACE_EVENT(mm_lru_insertion,
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index d21a9b8f40d15..7f8e99cdeacb8 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -4799,9 +4799,7 @@ static int split_huge_pages_pid(int pid, unsigned long vaddr_start,
* will try to drop it before split and then check if the folio
* can be split or not. So skip the check here.
*/
- if (!(folio_test_private(folio) &&
- !folio_test_swapcache(folio) &&
- !folio_test_hugetlb(folio)) &&
+ if (!folio_test_fs_private(folio) &&
folio_expected_ref_count(folio) != folio_ref_count(folio))
goto next;
diff --git a/mm/migrate.c b/mm/migrate.c
index ad5daef8721cf..014fcb745d210 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1330,8 +1330,7 @@ static int migrate_folio_unmap(new_folio_t get_new_folio,
* free the metadata, so the page can be freed.
*/
if (!src->mapping) {
- if (folio_test_private(src) && !folio_test_swapcache(src) &&
- !folio_test_hugetlb(src)) {
+ if (folio_test_fs_private(src)) {
try_to_free_buffers(src);
goto out;
}
--
2.53.0
^ permalink raw reply related [flat|nested] 15+ messages in thread* [PATCH RFC 11/14] treewide: remove folio_set/clear_private()
2026-08-01 2:13 [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
` (2 preceding siblings ...)
2026-08-01 2:13 ` [PATCH RFC 10/14] mm/page-flags: introduce folio_test_fs_private() Zi Yan
@ 2026-08-01 2:13 ` Zi Yan
2026-08-01 2:13 ` [PATCH RFC 14/14] mm/page-flags: remove PG_private Zi Yan
2026-08-03 9:07 ` [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Jürgen Groß
5 siblings, 0 replies; 15+ messages in thread
From: Zi Yan @ 2026-08-01 2:13 UTC (permalink / raw)
To: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song
Cc: linux-mm, linux-kernel, Zi Yan, Trond Myklebust, Anna Schumaker,
Jan Kara, Matthew Brost, Joshua Hahn, Rakie Kim, Byungchul Park,
linux-nfs, linux-fsdevel
They are no-ops now. Remove them.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Trond Myklebust <trondmy@kernel.org>
To: Anna Schumaker <anna@kernel.org>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
To: Andrew Morton <akpm@linux-foundation.org>
To: David Hildenbrand <david@kernel.org>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Rakie Kim <rakie.kim@sk.com>
Cc: Byungchul Park <byungchul@sk.com>
Cc: Gregory Price <gourry@gourry.net>
Cc: Ying Huang <ying.huang@linux.alibaba.com>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: linux-nfs@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
---
fs/nfs/write.c | 2 --
include/linux/pagemap.h | 4 +---
mm/migrate.c | 1 -
3 files changed, 1 insertion(+), 6 deletions(-)
diff --git a/fs/nfs/write.c b/fs/nfs/write.c
index d2b03ceaeb4f1..d7e50ecb4edc3 100644
--- a/fs/nfs/write.c
+++ b/fs/nfs/write.c
@@ -717,7 +717,6 @@ static void nfs_inode_add_request(struct nfs_page *req)
nfs_lock_request(req);
spin_lock(&mapping->i_private_lock);
set_bit(PG_MAPPED, &req->wb_flags);
- folio_set_private(folio);
folio->private = req;
spin_unlock(&mapping->i_private_lock);
atomic_long_inc(&nfsi->nrequests);
@@ -744,7 +743,6 @@ static void nfs_inode_remove_request(struct nfs_page *req)
spin_lock(&mapping->i_private_lock);
if (likely(folio)) {
folio->private = NULL;
- folio_clear_private(folio);
clear_bit(PG_MAPPED, &req->wb_head->wb_flags);
}
spin_unlock(&mapping->i_private_lock);
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 90904a4d173b7..5ca5aa365f319 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -594,7 +594,6 @@ static inline void folio_attach_private(struct folio *folio, void *data)
{
folio_get(folio);
folio->private = data;
- folio_set_private(folio);
}
/**
@@ -629,9 +628,8 @@ static inline void *folio_detach_private(struct folio *folio)
{
void *data = folio_get_private(folio);
- if (!folio_test_private(folio))
+ if (!data)
return NULL;
- folio_clear_private(folio);
folio->private = NULL;
folio_put(folio);
diff --git a/mm/migrate.c b/mm/migrate.c
index 014fcb745d210..1d84766f12a18 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -838,7 +838,6 @@ void folio_migrate_flags(struct folio *newfolio, struct folio *folio)
*/
if (folio_test_swapcache(folio))
folio_clear_swapcache(folio);
- folio_clear_private(folio);
/* page->private contains hugetlb specific flags */
if (!folio_test_hugetlb(folio))
--
2.53.0
^ permalink raw reply related [flat|nested] 15+ messages in thread* [PATCH RFC 14/14] mm/page-flags: remove PG_private
2026-08-01 2:13 [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
` (3 preceding siblings ...)
2026-08-01 2:13 ` [PATCH RFC 11/14] treewide: remove folio_set/clear_private() Zi Yan
@ 2026-08-01 2:13 ` Zi Yan
2026-08-03 9:07 ` [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Jürgen Groß
5 siblings, 0 replies; 15+ messages in thread
From: Zi Yan @ 2026-08-01 2:13 UTC (permalink / raw)
To: David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song
Cc: linux-mm, linux-kernel, Zi Yan, Baoquan He, Pasha Tatashin,
Pratyush Yadav, Jonathan Corbet, Jan Kara, Steven Rostedt,
Masami Hiramatsu, Dave Young, Shuah Khan, Mathieu Desnoyers,
kexec, linux-doc, linux-fsdevel, linux-trace-kernel
folio->private != NULL indicates a folio carries private data, replacing
PG_private. All PG_private users are converted. Remove PG_private and
reserve the space as __PG_folio for future use.
Also update files in Documentation. hugetlbfs_reserv.rst is outdated and
left unchanged. It should be rewritten.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Andrew Morton <akpm@linux-foundation.org>
To: Baoquan He <baoquan.he@linux.dev>
To: Mike Rapoport <rppt@kernel.org>
To: Pasha Tatashin <pasha.tatashin@soleen.com>
To: Pratyush Yadav <pratyush@kernel.org>
To: Jonathan Corbet <corbet@lwn.net>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
To: David Hildenbrand <david@kernel.org>
To: Steven Rostedt <rostedt@goodmis.org>
To: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Dave Young <ruirui.yang@linux.dev>
Cc: Shuah Khan <skhan@linuxfoundation.org>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: "Liam R. Howlett" <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: kexec@lists.infradead.org
Cc: linux-doc@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: linux-trace-kernel@vger.kernel.org
---
Documentation/admin-guide/kdump/vmcoreinfo.rst | 2 +-
Documentation/filesystems/vfs.rst | 6 +++---
include/linux/page-flags.h | 19 ++-----------------
include/trace/events/mmflags.h | 2 +-
kernel/vmcore_info.c | 1 -
5 files changed, 7 insertions(+), 23 deletions(-)
diff --git a/Documentation/admin-guide/kdump/vmcoreinfo.rst b/Documentation/admin-guide/kdump/vmcoreinfo.rst
index 7663c610fe901..5f1df6d080508 100644
--- a/Documentation/admin-guide/kdump/vmcoreinfo.rst
+++ b/Documentation/admin-guide/kdump/vmcoreinfo.rst
@@ -325,7 +325,7 @@ NR_FREE_PAGES
On linux-2.6.21 or later, the number of free pages is in
vm_stat[NR_FREE_PAGES]. Used to get the number of free pages.
-PG_lru|PG_private|PG_swapcache|PG_swapbacked|PG_hwpoison|PG_head_mask
+PG_lru|PG_swapcache|PG_swapbacked|PG_hwpoison|PG_head_mask
--------------------------------------------------------------------------
Page attributes. These flags are used to filter various unnecessary for
diff --git a/Documentation/filesystems/vfs.rst b/Documentation/filesystems/vfs.rst
index e7677423a20f7..5cc07f82fe371 100644
--- a/Documentation/filesystems/vfs.rst
+++ b/Documentation/filesystems/vfs.rst
@@ -649,8 +649,8 @@ Writeback.
The first can be used independently to the others. The VM can try to
release clean pages in order to reuse them. To do this it can call
-->release_folio on clean folios with the private
-flag set. Clean pages without PagePrivate and with no external references
+->release_folio on clean folios with folio->private set. Clean pages
+without folio->private set and with no external references
will be released without notice being given to the address_space.
To achieve this functionality, pages need to be placed on an LRU with
@@ -674,7 +674,7 @@ filemap_fdatawait_range, to wait for all writeback to complete.
An address_space handler may attach extra information to a page,
typically using the 'private' field in the 'struct page'. If such
-information is attached, the PG_Private flag should be set. This will
+information is attached, non-NULL 'private' field will
cause various VM routines to make extra calls into the address_space
handler to deal with that data.
diff --git a/include/linux/page-flags.h b/include/linux/page-flags.h
index 0e3628ea080c4..eb2961ed61018 100644
--- a/include/linux/page-flags.h
+++ b/include/linux/page-flags.h
@@ -44,10 +44,6 @@
* Consequently, PG_reserved for a page mapped into user space can indicate
* the zero page, the vDSO, MMIO pages or device memory.
*
- * The PG_private bitflag is set on pagecache pages if they contain filesystem
- * specific data (which is normally at page->private). It can be used by
- * private allocations for its own usage.
- *
* During initiation of disk I/O, PG_locked is set. This bit is set before I/O
* and cleared when writeback _starts_ or when read _completes_. PG_writeback
* is set before writeback starts and cleared when it finishes.
@@ -105,7 +101,7 @@ enum pageflags {
PG_owner_2, /* Owner use. If pagecache, fs may use */
PG_arch_1,
PG_reserved,
- PG_private, /* If pagecache, has fs-private data */
+ __PG_folio, /* Do not use: reserved for folio identification */
PG_private_2, /* If pagecache, has fs aux data */
PG_reclaim, /* To be reclaimed asap */
PG_swapbacked, /* Page is backed by RAM/swap */
@@ -576,7 +572,7 @@ FOLIO_FLAG(swapbacked, FOLIO_HEAD_PAGE)
/*
* Private page markings that may be used by the filesystem that owns the page
* for its own purposes.
- * - PG_private and PG_private_2 cause release_folio() and co to be invoked
+ * - folio->private and PG_private_2 cause release_folio() and co to be invoked
*/
static __always_inline bool folio_test_private(const struct folio *folio)
@@ -584,17 +580,6 @@ static __always_inline bool folio_test_private(const struct folio *folio)
return folio->private;
}
-static __always_inline int PagePrivate(const struct page *page)
-{
- return !!page_private(page);
-}
-
-/* no-ops during transition */
-static __always_inline void folio_set_private(struct folio *folio) { }
-static __always_inline void folio_clear_private(struct folio *folio) { }
-static __always_inline void SetPagePrivate(struct page *page) { }
-static __always_inline void ClearPagePrivate(struct page *page) { }
-
FOLIO_FLAG(private_2, FOLIO_HEAD_PAGE)
/* owner_2 can be set on tail pages for anon memory */
diff --git a/include/trace/events/mmflags.h b/include/trace/events/mmflags.h
index 935893e5ea53b..caf090cd6f85e 100644
--- a/include/trace/events/mmflags.h
+++ b/include/trace/events/mmflags.h
@@ -144,7 +144,7 @@ TRACE_DEFINE_ENUM(___GFP_LAST_BIT);
DEF_PAGEFLAG_NAME(owner_2), \
DEF_PAGEFLAG_NAME(arch_1), \
DEF_PAGEFLAG_NAME(reserved), \
- DEF_PAGEFLAG_NAME(private), \
+ { 1UL << __PG_folio, "folio" }, \
DEF_PAGEFLAG_NAME(private_2), \
DEF_PAGEFLAG_NAME(writeback), \
DEF_PAGEFLAG_NAME(head), \
diff --git a/kernel/vmcore_info.c b/kernel/vmcore_info.c
index 8614430ca212a..5a417f8a922ab 100644
--- a/kernel/vmcore_info.c
+++ b/kernel/vmcore_info.c
@@ -216,7 +216,6 @@ static int __init crash_save_vmcoreinfo_init(void)
VMCOREINFO_LENGTH(free_area.free_list, MIGRATE_TYPES);
VMCOREINFO_NUMBER(NR_FREE_PAGES);
VMCOREINFO_NUMBER(PG_lru);
- VMCOREINFO_NUMBER(PG_private);
VMCOREINFO_NUMBER(PG_swapcache);
VMCOREINFO_NUMBER(PG_swapbacked);
#define PAGE_SLAB_MAPCOUNT_VALUE (PGTY_slab << 24)
--
2.53.0
^ permalink raw reply related [flat|nested] 15+ messages in thread* Re: [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead
2026-08-01 2:13 [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
` (4 preceding siblings ...)
2026-08-01 2:13 ` [PATCH RFC 14/14] mm/page-flags: remove PG_private Zi Yan
@ 2026-08-03 9:07 ` Jürgen Groß
2026-08-03 18:13 ` Zi Yan
5 siblings, 1 reply; 15+ messages in thread
From: Jürgen Groß @ 2026-08-03 9:07 UTC (permalink / raw)
To: Zi Yan, David Hildenbrand, Matthew Wilcox (Oracle), Andrew Morton,
Muchun Song, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Baolin Wang,
Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
Usama Arif, Gregory Price, Ying Huang, Alistair Popple,
Johannes Weiner, Qi Zheng, Shakeel Butt, Kairui Song
Cc: linux-mm, linux-kernel, Minchan Kim, Sergey Senozhatsky,
Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers,
Adrian Hunter, James Clark, H. Peter Anvin, linux-perf-users,
Stefano Stabellini, Oleksandr Tyshchenko, xen-devel, Eric Biggers,
Theodore Y. Ts'o, Jaegeuk Kim, linux-fscrypt, Oscar Salvador,
Chao Yu, linux-f2fs-devel, Gao Xiang, Jan Kara, Yue Hu, Jeffle Xu,
Sandeep Dhavale, Hongbo Li, Chunhai Guo, linux-erofs,
linux-fsdevel, Steven Rostedt, Masami Hiramatsu,
Mathieu Desnoyers, Matthew Brost, Joshua Hahn, Rakie Kim,
Byungchul Park, Axel Rasmussen, Yuanchu Xie, Wei Xu,
linux-trace-kernel, Trond Myklebust, Anna Schumaker, linux-nfs,
Song Liu, Yu Kuai, Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Li Nan, Xiao Ni, linux-raid, ceph-devel, Richard Weinberger,
Zhihao Cheng, linux-mtd, Baoquan He, Pasha Tatashin,
Pratyush Yadav, Jonathan Corbet, Dave Young, Shuah Khan, kexec,
linux-doc
[-- Attachment #1.1.1: Type: text/plain, Size: 637 bytes --]
On 01.08.26 04:13, Zi Yan wrote:
> Hi all,
>
> This patchset removes PG_private to make space for upcoming PG_folio
> (reserved as __PG_folio) for identifying pages from a folio (more details
> in Note below). Instead of checking PG_private, all code is changed to
> check page/folio->private != NULL instead.
I'm a little bit worried that page/folio->private is in a union, so today
it could (in theory) be != NULL while PG_private isn't set.
Is it really not possible to enter a path where PG_private is tested while
page/folio->private != NULL due to the union being used otherwise (PG_private
not set)?
Juergen
[-- Attachment #1.1.2: OpenPGP public key --]
[-- Type: application/pgp-keys, Size: 3743 bytes --]
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 495 bytes --]
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead
2026-08-03 9:07 ` [PATCH RFC 00/14] Remove PG_private by using page/folio->private checks instead Jürgen Groß
@ 2026-08-03 18:13 ` Zi Yan
0 siblings, 0 replies; 15+ messages in thread
From: Zi Yan @ 2026-08-03 18:13 UTC (permalink / raw)
To: Jürgen Groß, David Hildenbrand, Matthew Wilcox (Oracle),
Andrew Morton, Muchun Song, Lorenzo Stoakes, Liam R. Howlett,
Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
Baolin Wang, Nico Pache, Ryan Roberts, Dev Jain, Barry Song,
Lance Yang, Usama Arif, Gregory Price, Ying Huang,
Alistair Popple, Johannes Weiner, Qi Zheng, Shakeel Butt,
Kairui Song
Cc: linux-mm, linux-kernel, Minchan Kim, Sergey Senozhatsky,
Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers,
Adrian Hunter, James Clark, H. Peter Anvin, linux-perf-users,
Stefano Stabellini, Oleksandr Tyshchenko, xen-devel, Eric Biggers,
Theodore Y. Ts'o, Jaegeuk Kim, linux-fscrypt, Oscar Salvador,
Chao Yu, linux-f2fs-devel, Gao Xiang, Jan Kara, Yue Hu, Jeffle Xu,
Sandeep Dhavale, Hongbo Li, Chunhai Guo, linux-erofs,
linux-fsdevel, Steven Rostedt, Masami Hiramatsu,
Mathieu Desnoyers, Matthew Brost, Joshua Hahn, Rakie Kim,
Byungchul Park, Axel Rasmussen, Yuanchu Xie, Wei Xu,
linux-trace-kernel, Trond Myklebust, Anna Schumaker, linux-nfs,
Song Liu, Yu Kuai, Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Li Nan, Xiao Ni, linux-raid, ceph-devel, Richard Weinberger,
Zhihao Cheng, linux-mtd, Baoquan He, Pasha Tatashin,
Pratyush Yadav, Jonathan Corbet, Dave Young, Shuah Khan, kexec,
linux-doc
On Mon Aug 3, 2026 at 5:07 AM EDT, Jürgen Groß wrote:
> On 01.08.26 04:13, Zi Yan wrote:
>> Hi all,
>>
>> This patchset removes PG_private to make space for upcoming PG_folio
>> (reserved as __PG_folio) for identifying pages from a folio (more details
>> in Note below). Instead of checking PG_private, all code is changed to
>> check page/folio->private != NULL instead.
>
> I'm a little bit worried that page/folio->private is in a union, so today
> it could (in theory) be != NULL while PG_private isn't set.
>
> Is it really not possible to enter a path where PG_private is tested while
> page/folio->private != NULL due to the union being used otherwise (PG_private
> not set)?
Yes, it is possible. See: #5 in the exceptional users: erofs uses
->private for reverse linked list and in-flight counters without setting
PG_private. I get rid of the first one and converted the second one to
use folio_attach/detach/get_private() to follow the general ->private
use pattern..
For non file system folios, anon swapcache puts swap_entry_t in
->private and hugetlb puts its flags in ->private. I added
folio_test_fs_private() to exclude them, but this helper is planned to
be used by core MM, since filesystem code should not encounter these
two.
--
Best Regards,
Yan, Zi
^ permalink raw reply [flat|nested] 15+ messages in thread