* [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file()
@ 2026-08-11 13:36 Vernon Yang
2026-08-11 15:19 ` David Hildenbrand (Arm)
2026-08-11 19:12 ` Andrew Morton
0 siblings, 2 replies; 5+ messages in thread
From: Vernon Yang @ 2026-08-11 13:36 UTC (permalink / raw)
To: akpm, david, ljs
Cc: nico.pache, ryan.roberts, dev.jain, baohua, lance.yang,
usama.arif, zokeefe, linux-kernel, linux-mm, Vernon Yang, stable
From: Vernon Yang <yanglincheng@kylinos.cn>
When the swap entries found exceed max_ptes_swap, the loop is left via
break with folio still holding the xarray value that encodes the swap
entry, not valid folio pointer.
That value is passed to trace_mm_khugepaged_scan_file(), which feeds it
to folio_pfn(). On FLATMEM and SPARSEMEM_VMEMMAP, the page_to_pfn() is
plain pointer arithmetic, so the trace event merely prints bogus
scan_pfn. On classic SPARSEMEM, the page_to_pfn() reads page->flags,
dereferencing the tiny encoded integer and oopsing khugepaged whenever
the trace event is enabled.
So set folio to NULL before breaking out, the tracepoint maps NULL to
scan_pfn of -1, just like exhausted scan naturally.
Fixes: d41fd2016ed0 ("mm/khugepaged: add tracepoint to hpage_collapse_scan_file()")
Cc: stable@vger.kernel.org
Signed-off-by: Vernon Yang <yanglincheng@kylinos.cn>
---
mm/khugepaged.c | 1 +
1 file changed, 1 insertion(+)
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 617bca76db49..bc0d04c9162d 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -2696,6 +2696,7 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
if (xa_is_value(folio)) {
swap += 1 << xas_get_order(&xas);
if (swap > max_ptes_swap) {
+ folio = NULL;
result = SCAN_EXCEED_SWAP_PTE;
count_vm_event(THP_SCAN_EXCEED_SWAP_PTE);
break;
base-commit: 075b74841bd0065a3bda3440873c747938e69b68
--
2.53.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file()
2026-08-11 13:36 [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file() Vernon Yang
@ 2026-08-11 15:19 ` David Hildenbrand (Arm)
2026-08-12 14:03 ` Vernon Yang
2026-08-11 19:12 ` Andrew Morton
1 sibling, 1 reply; 5+ messages in thread
From: David Hildenbrand (Arm) @ 2026-08-11 15:19 UTC (permalink / raw)
To: Vernon Yang, akpm, ljs
Cc: nico.pache, ryan.roberts, dev.jain, baohua, lance.yang,
usama.arif, zokeefe, linux-kernel, linux-mm, Vernon Yang, stable
On 8/11/26 15:36, Vernon Yang wrote:
> From: Vernon Yang <yanglincheng@kylinos.cn>
>
> When the swap entries found exceed max_ptes_swap, the loop is left via
> break with folio still holding the xarray value that encodes the swap
> entry, not valid folio pointer.
>
> That value is passed to trace_mm_khugepaged_scan_file(), which feeds it
> to folio_pfn(). On FLATMEM and SPARSEMEM_VMEMMAP, the page_to_pfn() is
> plain pointer arithmetic, so the trace event merely prints bogus
> scan_pfn. On classic SPARSEMEM, the page_to_pfn() reads page->flags,
> dereferencing the tiny encoded integer and oopsing khugepaged whenever
> the trace event is enabled.
>
> So set folio to NULL before breaking out, the tracepoint maps NULL to
> scan_pfn of -1, just like exhausted scan naturally.
>
> Fixes: d41fd2016ed0 ("mm/khugepaged: add tracepoint to hpage_collapse_scan_file()")
> Cc: stable@vger.kernel.org
> Signed-off-by: Vernon Yang <yanglincheng@kylinos.cn>
> ---
> mm/khugepaged.c | 1 +
> 1 file changed, 1 insertion(+)
>
> diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> index 617bca76db49..bc0d04c9162d 100644
> --- a/mm/khugepaged.c
> +++ b/mm/khugepaged.c
> @@ -2696,6 +2696,7 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
> if (xa_is_value(folio)) {
> swap += 1 << xas_get_order(&xas);
> if (swap > max_ptes_swap) {
> + folio = NULL;
> result = SCAN_EXCEED_SWAP_PTE;
> count_vm_event(THP_SCAN_EXCEED_SWAP_PTE);
> break;
Yes, we'll do a folio_pfn(), and used to do a page_to_pfn().
Using the folio after dropping the reference is rather nasty.
Instead of passing the folio, should we just pass the pfn directly?
--
Cheers,
David
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file()
2026-08-11 13:36 [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file() Vernon Yang
2026-08-11 15:19 ` David Hildenbrand (Arm)
@ 2026-08-11 19:12 ` Andrew Morton
2026-08-12 14:08 ` Vernon Yang
1 sibling, 1 reply; 5+ messages in thread
From: Andrew Morton @ 2026-08-11 19:12 UTC (permalink / raw)
To: Vernon Yang
Cc: david, ljs, nico.pache, ryan.roberts, dev.jain, baohua,
lance.yang, usama.arif, zokeefe, linux-kernel, linux-mm,
Vernon Yang, stable
On Tue, 11 Aug 2026 21:36:55 +0800 Vernon Yang <vernon2gm@gmail.com> wrote:
> When the swap entries found exceed max_ptes_swap, the loop is left via
> break with folio still holding the xarray value that encodes the swap
> entry, not valid folio pointer.
>
> That value is passed to trace_mm_khugepaged_scan_file(), which feeds it
> to folio_pfn(). On FLATMEM and SPARSEMEM_VMEMMAP, the page_to_pfn() is
> plain pointer arithmetic, so the trace event merely prints bogus
> scan_pfn. On classic SPARSEMEM, the page_to_pfn() reads page->flags,
> dereferencing the tiny encoded integer and oopsing khugepaged whenever
> the trace event is enabled.
>
> So set folio to NULL before breaking out, the tracepoint maps NULL to
> scan_pfn of -1, just like exhausted scan naturally.
>
> Fixes: d41fd2016ed0 ("mm/khugepaged: add tracepoint to hpage_collapse_scan_file()")
Added in 2022. Why so long - do people not use tracing?
Sashiko might have a found a couple of other tracing bugs in this code,
which I suggest are on-topic for your patch:
https://sashiko.dev/#/patchset/20260811133655.267739-1-vernon2gm@gmail.com
Also a possible bug mapping large folios which straddle i_size, which
is a separate thing.
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file()
2026-08-11 15:19 ` David Hildenbrand (Arm)
@ 2026-08-12 14:03 ` Vernon Yang
0 siblings, 0 replies; 5+ messages in thread
From: Vernon Yang @ 2026-08-12 14:03 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: akpm, ljs, nico.pache, ryan.roberts, dev.jain, baohua, lance.yang,
usama.arif, zokeefe, linux-kernel, linux-mm, Vernon Yang, stable
On Tue, Aug 11, 2026 at 05:19:38PM +0200, David Hildenbrand (Arm) wrote:
> On 8/11/26 15:36, Vernon Yang wrote:
> > From: Vernon Yang <yanglincheng@kylinos.cn>
> >
> > When the swap entries found exceed max_ptes_swap, the loop is left via
> > break with folio still holding the xarray value that encodes the swap
> > entry, not valid folio pointer.
> >
> > That value is passed to trace_mm_khugepaged_scan_file(), which feeds it
> > to folio_pfn(). On FLATMEM and SPARSEMEM_VMEMMAP, the page_to_pfn() is
> > plain pointer arithmetic, so the trace event merely prints bogus
> > scan_pfn. On classic SPARSEMEM, the page_to_pfn() reads page->flags,
> > dereferencing the tiny encoded integer and oopsing khugepaged whenever
> > the trace event is enabled.
> >
> > So set folio to NULL before breaking out, the tracepoint maps NULL to
> > scan_pfn of -1, just like exhausted scan naturally.
> >
> > Fixes: d41fd2016ed0 ("mm/khugepaged: add tracepoint to hpage_collapse_scan_file()")
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Vernon Yang <yanglincheng@kylinos.cn>
> > ---
> > mm/khugepaged.c | 1 +
> > 1 file changed, 1 insertion(+)
> >
> > diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> > index 617bca76db49..bc0d04c9162d 100644
> > --- a/mm/khugepaged.c
> > +++ b/mm/khugepaged.c
> > @@ -2696,6 +2696,7 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
> > if (xa_is_value(folio)) {
> > swap += 1 << xas_get_order(&xas);
> > if (swap > max_ptes_swap) {
> > + folio = NULL;
> > result = SCAN_EXCEED_SWAP_PTE;
> > count_vm_event(THP_SCAN_EXCEED_SWAP_PTE);
> > break;
>
> Yes, we'll do a folio_pfn(), and used to do a page_to_pfn().
>
> Using the folio after dropping the reference is rather nasty.
>
> Instead of passing the folio, should we just pass the pfn directly?
Yes, LGTM.
Would similar modifications like the following match the effect you want?
If so, I'll make these changes in the next version.
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 617bca76db49..e7830761d3a2 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -2683,6 +2683,7 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
int present, swap;
int node = NUMA_NO_NODE;
enum scan_result result = SCAN_SUCCEED;
+ unsigned long pfn;
present = 0;
swap = 0;
@@ -2720,27 +2721,23 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
* PMD-sized THP implies that we can only try
* retracting the PTE table.
*/
- folio_put(folio);
break;
}
node = folio_nid(folio);
if (collapse_scan_abort(node, cc)) {
result = SCAN_SCAN_ABORT;
- folio_put(folio);
break;
}
cc->node_load[node]++;
if (!folio_test_lru(folio)) {
result = SCAN_PAGE_LRU;
- folio_put(folio);
break;
}
if (folio_expected_ref_count(folio) + 1 != folio_ref_count(folio)) {
result = SCAN_PAGE_COUNT;
- folio_put(folio);
break;
}
@@ -2759,7 +2756,14 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
cond_resched_rcu();
}
}
+ if (!folio || xa_is_value(folio)) {
+ pfn = -1;
+ } else {
+ pfn = folio_pfn(folio);
+ folio_put(folio);
+ }
rcu_read_unlock();
+
if (result == SCAN_PTE_MAPPED_HUGEPAGE)
cc->progress++;
else
@@ -2774,7 +2778,7 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
}
}
- trace_mm_khugepaged_scan_file(mm, folio, file, present, swap, result);
+ trace_mm_khugepaged_scan_file(mm, pfn, file, present, swap, result);
return result;
}
--
Cheers,
Vernon
^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file()
2026-08-11 19:12 ` Andrew Morton
@ 2026-08-12 14:08 ` Vernon Yang
0 siblings, 0 replies; 5+ messages in thread
From: Vernon Yang @ 2026-08-12 14:08 UTC (permalink / raw)
To: Andrew Morton
Cc: david, ljs, nico.pache, ryan.roberts, dev.jain, baohua,
lance.yang, usama.arif, zokeefe, linux-kernel, linux-mm,
Vernon Yang, stable
On Tue, Aug 11, 2026 at 12:12:25PM -0700, Andrew Morton wrote:
> On Tue, 11 Aug 2026 21:36:55 +0800 Vernon Yang <vernon2gm@gmail.com> wrote:
>
> > When the swap entries found exceed max_ptes_swap, the loop is left via
> > break with folio still holding the xarray value that encodes the swap
> > entry, not valid folio pointer.
> >
> > That value is passed to trace_mm_khugepaged_scan_file(), which feeds it
> > to folio_pfn(). On FLATMEM and SPARSEMEM_VMEMMAP, the page_to_pfn() is
> > plain pointer arithmetic, so the trace event merely prints bogus
> > scan_pfn. On classic SPARSEMEM, the page_to_pfn() reads page->flags,
> > dereferencing the tiny encoded integer and oopsing khugepaged whenever
> > the trace event is enabled.
> >
> > So set folio to NULL before breaking out, the tracepoint maps NULL to
> > scan_pfn of -1, just like exhausted scan naturally.
> >
> > Fixes: d41fd2016ed0 ("mm/khugepaged: add tracepoint to hpage_collapse_scan_file()")
>
> Added in 2022. Why so long - do people not use tracing?
For most architectures, SPARSEMEM_VMEMMAP is the default, where
page_to_pfn() is plain pointer arithmetic, merely printing bogus
scan_pfn. Only on classic SPARSEMEM, the swap entry value cause
khugepaged to oops.
> Sashiko might have a found a couple of other tracing bugs in this code,
> which I suggest are on-topic for your patch:
>
> https://sashiko.dev/#/patchset/20260811133655.267739-1-vernon2gm@gmail.com
>
David also mentioned "Using the folio after dropping the reference",
and the same issue exists in
trace_mm_khugepaged_collapse_file()/trace_mm_khugepaged_scan_pmd(),
which merely prints a overdue scan_pfn without oopsing khugepaged.
However, this approach does have potential issues, just that they
haven't been triggered yet. I can fix them together.
I will add PATCH#2 and PATCH#3 in the next version of the patchset to
fix the other two trace_xxx() functions, rather than mixing them
together into one large patch.
> Also a possible bug mapping large folios which straddle i_size, which
> is a separate thing.
Yes, I will submit a separate fix patch later to resolve this issue.
--
Cheers,
Vernon
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-08-12 14:09 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-11 13:36 [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file() Vernon Yang
2026-08-11 15:19 ` David Hildenbrand (Arm)
2026-08-12 14:03 ` Vernon Yang
2026-08-11 19:12 ` Andrew Morton
2026-08-12 14:08 ` Vernon Yang
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.