* Re: [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio
2026-10-07 4:10 [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio Kyle Zeng
@ 2026-10-07 10:21 ` David Hildenbrand (Arm)
2026-10-07 11:14 ` Zi Yan
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-07 10:21 UTC (permalink / raw)
To: Kyle Zeng, linux-mm
Cc: linux-kernel, Andrew Morton, Zi Yan, Baolin Wang,
outbounddisclosures, stable
On 10/7/26 06:10, Kyle Zeng wrote:
> collapse_file() can fail its reference-count or dirty-folio check after
> unmapping with TTU_BATCH_FLUSH. Both paths put back the isolated folio,
> then unlock it and drop the lookup reference before reaching the common
> try_to_unmap_flush().
>
> The page-cache reference does not keep the folio stable once the lock is
> released. Another collapse can replace and free it. With its PTEs
> already gone, that collapse cannot flush the first task's per-task TLB
> batch, and retract_page_tables() skips short or unaligned VMAs. A CPU
> can therefore retain a user translation to the freed folio. This has
> been reproduced with unprivileged MADV_COLLAPSE on a memfd.
>
> Flush at out_unlock while the lookup reference and folio lock are still
> held. The common flush continues to cover the accumulated pagelist on
> both success and rollback, preserving batching on successful collapses.
>
> Fixes: 6d9df8a5889c ("mm/thp: collapse_file() do try_to_unmap(TTU_BATCH_FLUSH)")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Kyle Zeng <kylebot@openai.com>
> ---
LGTM
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio
2026-10-07 4:10 [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio Kyle Zeng
2026-10-07 10:21 ` David Hildenbrand (Arm)
@ 2026-10-07 11:14 ` Zi Yan
2026-10-08 6:20 ` Baolin Wang
2026-10-09 6:16 ` Hugh Dickins
3 siblings, 0 replies; 5+ messages in thread
From: Zi Yan @ 2026-10-07 11:14 UTC (permalink / raw)
To: Kyle Zeng, linux-mm
Cc: linux-kernel, Andrew Morton, David Hildenbrand, Baolin Wang,
outbounddisclosures, stable
On Wed Oct 7, 2026 at 12:10 AM EDT, Kyle Zeng wrote:
> collapse_file() can fail its reference-count or dirty-folio check after
> unmapping with TTU_BATCH_FLUSH. Both paths put back the isolated folio,
> then unlock it and drop the lookup reference before reaching the common
> try_to_unmap_flush().
>
> The page-cache reference does not keep the folio stable once the lock is
> released. Another collapse can replace and free it. With its PTEs
> already gone, that collapse cannot flush the first task's per-task TLB
> batch, and retract_page_tables() skips short or unaligned VMAs. A CPU
> can therefore retain a user translation to the freed folio. This has
> been reproduced with unprivileged MADV_COLLAPSE on a memfd.
>
> Flush at out_unlock while the lookup reference and folio lock are still
> held. The common flush continues to cover the accumulated pagelist on
> both success and rollback, preserving batching on successful collapses.
>
> Fixes: 6d9df8a5889c ("mm/thp: collapse_file() do try_to_unmap(TTU_BATCH_FLUSH)")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Kyle Zeng <kylebot@openai.com>
> ---
> Changes in v2:
> - Use Assisted-by: LLM.
> - Explain that the common flush is a no-op after out_unlock flushes.
>
> mm/khugepaged.c | 10 +++++++---
> 1 file changed, 7 insertions(+), 3 deletions(-)
>
LGTM. Thanks.
Reviewed-by: Zi Yan <ziy@nvidia.com>
--
Best Regards,
Yan, Zi
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio
2026-10-07 4:10 [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio Kyle Zeng
2026-10-07 10:21 ` David Hildenbrand (Arm)
2026-10-07 11:14 ` Zi Yan
@ 2026-10-08 6:20 ` Baolin Wang
2026-10-09 6:16 ` Hugh Dickins
3 siblings, 0 replies; 5+ messages in thread
From: Baolin Wang @ 2026-10-08 6:20 UTC (permalink / raw)
To: Kyle Zeng, linux-mm
Cc: linux-kernel, Andrew Morton, David Hildenbrand, Zi Yan,
outbounddisclosures, stable
On 10/7/26 12:10 PM, Kyle Zeng wrote:
> collapse_file() can fail its reference-count or dirty-folio check after
> unmapping with TTU_BATCH_FLUSH. Both paths put back the isolated folio,
> then unlock it and drop the lookup reference before reaching the common
> try_to_unmap_flush().
>
> The page-cache reference does not keep the folio stable once the lock is
> released. Another collapse can replace and free it. With its PTEs
> already gone, that collapse cannot flush the first task's per-task TLB
> batch, and retract_page_tables() skips short or unaligned VMAs. A CPU
> can therefore retain a user translation to the freed folio. This has
> been reproduced with unprivileged MADV_COLLAPSE on a memfd.
>
> Flush at out_unlock while the lookup reference and folio lock are still
> held. The common flush continues to cover the accumulated pagelist on
> both success and rollback, preserving batching on successful collapses.
>
> Fixes: 6d9df8a5889c ("mm/thp: collapse_file() do try_to_unmap(TTU_BATCH_FLUSH)")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Kyle Zeng <kylebot@openai.com>
> ---
Good catch. LGTM.
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio
2026-10-07 4:10 [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio Kyle Zeng
` (2 preceding siblings ...)
2026-10-08 6:20 ` Baolin Wang
@ 2026-10-09 6:16 ` Hugh Dickins
3 siblings, 0 replies; 5+ messages in thread
From: Hugh Dickins @ 2026-10-09 6:16 UTC (permalink / raw)
To: Kyle Zeng
Cc: linux-mm, linux-kernel, Andrew Morton, David Hildenbrand, Zi Yan,
Baolin Wang, outbounddisclosures, stable
On Tue, 6 Oct 2026, Kyle Zeng wrote:
> collapse_file() can fail its reference-count or dirty-folio check after
> unmapping with TTU_BATCH_FLUSH. Both paths put back the isolated folio,
> then unlock it and drop the lookup reference before reaching the common
> try_to_unmap_flush().
>
> The page-cache reference does not keep the folio stable once the lock is
> released. Another collapse can replace and free it. With its PTEs
> already gone, that collapse cannot flush the first task's per-task TLB
> batch, and retract_page_tables() skips short or unaligned VMAs. A CPU
> can therefore retain a user translation to the freed folio. This has
> been reproduced with unprivileged MADV_COLLAPSE on a memfd.
>
> Flush at out_unlock while the lookup reference and folio lock are still
> held. The common flush continues to cover the accumulated pagelist on
> both success and rollback, preserving batching on successful collapses.
>
> Fixes: 6d9df8a5889c ("mm/thp: collapse_file() do try_to_unmap(TTU_BATCH_FLUSH)")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Kyle Zeng <kylebot@openai.com>
Good catch, yes, thanks: the folios already on the pagelist were correctly
flushed before reference dropped, but the one failing folio had its
reference dropped too soon.
I might have chosen to fix it differently (holding the final reference),
rather than duplicating the try_to_unmap_flush(); but you've put a nice
comment on its no-op when already flushed (thanks to Zi Yan), so I don't
think it's worth messing around with this good fix you've already tested.
Acked-by: Hugh Dickins <hughd@google.com>
> ---
> Changes in v2:
> - Use Assisted-by: LLM.
> - Explain that the common flush is a no-op after out_unlock flushes.
>
> mm/khugepaged.c | 10 +++++++---
> 1 file changed, 7 insertions(+), 3 deletions(-)
>
> diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> index 75639298efc2..e1a5890818ad 100644
> --- a/mm/khugepaged.c
> +++ b/mm/khugepaged.c
> @@ -2478,6 +2478,11 @@ static enum scan_result collapse_file(struct mm_struct *mm, unsigned long addr,
> index += folio_nr_pages(folio);
> continue;
> out_unlock:
> + /*
> + * The folio may have been unmapped with TTU_BATCH_FLUSH.
> + * Flush before releasing the lock and our last reference.
> + */
> + try_to_unmap_flush();
> folio_unlock(folio);
> folio_put(folio);
> goto xa_unlocked;
> @@ -2488,9 +2493,8 @@ static enum scan_result collapse_file(struct mm_struct *mm, unsigned long addr,
> xa_unlocked:
>
> /*
> - * If collapse is successful, flush must be done now before copying.
> - * If collapse is unsuccessful, does flush actually need to be done?
> - * Do it anyway, to clear the state.
> + * Flush before copying the folios, or releasing them in rollback.
> + * This is a no-op if out_unlock already flushed the batch.
> */
> try_to_unmap_flush();
>
^ permalink raw reply [flat|nested] 5+ messages in thread