The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios
@ 2026-07-16  9:54 Kiryl Shutsemau
  2026-07-18  2:58 ` Andrew Morton
  2026-07-20  3:26 ` Miaohe Lin
  0 siblings, 2 replies; 5+ messages in thread
From: Kiryl Shutsemau @ 2026-07-16  9:54 UTC (permalink / raw)
  To: Andrew Morton, David Hildenbrand, Lorenzo Stoakes, Miaohe Lin,
	Naoya Horiguchi
  Cc: Zi Yan, Baolin Wang, Liam R . Howlett, Nico Pache, Ryan Roberts,
	Dev Jain, Barry Song, Lance Yang, Usama Arif, Hao Zhang,
	Hao Zhang, linux-mm, linux-kernel, Kiryl Shutsemau (Meta), stable

From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>

__folio_split() keeps dereferencing the mapping after the split:
shmem_uncharge(mapping->host) and remap_page() while the folios are still
frozen/locked, and i_mmap_unlock_read(mapping) at the very end, after the
after-split folios have been unlocked and freed.

Nothing holds an inode reference across that. The split relies on @folio
-- which the beyond-EOF drop loop never removes, as it starts at
folio_next(folio) -- staying locked and in the page cache to hold off
eviction. But the unlock loop unlocks @folio before i_mmap_unlock_read()
runs. If the caller's @lock_at is a tail beyond EOF, as memory_failure()
passes when splitting a poisoned tail of a shmem THP that reaches past
i_size during truncation, it too is gone from the page cache; so once
@folio is unlocked no locked, in-cache folio pins the inode, and a
concurrent final iput() can evict and RCU-free it before
i_mmap_unlock_read() touches i_mmap_rwsem:

  BUG: KASAN: slab-use-after-free in __up_read+0x634/0x790
   i_mmap_unlock_read include/linux/fs.h:537 [inline]
   __folio_split+0x732/0x1640 mm/huge_memory.c:4100
   try_to_split_thp_page+0xab/0x390 mm/memory-failure.c:1675
   memory_failure+0x1394/0x26e0 mm/memory-failure.c:2470

  Freed by task 4601:
   shmem_free_in_core_inode+0x54/0xb0 mm/shmem.c:5177
   evict+0x57f/0xac0 fs/inode.c:870

Do every mapping dereference while @folio still pins the inode: drop
i_mmap_rwsem right after remap_page(), before the loop that unlocks and
frees the after-split folios, and clear @mapping so the exit path does not
unlock it again. shmem_uncharge() and remap_page() already run before that
point, so after this nothing past the unlock loop touches the inode or the
mapping.

This is now a rule the split depends on, alongside keeping @folio frozen
until the page cache is updated: no inode or mapping dereference once the
after-split folios start being unlocked.

Reported-by: Hao Zhang <zhanghao1@kylinos.cn>
Closes: https://lore.kernel.org/linux-mm/20260710071344.GA106129@zh-pc
Fixes: baa355fd3314 ("thp: file pages support for split_huge_page()")
Cc: <stable@vger.kernel.org>
Co-developed-by: Hao Zhang <zhanghao1@kylinos.cn>
Signed-off-by: Hao Zhang <zhanghao1@kylinos.cn>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Zi Yan <ziy@nvidia.com>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
---

Same patch as posted in the v2 discussion [1], now with tags collected.

v3:
 - Move i_mmap_unlock_read() before the after-split unlock loop instead
   of pinning the inode or refusing the split; based on Hao's original
   patch [2] with the race analysis corrected.
v2 [3]:
 - Anchor the split on the head in memory_failure() (broken: the caller's
   reference must land on the poisoned page) + -EBUSY guard in
   __folio_split() (over-rejects for new_order > 0).
v1 [4]:
 - igrab()/iput() around the split (self-deadlocks on final iput() under
   the folio lock).

[1] https://lore.kernel.org/linux-mm/aldjhtfVByHDQXe6@thinkstation
[2] https://lore.kernel.org/linux-mm/20260710071344.GA106129@zh-pc
[3] https://lore.kernel.org/linux-mm/20260714122344.351895-1-kirill@shutemov.name
[4] https://lore.kernel.org/linux-mm/20260713170915.239819-1-kirill@shutemov.name

 mm/huge_memory.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 2bccb0a53a0a..abaea34ef558 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -4109,6 +4109,18 @@ static int __folio_split(struct folio *folio, unsigned int new_order,
 
 	remap_page(folio, 1 << old_order, ttu_flags);
 
+	/*
+	 * Drop the mapping while the inode is still pinned. @folio stays
+	 * locked and present in the page cache until the loop below, so
+	 * eviction cannot free the inode yet; @lock_at is not enough, it may
+	 * be a tail beyond EOF that the split already dropped from the page
+	 * cache. Nothing past this point may touch the inode or the mapping.
+	 */
+	if (mapping) {
+		i_mmap_unlock_read(mapping);
+		mapping = NULL;
+	}
+
 	/*
 	 * Unlock all after-split folios except the one containing
 	 * @lock_at page. If @folio is not split, it will be kept locked.

base-commit: 0e35b9b6ec0ffcc5e23cbdec09f5c622ad532b53
-- 
2.54.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios
  2026-07-16  9:54 [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios Kiryl Shutsemau
@ 2026-07-18  2:58 ` Andrew Morton
  2026-07-20  3:34   ` Matthew Wilcox
  2026-07-20  3:26 ` Miaohe Lin
  1 sibling, 1 reply; 5+ messages in thread
From: Andrew Morton @ 2026-07-18  2:58 UTC (permalink / raw)
  To: Kiryl Shutsemau
  Cc: David Hildenbrand, Lorenzo Stoakes, Miaohe Lin, Naoya Horiguchi,
	Zi Yan, Baolin Wang, Liam R . Howlett, Nico Pache, Ryan Roberts,
	Dev Jain, Barry Song, Lance Yang, Usama Arif, Hao Zhang,
	Hao Zhang, linux-mm, linux-kernel, Kiryl Shutsemau (Meta), stable

On Thu, 16 Jul 2026 10:54:24 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:

> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> 
> __folio_split() keeps dereferencing the mapping after the split:
> shmem_uncharge(mapping->host) and remap_page() while the folios are still
> frozen/locked, and i_mmap_unlock_read(mapping) at the very end, after the
> after-split folios have been unlocked and freed.
> 
> Nothing holds an inode reference across that. The split relies on @folio
> -- which the beyond-EOF drop loop never removes, as it starts at
> folio_next(folio) -- staying locked and in the page cache to hold off
> eviction. But the unlock loop unlocks @folio before i_mmap_unlock_read()
> runs. If the caller's @lock_at is a tail beyond EOF, as memory_failure()
> passes when splitting a poisoned tail of a shmem THP that reaches past
> i_size during truncation, it too is gone from the page cache; so once
> @folio is unlocked no locked, in-cache folio pins the inode, and a
> concurrent final iput() can evict and RCU-free it before
> i_mmap_unlock_read() touches i_mmap_rwsem:
> 
>   BUG: KASAN: slab-use-after-free in __up_read+0x634/0x790
>    i_mmap_unlock_read include/linux/fs.h:537 [inline]
>    __folio_split+0x732/0x1640 mm/huge_memory.c:4100
>    try_to_split_thp_page+0xab/0x390 mm/memory-failure.c:1675
>    memory_failure+0x1394/0x26e0 mm/memory-failure.c:2470

Added as a hotfix. thanks.

>  mm/huge_memory.c | 12 ++++++++++++
>  1 file changed, 12 insertions(+)

Sashiko might have found a pre-existing bug in there:

	https://sashiko.dev/#/patchset/20260716095424.471052-1-kirill@shutemov.name

(why is mapping_set_update() a) a macro and b) undocumented?)


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios
  2026-07-16  9:54 [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios Kiryl Shutsemau
  2026-07-18  2:58 ` Andrew Morton
@ 2026-07-20  3:26 ` Miaohe Lin
  1 sibling, 0 replies; 5+ messages in thread
From: Miaohe Lin @ 2026-07-20  3:26 UTC (permalink / raw)
  To: Kiryl Shutsemau
  Cc: Zi Yan, Baolin Wang, Liam R . Howlett, Nico Pache, Ryan Roberts,
	Dev Jain, Barry Song, Lance Yang, Usama Arif, Hao Zhang,
	Hao Zhang, linux-mm, linux-kernel, Kiryl Shutsemau (Meta), stable,
	Andrew Morton, David Hildenbrand, Lorenzo Stoakes,
	Naoya Horiguchi

On 2026/7/16 17:54, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> 
> __folio_split() keeps dereferencing the mapping after the split:
> shmem_uncharge(mapping->host) and remap_page() while the folios are still
> frozen/locked, and i_mmap_unlock_read(mapping) at the very end, after the
> after-split folios have been unlocked and freed.
> 
> Nothing holds an inode reference across that. The split relies on @folio
> -- which the beyond-EOF drop loop never removes, as it starts at
> folio_next(folio) -- staying locked and in the page cache to hold off
> eviction. But the unlock loop unlocks @folio before i_mmap_unlock_read()
> runs. If the caller's @lock_at is a tail beyond EOF, as memory_failure()
> passes when splitting a poisoned tail of a shmem THP that reaches past
> i_size during truncation, it too is gone from the page cache; so once
> @folio is unlocked no locked, in-cache folio pins the inode, and a
> concurrent final iput() can evict and RCU-free it before
> i_mmap_unlock_read() touches i_mmap_rwsem:
> 
>   BUG: KASAN: slab-use-after-free in __up_read+0x634/0x790
>    i_mmap_unlock_read include/linux/fs.h:537 [inline]
>    __folio_split+0x732/0x1640 mm/huge_memory.c:4100
>    try_to_split_thp_page+0xab/0x390 mm/memory-failure.c:1675
>    memory_failure+0x1394/0x26e0 mm/memory-failure.c:2470
> 
>   Freed by task 4601:
>    shmem_free_in_core_inode+0x54/0xb0 mm/shmem.c:5177
>    evict+0x57f/0xac0 fs/inode.c:870
> 
> Do every mapping dereference while @folio still pins the inode: drop
> i_mmap_rwsem right after remap_page(), before the loop that unlocks and
> frees the after-split folios, and clear @mapping so the exit path does not
> unlock it again. shmem_uncharge() and remap_page() already run before that
> point, so after this nothing past the unlock loop touches the inode or the
> mapping.
> 
> This is now a rule the split depends on, alongside keeping @folio frozen
> until the page cache is updated: no inode or mapping dereference once the
> after-split folios start being unlocked.
> 
> Reported-by: Hao Zhang <zhanghao1@kylinos.cn>
> Closes: https://lore.kernel.org/linux-mm/20260710071344.GA106129@zh-pc
> Fixes: baa355fd3314 ("thp: file pages support for split_huge_page()")
> Cc: <stable@vger.kernel.org>
> Co-developed-by: Hao Zhang <zhanghao1@kylinos.cn>
> Signed-off-by: Hao Zhang <zhanghao1@kylinos.cn>
> Acked-by: David Hildenbrand (Arm) <david@kernel.org>
> Reviewed-by: Zi Yan <ziy@nvidia.com>
> Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>

Reviewed-by: Miaohe Lin <linmiaohe@huawei.com>

Thanks.
.

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios
  2026-07-18  2:58 ` Andrew Morton
@ 2026-07-20  3:34   ` Matthew Wilcox
  2026-07-20  4:35     ` Kairui Song
  0 siblings, 1 reply; 5+ messages in thread
From: Matthew Wilcox @ 2026-07-20  3:34 UTC (permalink / raw)
  To: Andrew Morton
  Cc: Kiryl Shutsemau, David Hildenbrand, Lorenzo Stoakes, Miaohe Lin,
	Naoya Horiguchi, Zi Yan, Baolin Wang, Liam R . Howlett,
	Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang,
	Usama Arif, Hao Zhang, Hao Zhang, linux-mm, linux-kernel,
	Kiryl Shutsemau (Meta), stable, Kairui Song

On Fri, Jul 17, 2026 at 07:58:51PM -0700, Andrew Morton wrote:
> (why is mapping_set_update() a) a macro and b) undocumented?)

commit 62e72d2cf702

    So make madvise(..., MADV_COLLAPSE) also call xas_set_lru() to pass the
    list_lru which we may want to insert xa_node into later.  And move
    mapping_set_update to mm/internal.h, and turn into a macro to avoid
    including extra headers in mm/internal.h.

Take it up with Kairui.  And whoever it was added that patch to
linux-mm.

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios
  2026-07-20  3:34   ` Matthew Wilcox
@ 2026-07-20  4:35     ` Kairui Song
  0 siblings, 0 replies; 5+ messages in thread
From: Kairui Song @ 2026-07-20  4:35 UTC (permalink / raw)
  To: Matthew Wilcox
  Cc: Andrew Morton, Kiryl Shutsemau, David Hildenbrand,
	Lorenzo Stoakes, Miaohe Lin, Naoya Horiguchi, Zi Yan, Baolin Wang,
	Liam R . Howlett, Nico Pache, Ryan Roberts, Dev Jain, Barry Song,
	Lance Yang, Usama Arif, Hao Zhang, Hao Zhang, linux-mm,
	linux-kernel, Kiryl Shutsemau (Meta), stable, Kairui Song

On Mon, Jul 20, 2026 at 04:34:37AM +0800, Matthew Wilcox wrote:
> On Fri, Jul 17, 2026 at 07:58:51PM -0700, Andrew Morton wrote:
> > (why is mapping_set_update() a) a macro and b) undocumented?)
> 
> commit 62e72d2cf702
> 
>     So make madvise(..., MADV_COLLAPSE) also call xas_set_lru() to pass the
>     list_lru which we may want to insert xa_node into later.  And move
>     mapping_set_update to mm/internal.h, and turn into a macro to avoid
>     including extra headers in mm/internal.h.
> 
> Take it up with Kairui.  And whoever it was added that patch to
> linux-mm.

Right, I initially wanted to make it a inline helper, since this function
were a macro and later become a static helper, so it should be super cheap
in the hot path. That requires to include extra headers in internal.h
and that looked a bit more uglier and every internel.h user will have
extra header dependency. Moving it to pagemap.h also looked confusing
along size other mapping_set_* helper, they are completely different
things, and this help should stay in mm/.

So I just partially restored how it was before b64e74e95aa6, It was
never documented though.

We can make it a function for sure, open to suggestions on this.

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-07-20  4:35 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-16  9:54 [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios Kiryl Shutsemau
2026-07-18  2:58 ` Andrew Morton
2026-07-20  3:34   ` Matthew Wilcox
2026-07-20  4:35     ` Kairui Song
2026-07-20  3:26 ` Miaohe Lin

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox