* Re: [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios [not found] <20260716095424.471052-1-kirill@shutemov.name> @ 2026-07-18 2:58 ` Andrew Morton 2026-07-20 3:34 ` Matthew Wilcox 2026-07-20 3:26 ` Miaohe Lin 1 sibling, 1 reply; 4+ messages in thread From: Andrew Morton @ 2026-07-18 2:58 UTC (permalink / raw) To: Kiryl Shutsemau Cc: David Hildenbrand, Lorenzo Stoakes, Miaohe Lin, Naoya Horiguchi, Zi Yan, Baolin Wang, Liam R . Howlett, Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang, Usama Arif, Hao Zhang, Hao Zhang, linux-mm, linux-kernel, Kiryl Shutsemau (Meta), stable On Thu, 16 Jul 2026 10:54:24 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote: > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org> > > __folio_split() keeps dereferencing the mapping after the split: > shmem_uncharge(mapping->host) and remap_page() while the folios are still > frozen/locked, and i_mmap_unlock_read(mapping) at the very end, after the > after-split folios have been unlocked and freed. > > Nothing holds an inode reference across that. The split relies on @folio > -- which the beyond-EOF drop loop never removes, as it starts at > folio_next(folio) -- staying locked and in the page cache to hold off > eviction. But the unlock loop unlocks @folio before i_mmap_unlock_read() > runs. If the caller's @lock_at is a tail beyond EOF, as memory_failure() > passes when splitting a poisoned tail of a shmem THP that reaches past > i_size during truncation, it too is gone from the page cache; so once > @folio is unlocked no locked, in-cache folio pins the inode, and a > concurrent final iput() can evict and RCU-free it before > i_mmap_unlock_read() touches i_mmap_rwsem: > > BUG: KASAN: slab-use-after-free in __up_read+0x634/0x790 > i_mmap_unlock_read include/linux/fs.h:537 [inline] > __folio_split+0x732/0x1640 mm/huge_memory.c:4100 > try_to_split_thp_page+0xab/0x390 mm/memory-failure.c:1675 > memory_failure+0x1394/0x26e0 mm/memory-failure.c:2470 Added as a hotfix. thanks. > mm/huge_memory.c | 12 ++++++++++++ > 1 file changed, 12 insertions(+) Sashiko might have found a pre-existing bug in there: https://sashiko.dev/#/patchset/20260716095424.471052-1-kirill@shutemov.name (why is mapping_set_update() a) a macro and b) undocumented?) ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios 2026-07-18 2:58 ` [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios Andrew Morton @ 2026-07-20 3:34 ` Matthew Wilcox 2026-07-20 4:35 ` Kairui Song 0 siblings, 1 reply; 4+ messages in thread From: Matthew Wilcox @ 2026-07-20 3:34 UTC (permalink / raw) To: Andrew Morton Cc: Kiryl Shutsemau, David Hildenbrand, Lorenzo Stoakes, Miaohe Lin, Naoya Horiguchi, Zi Yan, Baolin Wang, Liam R . Howlett, Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang, Usama Arif, Hao Zhang, Hao Zhang, linux-mm, linux-kernel, Kiryl Shutsemau (Meta), stable, Kairui Song On Fri, Jul 17, 2026 at 07:58:51PM -0700, Andrew Morton wrote: > (why is mapping_set_update() a) a macro and b) undocumented?) commit 62e72d2cf702 So make madvise(..., MADV_COLLAPSE) also call xas_set_lru() to pass the list_lru which we may want to insert xa_node into later. And move mapping_set_update to mm/internal.h, and turn into a macro to avoid including extra headers in mm/internal.h. Take it up with Kairui. And whoever it was added that patch to linux-mm. ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios 2026-07-20 3:34 ` Matthew Wilcox @ 2026-07-20 4:35 ` Kairui Song 0 siblings, 0 replies; 4+ messages in thread From: Kairui Song @ 2026-07-20 4:35 UTC (permalink / raw) To: Matthew Wilcox Cc: Andrew Morton, Kiryl Shutsemau, David Hildenbrand, Lorenzo Stoakes, Miaohe Lin, Naoya Horiguchi, Zi Yan, Baolin Wang, Liam R . Howlett, Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang, Usama Arif, Hao Zhang, Hao Zhang, linux-mm, linux-kernel, Kiryl Shutsemau (Meta), stable, Kairui Song On Mon, Jul 20, 2026 at 04:34:37AM +0800, Matthew Wilcox wrote: > On Fri, Jul 17, 2026 at 07:58:51PM -0700, Andrew Morton wrote: > > (why is mapping_set_update() a) a macro and b) undocumented?) > > commit 62e72d2cf702 > > So make madvise(..., MADV_COLLAPSE) also call xas_set_lru() to pass the > list_lru which we may want to insert xa_node into later. And move > mapping_set_update to mm/internal.h, and turn into a macro to avoid > including extra headers in mm/internal.h. > > Take it up with Kairui. And whoever it was added that patch to > linux-mm. Right, I initially wanted to make it a inline helper, since this function were a macro and later become a static helper, so it should be super cheap in the hot path. That requires to include extra headers in internal.h and that looked a bit more uglier and every internel.h user will have extra header dependency. Moving it to pagemap.h also looked confusing along size other mapping_set_* helper, they are completely different things, and this help should stay in mm/. So I just partially restored how it was before b64e74e95aa6, It was never documented though. We can make it a function for sure, open to suggestions on this. ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios [not found] <20260716095424.471052-1-kirill@shutemov.name> 2026-07-18 2:58 ` [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios Andrew Morton @ 2026-07-20 3:26 ` Miaohe Lin 1 sibling, 0 replies; 4+ messages in thread From: Miaohe Lin @ 2026-07-20 3:26 UTC (permalink / raw) To: Kiryl Shutsemau Cc: Zi Yan, Baolin Wang, Liam R . Howlett, Nico Pache, Ryan Roberts, Dev Jain, Barry Song, Lance Yang, Usama Arif, Hao Zhang, Hao Zhang, linux-mm, linux-kernel, Kiryl Shutsemau (Meta), stable, Andrew Morton, David Hildenbrand, Lorenzo Stoakes, Naoya Horiguchi On 2026/7/16 17:54, Kiryl Shutsemau wrote: > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org> > > __folio_split() keeps dereferencing the mapping after the split: > shmem_uncharge(mapping->host) and remap_page() while the folios are still > frozen/locked, and i_mmap_unlock_read(mapping) at the very end, after the > after-split folios have been unlocked and freed. > > Nothing holds an inode reference across that. The split relies on @folio > -- which the beyond-EOF drop loop never removes, as it starts at > folio_next(folio) -- staying locked and in the page cache to hold off > eviction. But the unlock loop unlocks @folio before i_mmap_unlock_read() > runs. If the caller's @lock_at is a tail beyond EOF, as memory_failure() > passes when splitting a poisoned tail of a shmem THP that reaches past > i_size during truncation, it too is gone from the page cache; so once > @folio is unlocked no locked, in-cache folio pins the inode, and a > concurrent final iput() can evict and RCU-free it before > i_mmap_unlock_read() touches i_mmap_rwsem: > > BUG: KASAN: slab-use-after-free in __up_read+0x634/0x790 > i_mmap_unlock_read include/linux/fs.h:537 [inline] > __folio_split+0x732/0x1640 mm/huge_memory.c:4100 > try_to_split_thp_page+0xab/0x390 mm/memory-failure.c:1675 > memory_failure+0x1394/0x26e0 mm/memory-failure.c:2470 > > Freed by task 4601: > shmem_free_in_core_inode+0x54/0xb0 mm/shmem.c:5177 > evict+0x57f/0xac0 fs/inode.c:870 > > Do every mapping dereference while @folio still pins the inode: drop > i_mmap_rwsem right after remap_page(), before the loop that unlocks and > frees the after-split folios, and clear @mapping so the exit path does not > unlock it again. shmem_uncharge() and remap_page() already run before that > point, so after this nothing past the unlock loop touches the inode or the > mapping. > > This is now a rule the split depends on, alongside keeping @folio frozen > until the page cache is updated: no inode or mapping dereference once the > after-split folios start being unlocked. > > Reported-by: Hao Zhang <zhanghao1@kylinos.cn> > Closes: https://lore.kernel.org/linux-mm/20260710071344.GA106129@zh-pc > Fixes: baa355fd3314 ("thp: file pages support for split_huge_page()") > Cc: <stable@vger.kernel.org> > Co-developed-by: Hao Zhang <zhanghao1@kylinos.cn> > Signed-off-by: Hao Zhang <zhanghao1@kylinos.cn> > Acked-by: David Hildenbrand (Arm) <david@kernel.org> > Reviewed-by: Zi Yan <ziy@nvidia.com> > Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com> > Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org> Reviewed-by: Miaohe Lin <linmiaohe@huawei.com> Thanks. . ^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-07-20 4:35 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <20260716095424.471052-1-kirill@shutemov.name>
2026-07-18 2:58 ` [PATCH v3] mm/huge_memory: unlock i_mmap_rwsem before releasing after-split folios Andrew Morton
2026-07-20 3:34 ` Matthew Wilcox
2026-07-20 4:35 ` Kairui Song
2026-07-20 3:26 ` Miaohe Lin
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox