From: Will Deacon <will@kernel.org>
To: Zizhi Wo <wozizhi@huaweicloud.com>
Cc: jack@suse.com, brauner@kernel.org, hch@lst.de,
akpm@linux-foundation.org, linux@armlinux.org.uk,
linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-mm@kvack.org, linux-arm-kernel@lists.infradead.org,
yangerkun@huawei.com, wangkefeng.wang@huawei.com,
pangliyuan1@huawei.com, xieyuanbin1@huawei.com
Subject: Re: [Bug report] hash_name() may cross page boundary and trigger sleep in RCU context
Date: Fri, 28 Nov 2025 12:25:01 +0000 [thread overview]
Message-ID: <aSmUnZZATTn3JD7m@willie-the-truck> (raw)
In-Reply-To: <61757d05-ffce-476d-9b07-88332e5db1b9@huaweicloud.com>
On Fri, Nov 28, 2025 at 09:39:45AM +0800, Zizhi Wo wrote:
> 在 2025/11/28 9:18, Zizhi Wo 写道:
> > 在 2025/11/28 9:17, Zizhi Wo 写道:
> > > 在 2025/11/27 20:59, Will Deacon 写道:
> > > > On Wed, Nov 26, 2025 at 05:05:05PM +0800, Zizhi Wo wrote:
> > > > > We're running into the following issue on an ARM32 platform
> > > > > with the linux
> > > > > 5.10 kernel:
> > > > >
> > > > > [<c0300b78>] (__dabt_svc) from [<c0529cb8>]
> > > > > (link_path_walk.part.7+0x108/0x45c)
> > > > > [<c0529cb8>] (link_path_walk.part.7) from [<c052a948>]
> > > > > (path_openat+0xc4/0x10ec)
> > > > > [<c052a948>] (path_openat) from [<c052cf90>] (do_filp_open+0x9c/0x114)
> > > > > [<c052cf90>] (do_filp_open) from [<c0511e4c>]
> > > > > (do_sys_openat2+0x418/0x528)
> > > > > [<c0511e4c>] (do_sys_openat2) from [<c0513d98>] (do_sys_open+0x88/0xe4)
> > > > > [<c0513d98>] (do_sys_open) from [<c03000c0>]
> > > > > (ret_fast_syscall+0x0/0x58)
> > > > > ...
> > > > > [<c0315e34>] (unwind_backtrace) from [<c030f2b0>]
> > > > > (show_stack+0x20/0x24)
> > > > > [<c030f2b0>] (show_stack) from [<c14239f4>] (dump_stack+0xd8/0xf8)
> > > > > [<c14239f4>] (dump_stack) from [<c038d188>]
> > > > > (___might_sleep+0x19c/0x1e4)
> > > > > [<c038d188>] (___might_sleep) from [<c031b6fc>]
> > > > > (do_page_fault+0x2f8/0x51c)
> > > > > [<c031b6fc>] (do_page_fault) from [<c031bb44>]
> > > > > (do_DataAbort+0x90/0x118)
> > > > > [<c031bb44>] (do_DataAbort) from [<c0300b78>] (__dabt_svc+0x58/0x80)
> > > > > ...
> > > > >
> > > > > During the execution of
> > > > > hash_name()->load_unaligned_zeropad(), a potential
> > > > > memory access beyond the PAGE boundary may occur. For example, when the
> > > > > filename length is near the PAGE_SIZE boundary. This
> > > > > triggers a page fault,
> > > > > which leads to a call to
> > > > > do_page_fault()->mmap_read_trylock(). If we can't
> > > > > acquire the lock, we have to fall back to the
> > > > > mmap_read_lock() path, which
> > > > > calls might_sleep(). This breaks RCU semantics because path
> > > > > lookup occurs
> > > > > under an RCU read-side critical section. In linux-mainline, arm/arm64
> > > > > do_page_fault() still has this problem:
> > > > >
> > > > > lock_mm_and_find_vma->get_mmap_lock_carefully->mmap_read_lock_killable.
> > > > >
> > > > > And before commit bfcfaa77bdf0 ("vfs: use 'unsigned long' accesses for
> > > > > dcache name comparison and hashing"), hash_name accessed the
> > > > > name byte by
> > > > > byte.
> > > > >
> > > > > To prevent load_unaligned_zeropad() from accessing beyond
> > > > > the valid memory
> > > > > region, we would need to intercept such cases beforehand? But doing so
> > > > > would require replicating the internal logic of
> > > > > load_unaligned_zeropad(),
> > > > > including handling endianness and constructing the correct
> > > > > value manually.
> > > > > Given that load_unaligned_zeropad() is used in many places across the
> > > > > kernel, we currently haven't found a good solution to
> > > > > address this cleanly.
> > > > >
> > > > > What would be the recommended way to handle this situation? Would
> > > > > appreciate any feedback and guidance from the community. Thanks!
> > > >
> > > > Does it help if you bodge the translation fault handler along the lines
> > > > of the untested diff below?
>
> I tried it out and it works — thank you for the solution you provided.
Thanks for giving it a spin.
> At the same time, since I’m a beginner in this area, I’d like to ask a
> question.
>
> The comment above do_translation_fault() says:
> “We enter here because the first level page table doesn't contain a
> valid entry for the address.”
>
> However, after modifying the code, it seems that when encountering
> FSR_FS_INVALID_PAGE, the kernel no longer creates a page table entry,
> but instead directly jumps to bad_area.
FSR_FS_INVALID_PAGE indicates a last level translation fault (that's the
"page" part) so it's only applicable in the case where the other levels
of page-table have been populated already.
I wondered about checking !is_vmalloc_addr() too, but I couldn't
convince myself that load_unaligned_zeropad() is only ever used with the
linear map.
> I'd like to ask — could this change potentially cause any other side
> effects?
There's always the possibility but I personally think it's more
self-contained than the other patches doing the rounds. For example, I
don't make any changes to the permission fault handling path.
Will
next prev parent reply other threads:[~2025-11-28 12:25 UTC|newest]
Thread overview: 65+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-11-26 9:05 [Bug report] hash_name() may cross page boundary and trigger sleep in RCU context Zizhi Wo
2025-11-26 10:19 ` [RFC PATCH] vfs: Fix might sleep in load_unaligned_zeropad() with rcu read lock held Xie Yuanbin
2025-11-26 18:10 ` Al Viro
2025-11-26 18:48 ` Al Viro
2025-11-26 19:05 ` Russell King (Oracle)
2025-11-26 19:26 ` Al Viro
2025-11-26 19:51 ` Russell King (Oracle)
2025-11-26 20:02 ` Al Viro
2025-11-26 22:25 ` david laight
2025-11-26 23:51 ` Al Viro
2025-11-26 23:31 ` Russell King (Oracle)
2025-11-27 3:03 ` Xie Yuanbin
2025-11-27 7:20 ` Sebastian Andrzej Siewior
2025-11-27 11:20 ` Xie Yuanbin
2025-11-28 1:39 ` Xie Yuanbin
2025-11-26 20:42 ` Al Viro
2025-11-26 10:27 ` [Bug report] hash_name() may cross page boundary and trigger sleep in RCU context Zizhi Wo
2025-11-26 21:12 ` Linus Torvalds
2025-11-27 10:27 ` Will Deacon
2025-11-27 10:57 ` Russell King (Oracle)
2025-11-28 17:06 ` Linus Torvalds
2025-11-29 1:01 ` Zizhi Wo
2025-11-29 1:35 ` Linus Torvalds
2025-11-29 4:08 ` [Bug report] hash_name() may cross page boundary and trigger Xie Yuanbin
2025-11-29 9:08 ` Al Viro
2025-11-29 9:25 ` Xie Yuanbin
2025-11-29 9:44 ` Al Viro
2025-11-29 10:05 ` Xie Yuanbin
2025-11-29 10:45 ` david laight
2025-11-29 8:54 ` [Bug report] hash_name() may cross page boundary and trigger sleep in RCU context Al Viro
2025-12-01 2:08 ` Zizhi Wo
2025-11-29 2:18 ` [Bug report] hash_name() may cross page boundary and trigger Xie Yuanbin
2025-12-01 13:28 ` [Bug report] hash_name() may cross page boundary and trigger sleep in RCU context Will Deacon
2025-12-02 12:43 ` Russell King (Oracle)
2025-12-02 13:02 ` Xie Yuanbin
2025-12-02 22:07 ` Linus Torvalds
2025-12-03 1:48 ` Xie Yuanbin
2025-12-05 12:08 ` Russell King (Oracle)
2025-12-08 2:32 ` Xie Yuanbin
2025-12-08 9:26 ` David Laight
2025-12-08 10:07 ` Russell King (Oracle)
2025-12-08 13:18 ` Xie Yuanbin
2025-12-08 15:43 ` Russell King (Oracle)
2025-12-09 1:30 ` Xie Yuanbin
2025-11-26 18:55 ` Al Viro
2025-11-27 2:24 ` Zizhi Wo
2025-11-29 3:37 ` Al Viro
2025-11-30 3:01 ` [RFC][alpha] saner vmalloc handling (was Re: [Bug report] hash_name() may cross page boundary and trigger sleep in RCU context) Al Viro
2025-11-30 11:32 ` david laight
2025-11-30 16:43 ` Al Viro
2025-11-30 18:14 ` Magnus Lindholm
2025-11-30 19:03 ` david laight
2025-11-30 20:31 ` Al Viro
2025-11-30 20:32 ` Al Viro
2025-11-30 22:16 ` Linus Torvalds
2025-11-30 23:37 ` Al Viro
2025-12-01 2:03 ` [Bug report] hash_name() may cross page boundary and trigger sleep in RCU context Zizhi Wo
2025-11-27 12:59 ` Will Deacon
2025-11-28 1:17 ` Zizhi Wo
2025-11-28 1:18 ` Zizhi Wo
2025-11-28 1:39 ` Zizhi Wo
2025-11-28 12:25 ` Will Deacon [this message]
2025-11-29 1:02 ` Zizhi Wo
2025-11-29 3:55 ` Al Viro
2025-12-01 2:38 ` Zizhi Wo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aSmUnZZATTn3JD7m@willie-the-truck \
--to=will@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=brauner@kernel.org \
--cc=hch@lst.de \
--cc=jack@suse.com \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux@armlinux.org.uk \
--cc=pangliyuan1@huawei.com \
--cc=wangkefeng.wang@huawei.com \
--cc=wozizhi@huaweicloud.com \
--cc=xieyuanbin1@huawei.com \
--cc=yangerkun@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.