* [RFC PATCH] arm64: mm: Fix kernel page tables incorrectly deleted during memory removal
@ 2023-07-17 11:51 Wupeng Ma
2023-07-21 10:36 ` Will Deacon
0 siblings, 1 reply; 4+ messages in thread
From: Wupeng Ma @ 2023-07-17 11:51 UTC (permalink / raw)
To: catalin.marinas, akpm, sudaraja
Cc: linux-mm, linux-kernel, mawupeng1, wangkefeng.wang, will,
linux-arm-kernel, mark.rutland, anshuman.khandual
From: Ma Wupeng <mawupeng1@huawei.com>
During our test, we found that kernel page table may be unexpectedly
cleared with rodata off. The root cause is that the kernel page is
initialized with pud size(1G block mapping) while offline is memory
block size(MIN_MEMORY_BLOCK_SIZE 128M), eg, if 2G memory is hot-added,
when offline a memory block, the call trace is shown below,
offline_and_remove_memory
try_remove_memory
arch_remove_memory
__remove_pgd_mapping
unmap_hotplug_range
unmap_hotplug_p4d_range
unmap_hotplug_pud_range
if (pud_sect(pud))
pud_clear(pudp);
There is no issue for block mapping with pmd level(2M) because the
memory block size is aligned with 2M.
Commit f0b13ee23241 ("arm64/sparsemem: reduce SECTION_SIZE_BITS") reduces
SECTION_SIZE_BITS from arm64, this make memory section size less than pud
size possible. Since only hotadded memory can be removed for arm64 due to
commit bbd6ec605c0f ("arm64/mm: Enable memory hot remove"), stop using pud
size kernel page entry during memory hot join can fix this.
Fixes: f0b13ee23241 ("arm64/sparsemem: reduce SECTION_SIZE_BITS")
Signed-off-by: Ma Wupeng <mawupeng1@huawei.com>
---
arch/arm64/mm/mmu.c | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
index 95d360805f8a..44c724ce4f70 100644
--- a/arch/arm64/mm/mmu.c
+++ b/arch/arm64/mm/mmu.c
@@ -44,6 +44,7 @@
#define NO_BLOCK_MAPPINGS BIT(0)
#define NO_CONT_MAPPINGS BIT(1)
#define NO_EXEC_MAPPINGS BIT(2) /* assumes FEAT_HPDS is not used */
+#define NO_PUD_MAPPINGS BIT(3)
int idmap_t0sz __ro_after_init;
@@ -344,7 +345,7 @@ static void alloc_init_pud(pgd_t *pgdp, unsigned long addr, unsigned long end,
*/
if (pud_sect_supported() &&
((addr | next | phys) & ~PUD_MASK) == 0 &&
- (flags & NO_BLOCK_MAPPINGS) == 0) {
+ (flags & (NO_BLOCK_MAPPINGS | NO_PUD_MAPPINGS)) == 0) {
pud_set_huge(pudp, phys, prot);
/*
@@ -1305,7 +1306,7 @@ struct range arch_get_mappable_range(void)
int arch_add_memory(int nid, u64 start, u64 size,
struct mhp_params *params)
{
- int ret, flags = NO_EXEC_MAPPINGS;
+ int ret, flags = NO_EXEC_MAPPINGS | NO_PUD_MAPPINGS;
VM_BUG_ON(!mhp_range_allowed(start, size, true));
--
2.25.1
_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel
^ permalink raw reply related [flat|nested] 4+ messages in thread* Re: [RFC PATCH] arm64: mm: Fix kernel page tables incorrectly deleted during memory removal 2023-07-17 11:51 [RFC PATCH] arm64: mm: Fix kernel page tables incorrectly deleted during memory removal Wupeng Ma @ 2023-07-21 10:36 ` Will Deacon [not found] ` <35a0dad6-4f3b-f2c3-f835-b13c1e899f8d@huawei.com> 0 siblings, 1 reply; 4+ messages in thread From: Will Deacon @ 2023-07-21 10:36 UTC (permalink / raw) To: Wupeng Ma Cc: catalin.marinas, akpm, sudaraja, linux-mm, linux-kernel, wangkefeng.wang, linux-arm-kernel, mark.rutland, anshuman.khandual On Mon, Jul 17, 2023 at 07:51:50PM +0800, Wupeng Ma wrote: > From: Ma Wupeng <mawupeng1@huawei.com> > > During our test, we found that kernel page table may be unexpectedly > cleared with rodata off. The root cause is that the kernel page is > initialized with pud size(1G block mapping) while offline is memory > block size(MIN_MEMORY_BLOCK_SIZE 128M), eg, if 2G memory is hot-added, > when offline a memory block, the call trace is shown below, > > offline_and_remove_memory > try_remove_memory > arch_remove_memory > __remove_pgd_mapping > unmap_hotplug_range > unmap_hotplug_p4d_range > unmap_hotplug_pud_range > if (pud_sect(pud)) > pud_clear(pudp); Sorry, but I'm struggling to understand the problem here. If we're adding and removing a 2G memory region, why _wouldn't_ we want to use large 1GiB mappings? Or are you saying that only a subset of the memory is removed, but we then accidentally unmap the whole thing? > diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c > index 95d360805f8a..44c724ce4f70 100644 > --- a/arch/arm64/mm/mmu.c > +++ b/arch/arm64/mm/mmu.c > @@ -44,6 +44,7 @@ > #define NO_BLOCK_MAPPINGS BIT(0) > #define NO_CONT_MAPPINGS BIT(1) > #define NO_EXEC_MAPPINGS BIT(2) /* assumes FEAT_HPDS is not used */ > +#define NO_PUD_MAPPINGS BIT(3) > > int idmap_t0sz __ro_after_init; > > @@ -344,7 +345,7 @@ static void alloc_init_pud(pgd_t *pgdp, unsigned long addr, unsigned long end, > */ > if (pud_sect_supported() && > ((addr | next | phys) & ~PUD_MASK) == 0 && > - (flags & NO_BLOCK_MAPPINGS) == 0) { > + (flags & (NO_BLOCK_MAPPINGS | NO_PUD_MAPPINGS)) == 0) { > pud_set_huge(pudp, phys, prot); > > /* > @@ -1305,7 +1306,7 @@ struct range arch_get_mappable_range(void) > int arch_add_memory(int nid, u64 start, u64 size, > struct mhp_params *params) > { > - int ret, flags = NO_EXEC_MAPPINGS; > + int ret, flags = NO_EXEC_MAPPINGS | NO_PUD_MAPPINGS; I think we should allow large mappings here and instead prevent partial removal of the block, if that's what is causing the issue. Will _______________________________________________ linux-arm-kernel mailing list linux-arm-kernel@lists.infradead.org http://lists.infradead.org/mailman/listinfo/linux-arm-kernel ^ permalink raw reply [flat|nested] 4+ messages in thread
[parent not found: <35a0dad6-4f3b-f2c3-f835-b13c1e899f8d@huawei.com>]
[parent not found: <c49c5f19-99d3-0a1f-88c6-03f60587b38c@arm.com>]
[parent not found: <732e0db0-eb41-6c58-85b7-46257b4ba0b7@redhat.com>]
[parent not found: <3149f5f8-7878-dfe1-5745-870fddcc1108@huawei.com>]
[parent not found: <a8775e0b-a206-3ec8-7499-a3c3cfd782e2@redhat.com>]
* Re: [RFC PATCH] arm64: mm: Fix kernel page tables incorrectly deleted during memory removal [not found] ` <a8775e0b-a206-3ec8-7499-a3c3cfd782e2@redhat.com> @ 2023-07-28 1:06 ` mawupeng 2023-07-28 4:01 ` Anshuman Khandual 1 sibling, 0 replies; 4+ messages in thread From: mawupeng @ 2023-07-28 1:06 UTC (permalink / raw) To: david, anshuman.khandual, will Cc: mawupeng1, catalin.marinas, akpm, sudaraja, linux-mm, linux-kernel, wangkefeng.wang, linux-arm-kernel, mark.rutland On 2023/7/26 15:50, David Hildenbrand wrote: > On 26.07.23 08:20, mawupeng wrote: >> >> >> On 2023/7/24 14:11, David Hildenbrand wrote: >>> On 24.07.23 07:54, Anshuman Khandual wrote: >>>> >>>> >>>> On 7/24/23 06:55, mawupeng wrote: >>>>> >>>>> On 2023/7/21 18:36, Will Deacon wrote: >>>>>> On Mon, Jul 17, 2023 at 07:51:50PM +0800, Wupeng Ma wrote: >>>>>>> From: Ma Wupeng <mawupeng1@huawei.com> >>>>>>> >>>>>>> During our test, we found that kernel page table may be unexpectedly >>>>>>> cleared with rodata off. The root cause is that the kernel page is >>>>>>> initialized with pud size(1G block mapping) while offline is memory >>>>>>> block size(MIN_MEMORY_BLOCK_SIZE 128M), eg, if 2G memory is hot-added, >>>>>>> when offline a memory block, the call trace is shown below, >>> >>> Is someone adding memory in 2 GiB granularity and then removing parts of it in 128 MiB granularity? That would be against what we support using the add_memory() / offline_and_remove_memory() API and that driver should be fixed instead. >> >> Yes, this kind of situation. >> >> The problem occurs in the following scenarios: >> 1. use mem=xxG to reserve memory. >> 2. add_momory to online memory. >> 3. offline part of the memroy via offline_and_remove_memory. >> >> During my research, ACPI memory removal use memory_subsys_offline to offline memory section and >> this will not delete page table entry which do not trigger this kind of problem. >> >> So I understand what you are talking about. >> 1. 3rd-party driver shouldn't use add_memory/offline_and_remove_memory to online/offline memory. >> If it have to use, this can be achieved by driver. >> 2. memory_subsys_offline is perfered to do such thing. > > No, my point is that > > 1) If you use add_memory() and offline_and_remove_memory() in the *same > granularity* it has to be working, otherwise it has to be fixed. > > 2) If you use add_memory() and offline_and_remove_memory() in different > granularity (especially, add_memory() in bigger granularity) , then > change your code to do add_memory() in the same granularity. > > > If you run into 1), then we populated a PUD for boot memory that also covers yet unpopulated physical memory ranges that are later populated by add_memory(). If that's the case, then we can either fix it by > > a) Not doing that. Use PMD tables instead for that piece of memory. > > b) Detecting that that PUD still covers memory and refusing to remove > that PUD. > > c) Rejecting to hotadd memory in this situation at that location. We > have mhp_get_pluggable_range() -> arch_get_mappable_range() to kind- > of handle something like that. Thank you for your patient answer. This I do understand and answer my question. > _______________________________________________ linux-arm-kernel mailing list linux-arm-kernel@lists.infradead.org http://lists.infradead.org/mailman/listinfo/linux-arm-kernel ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [RFC PATCH] arm64: mm: Fix kernel page tables incorrectly deleted during memory removal [not found] ` <a8775e0b-a206-3ec8-7499-a3c3cfd782e2@redhat.com> 2023-07-28 1:06 ` mawupeng @ 2023-07-28 4:01 ` Anshuman Khandual 1 sibling, 0 replies; 4+ messages in thread From: Anshuman Khandual @ 2023-07-28 4:01 UTC (permalink / raw) To: David Hildenbrand, mawupeng, will Cc: catalin.marinas, akpm, sudaraja, linux-mm, linux-kernel, wangkefeng.wang, linux-arm-kernel, mark.rutland On 7/26/23 13:20, David Hildenbrand wrote: > On 26.07.23 08:20, mawupeng wrote: >> >> >> On 2023/7/24 14:11, David Hildenbrand wrote: >>> On 24.07.23 07:54, Anshuman Khandual wrote: >>>> >>>> >>>> On 7/24/23 06:55, mawupeng wrote: >>>>> >>>>> On 2023/7/21 18:36, Will Deacon wrote: >>>>>> On Mon, Jul 17, 2023 at 07:51:50PM +0800, Wupeng Ma wrote: >>>>>>> From: Ma Wupeng <mawupeng1@huawei.com> >>>>>>> >>>>>>> During our test, we found that kernel page table may be unexpectedly >>>>>>> cleared with rodata off. The root cause is that the kernel page is >>>>>>> initialized with pud size(1G block mapping) while offline is memory >>>>>>> block size(MIN_MEMORY_BLOCK_SIZE 128M), eg, if 2G memory is hot-added, >>>>>>> when offline a memory block, the call trace is shown below, >>> >>> Is someone adding memory in 2 GiB granularity and then removing parts of it in 128 MiB granularity? That would be against what we support using the add_memory() / offline_and_remove_memory() API and that driver should be fixed instead. >> >> Yes, this kind of situation. >> >> The problem occurs in the following scenarios: >> 1. use mem=xxG to reserve memory. >> 2. add_momory to online memory. >> 3. offline part of the memroy via offline_and_remove_memory. >> >> During my research, ACPI memory removal use memory_subsys_offline to offline memory section and >> this will not delete page table entry which do not trigger this kind of problem. >> >> So I understand what you are talking about. >> 1. 3rd-party driver shouldn't use add_memory/offline_and_remove_memory to online/offline memory. >> If it have to use, this can be achieved by driver. >> 2. memory_subsys_offline is perfered to do such thing. > > No, my point is that > > 1) If you use add_memory() and offline_and_remove_memory() in the *same > granularity* it has to be working, otherwise it has to be fixed. > > 2) If you use add_memory() and offline_and_remove_memory() in different > granularity (especially, add_memory() in bigger granularity) , then > change your code to do add_memory() in the same granularity. > > > If you run into 1), then we populated a PUD for boot memory that also covers yet unpopulated physical memory ranges that are later populated by add_memory(). If that's the case, then we can either fix it by Is that case possible ? __create_pgd_mapping() is called to create the mapping both in hotplug and boot memory cases. alloc_init_pud() ensures [1], that both virtual and physical address ranges are PUD_MASK aligned, before creating a huge or block page entry. (addr | next | phys) & ~PUD_MASK) == 0 > > a) Not doing that. Use PMD tables instead for that piece of memory. > > b) Detecting that that PUD still covers memory and refusing to remove > that PUD. > > c) Rejecting to hotadd memory in this situation at that location. We > have mhp_get_pluggable_range() -> arch_get_mappable_range() to kind- > of handle something like that. > _______________________________________________ linux-arm-kernel mailing list linux-arm-kernel@lists.infradead.org http://lists.infradead.org/mailman/listinfo/linux-arm-kernel ^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2023-07-28 4:01 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2023-07-17 11:51 [RFC PATCH] arm64: mm: Fix kernel page tables incorrectly deleted during memory removal Wupeng Ma
2023-07-21 10:36 ` Will Deacon
[not found] ` <35a0dad6-4f3b-f2c3-f835-b13c1e899f8d@huawei.com>
[not found] ` <c49c5f19-99d3-0a1f-88c6-03f60587b38c@arm.com>
[not found] ` <732e0db0-eb41-6c58-85b7-46257b4ba0b7@redhat.com>
[not found] ` <3149f5f8-7878-dfe1-5745-870fddcc1108@huawei.com>
[not found] ` <a8775e0b-a206-3ec8-7499-a3c3cfd782e2@redhat.com>
2023-07-28 1:06 ` mawupeng
2023-07-28 4:01 ` Anshuman Khandual
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox