All of lore.kernel.org
 help / color / mirror / Atom feed
From: Muchun Song <muchun.song@linux.dev>
To: Jiakai Xu <xujiakai2025@iscas.ac.cn>
Cc: linux-kernel@vger.kernel.org, linux-riscv@lists.infradead.org,
	David Hildenbrand <david@kernel.org>, Guo Ren <guoren@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Vishal Moola <vishal.moola@gmail.com>,
	Albert Ou <aou@eecs.berkeley.edu>,
	Alexandre Ghiti <alex@ghiti.fr>,
	Andrew Morton <akpm@linux-foundation.org>,
	Junhui Liu <junhui.liu@pigmoral.tech>,
	Kiryl Shutsemau <kas@kernel.org>, Nam Cao <namcao@linutronix.de>,
	Palmer Dabbelt <palmer@dabbelt.com>,
	Paul Walmsley <pjw@kernel.org>,
	Vivian Wang <wangruikang@iscas.ac.cn>
Subject: Re: [PATCH] riscv/mm: use physical alignment for vmemmap_start_pfn
Date: Mon, 20 Jul 2026 17:20:16 +0800	[thread overview]
Message-ID: <599C4370-89C3-4C73-8296-5723D73DDBB8@linux.dev> (raw)
In-Reply-To: <20260716115326.3466926-1-xujiakai2025@iscas.ac.cn>



> On Jul 16, 2026, at 19:53, Jiakai Xu <xujiakai2025@iscas.ac.cn> wrote:
> 
> RISC-V computes vmemmap_start_pfn by rounding phys_ram_base down to
> VMEMMAP_ADDR_ALIGN. That alignment must therefore be expressed in the
> physical-address domain.
> 
> Commit 476849b0fba4 ("riscv/mm: align vmemmap to maximal folio size")
> attempted to account for the maximal folio alignment by feeding
> MAX_FOLIO_VMEMMAP_ALIGN directly into VMEMMAP_ADDR_ALIGN. However,
> MAX_FOLIO_VMEMMAP_ALIGN is measured in bytes of struct page storage,
> whereas VMEMMAP_ADDR_ALIGN is used to align a physical address.
> 
> The mask-based compound_info encoding requires pfn_to_page(0) to be
> naturally aligned to MAX_FOLIO_VMEMMAP_ALIGN. Commit 9f94db4c7eaa
> ("mm/sparse: check memmap alignment for compound_info_has_mask()")
> added a check for that requirement and exposed the unit mismatch on
> systems such as QEMU virt, where the DRAM base is not aligned to
> MAX_FOLIO_NR_PAGES * PAGE_SIZE.
> 
> Convert MAX_FOLIO_VMEMMAP_ALIGN to the equivalent physical alignment
> before using it in VMEMMAP_ADDR_ALIGN. This keeps the existing
> round_down() logic while making the resulting vmemmap base satisfy the
> mask-alignment requirement.
> 
> Fixes: 476849b0fba4 ("riscv/mm: align vmemmap to maximal folio size")
> Signed-off-by: Jiakai Xu <xujiakai2025@iscas.ac.cn>
> Assisted-by: YuanSheng:DeepSeek-V4-Flash

I've always wondered why RISC-V uses a complex logic to calculate the
mapping relationship between vmemmap and PFN. We could easily follow the
x86 approach to make it much simpler.

I previously ran an experiment and found that the system boots perfectly
fine using the following diff code. Granted, I wrote this quite a while ago,
so it might not be directly compatible with the current codebase, but it
should work with some minor adjustments.

In this scenario, the overall processing logic would become significantly
more straightforward, and handling alignment would be much easier as well.

I'm not sure if the following direction is correct (Perhaps this was a
deliberate choice by RISC-V; I'm just not aware of the background behind it),
but from my perspective, it makes the code much simpler. We might need to
reach out to the relevant RISC-V maintainers to confirm whether this aligns
with what's expected.

Thanks,
Muchun

diff --git a/arch/riscv/include/asm/page.h b/arch/riscv/include/asm/page.h
index ffe213ad65a4..f2f67007f482 100644
--- a/arch/riscv/include/asm/page.h
+++ b/arch/riscv/include/asm/page.h
@@ -119,7 +119,6 @@ struct kernel_mapping {

extern struct kernel_mapping kernel_map;
extern phys_addr_t phys_ram_base;
-extern unsigned long vmemmap_start_pfn;

#define is_kernel_mapping(x) \
	((x) >= kernel_map.virt_addr && (x) < (kernel_map.virt_addr + kernel_map.size))
diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgtable.h
index 8bd36ac842eb..067c87062bcf 100644
--- a/arch/riscv/include/asm/pgtable.h
+++ b/arch/riscv/include/asm/pgtable.h
@@ -91,7 +91,7 @@
* Define vmemmap for pfn_to_page & page_to_pfn calls. Needed if kernel
* is configured with CONFIG_SPARSEMEM_VMEMMAP enabled.
*/
-#define vmemmap ((struct page *)VMEMMAP_START - vmemmap_start_pfn)
+#define vmemmap ((struct page *)VMEMMAP_START)

#define PCI_IO_SIZE SZ_16M
#define PCI_IO_END VMEMMAP_START
diff --git a/arch/riscv/mm/init.c b/arch/riscv/mm/init.c
index addb8a9305be..8f86a6d07a53 100644
--- a/arch/riscv/mm/init.c
+++ b/arch/riscv/mm/init.c
@@ -62,13 +62,6 @@ EXPORT_SYMBOL(pgtable_l5_enabled);
phys_addr_t phys_ram_base __ro_after_init;
EXPORT_SYMBOL(phys_ram_base);

-#ifdef CONFIG_SPARSEMEM_VMEMMAP
-#define VMEMMAP_ADDR_ALIGN (1ULL << SECTION_SIZE_BITS)
-
-unsigned long vmemmap_start_pfn __ro_after_init;
-EXPORT_SYMBOL(vmemmap_start_pfn);
-#endif
-
unsigned long empty_zero_page[PAGE_SIZE / sizeof(unsigned long)]
__page_aligned_bss;
EXPORT_SYMBOL(empty_zero_page);
@@ -246,12 +239,8 @@ static void __init setup_bootmem(void)
	* Make sure we align the start of the memory on a PMD boundary so that
	* at worst, we map the linear mapping with PMD mappings.
	*/
- 	if (!IS_ENABLED(CONFIG_XIP_KERNEL)) {
+ 	if (!IS_ENABLED(CONFIG_XIP_KERNEL))
		phys_ram_base = memblock_start_of_DRAM() & PMD_MASK;
-#ifdef CONFIG_SPARSEMEM_VMEMMAP
- 		vmemmap_start_pfn = round_down(phys_ram_base, VMEMMAP_ADDR_ALIGN) >> PAGE_SHIFT;
-#endif
- 	}

	/*
	 * In 64-bit, any use of __va/__pa before this point is wrong as we
@@ -1121,9 +1110,6 @@ asmlinkage void __init setup_vm(uintptr_t dtb_pa)
	kernel_map.xiprom_sz = (uintptr_t)(&_exiprom) - (uintptr_t)(&_xiprom);

	phys_ram_base = CONFIG_PHYS_RAM_BASE;
-#ifdef CONFIG_SPARSEMEM_VMEMMAP
- 	vmemmap_start_pfn = round_down(phys_ram_base, VMEMMAP_ADDR_ALIGN) >> PAGE_SHIFT;
-#endif
	kernel_map.phys_addr = (uintptr_t)CONFIG_PHYS_RAM_BASE;
	kernel_map.size = (uintptr_t)(&_end) - (uintptr_t)(&_start);


_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv

WARNING: multiple messages have this Message-ID (diff)
From: Muchun Song <muchun.song@linux.dev>
To: Jiakai Xu <xujiakai2025@iscas.ac.cn>
Cc: linux-kernel@vger.kernel.org, linux-riscv@lists.infradead.org,
	David Hildenbrand <david@kernel.org>, Guo Ren <guoren@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Vishal Moola <vishal.moola@gmail.com>,
	Albert Ou <aou@eecs.berkeley.edu>,
	Alexandre Ghiti <alex@ghiti.fr>,
	Andrew Morton <akpm@linux-foundation.org>,
	Junhui Liu <junhui.liu@pigmoral.tech>,
	Kiryl Shutsemau <kas@kernel.org>, Nam Cao <namcao@linutronix.de>,
	Palmer Dabbelt <palmer@dabbelt.com>,
	Paul Walmsley <pjw@kernel.org>,
	Vivian Wang <wangruikang@iscas.ac.cn>
Subject: Re: [PATCH] riscv/mm: use physical alignment for vmemmap_start_pfn
Date: Mon, 20 Jul 2026 17:20:16 +0800	[thread overview]
Message-ID: <599C4370-89C3-4C73-8296-5723D73DDBB8@linux.dev> (raw)
In-Reply-To: <20260716115326.3466926-1-xujiakai2025@iscas.ac.cn>



> On Jul 16, 2026, at 19:53, Jiakai Xu <xujiakai2025@iscas.ac.cn> wrote:
> 
> RISC-V computes vmemmap_start_pfn by rounding phys_ram_base down to
> VMEMMAP_ADDR_ALIGN. That alignment must therefore be expressed in the
> physical-address domain.
> 
> Commit 476849b0fba4 ("riscv/mm: align vmemmap to maximal folio size")
> attempted to account for the maximal folio alignment by feeding
> MAX_FOLIO_VMEMMAP_ALIGN directly into VMEMMAP_ADDR_ALIGN. However,
> MAX_FOLIO_VMEMMAP_ALIGN is measured in bytes of struct page storage,
> whereas VMEMMAP_ADDR_ALIGN is used to align a physical address.
> 
> The mask-based compound_info encoding requires pfn_to_page(0) to be
> naturally aligned to MAX_FOLIO_VMEMMAP_ALIGN. Commit 9f94db4c7eaa
> ("mm/sparse: check memmap alignment for compound_info_has_mask()")
> added a check for that requirement and exposed the unit mismatch on
> systems such as QEMU virt, where the DRAM base is not aligned to
> MAX_FOLIO_NR_PAGES * PAGE_SIZE.
> 
> Convert MAX_FOLIO_VMEMMAP_ALIGN to the equivalent physical alignment
> before using it in VMEMMAP_ADDR_ALIGN. This keeps the existing
> round_down() logic while making the resulting vmemmap base satisfy the
> mask-alignment requirement.
> 
> Fixes: 476849b0fba4 ("riscv/mm: align vmemmap to maximal folio size")
> Signed-off-by: Jiakai Xu <xujiakai2025@iscas.ac.cn>
> Assisted-by: YuanSheng:DeepSeek-V4-Flash

I've always wondered why RISC-V uses a complex logic to calculate the
mapping relationship between vmemmap and PFN. We could easily follow the
x86 approach to make it much simpler.

I previously ran an experiment and found that the system boots perfectly
fine using the following diff code. Granted, I wrote this quite a while ago,
so it might not be directly compatible with the current codebase, but it
should work with some minor adjustments.

In this scenario, the overall processing logic would become significantly
more straightforward, and handling alignment would be much easier as well.

I'm not sure if the following direction is correct (Perhaps this was a
deliberate choice by RISC-V; I'm just not aware of the background behind it),
but from my perspective, it makes the code much simpler. We might need to
reach out to the relevant RISC-V maintainers to confirm whether this aligns
with what's expected.

Thanks,
Muchun

diff --git a/arch/riscv/include/asm/page.h b/arch/riscv/include/asm/page.h
index ffe213ad65a4..f2f67007f482 100644
--- a/arch/riscv/include/asm/page.h
+++ b/arch/riscv/include/asm/page.h
@@ -119,7 +119,6 @@ struct kernel_mapping {

extern struct kernel_mapping kernel_map;
extern phys_addr_t phys_ram_base;
-extern unsigned long vmemmap_start_pfn;

#define is_kernel_mapping(x) \
	((x) >= kernel_map.virt_addr && (x) < (kernel_map.virt_addr + kernel_map.size))
diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgtable.h
index 8bd36ac842eb..067c87062bcf 100644
--- a/arch/riscv/include/asm/pgtable.h
+++ b/arch/riscv/include/asm/pgtable.h
@@ -91,7 +91,7 @@
* Define vmemmap for pfn_to_page & page_to_pfn calls. Needed if kernel
* is configured with CONFIG_SPARSEMEM_VMEMMAP enabled.
*/
-#define vmemmap ((struct page *)VMEMMAP_START - vmemmap_start_pfn)
+#define vmemmap ((struct page *)VMEMMAP_START)

#define PCI_IO_SIZE SZ_16M
#define PCI_IO_END VMEMMAP_START
diff --git a/arch/riscv/mm/init.c b/arch/riscv/mm/init.c
index addb8a9305be..8f86a6d07a53 100644
--- a/arch/riscv/mm/init.c
+++ b/arch/riscv/mm/init.c
@@ -62,13 +62,6 @@ EXPORT_SYMBOL(pgtable_l5_enabled);
phys_addr_t phys_ram_base __ro_after_init;
EXPORT_SYMBOL(phys_ram_base);

-#ifdef CONFIG_SPARSEMEM_VMEMMAP
-#define VMEMMAP_ADDR_ALIGN (1ULL << SECTION_SIZE_BITS)
-
-unsigned long vmemmap_start_pfn __ro_after_init;
-EXPORT_SYMBOL(vmemmap_start_pfn);
-#endif
-
unsigned long empty_zero_page[PAGE_SIZE / sizeof(unsigned long)]
__page_aligned_bss;
EXPORT_SYMBOL(empty_zero_page);
@@ -246,12 +239,8 @@ static void __init setup_bootmem(void)
	* Make sure we align the start of the memory on a PMD boundary so that
	* at worst, we map the linear mapping with PMD mappings.
	*/
- 	if (!IS_ENABLED(CONFIG_XIP_KERNEL)) {
+ 	if (!IS_ENABLED(CONFIG_XIP_KERNEL))
		phys_ram_base = memblock_start_of_DRAM() & PMD_MASK;
-#ifdef CONFIG_SPARSEMEM_VMEMMAP
- 		vmemmap_start_pfn = round_down(phys_ram_base, VMEMMAP_ADDR_ALIGN) >> PAGE_SHIFT;
-#endif
- 	}

	/*
	 * In 64-bit, any use of __va/__pa before this point is wrong as we
@@ -1121,9 +1110,6 @@ asmlinkage void __init setup_vm(uintptr_t dtb_pa)
	kernel_map.xiprom_sz = (uintptr_t)(&_exiprom) - (uintptr_t)(&_xiprom);

	phys_ram_base = CONFIG_PHYS_RAM_BASE;
-#ifdef CONFIG_SPARSEMEM_VMEMMAP
- 	vmemmap_start_pfn = round_down(phys_ram_base, VMEMMAP_ADDR_ALIGN) >> PAGE_SHIFT;
-#endif
	kernel_map.phys_addr = (uintptr_t)CONFIG_PHYS_RAM_BASE;
	kernel_map.size = (uintptr_t)(&_end) - (uintptr_t)(&_start);


  parent reply	other threads:[~2026-07-20  9:21 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-16 11:53 [PATCH] riscv/mm: use physical alignment for vmemmap_start_pfn Jiakai Xu
2026-07-16 11:53 ` Jiakai Xu
2026-07-16 11:53 ` Jiakai Xu
2026-07-18  2:23 ` Andrew Morton
2026-07-18  2:23   ` Andrew Morton
2026-07-20  8:52 ` David Hildenbrand (Arm)
2026-07-20  8:52   ` David Hildenbrand (Arm)
2026-07-20  9:20 ` Muchun Song [this message]
2026-07-20  9:20   ` Muchun Song
2026-07-22 12:29   ` Kiryl Shutsemau
2026-07-22 12:29     ` Kiryl Shutsemau
2026-07-22 12:32 ` Kiryl Shutsemau
2026-07-22 12:32   ` Kiryl Shutsemau
  -- strict thread matches above, loose matches on Subject: below --
2026-07-16 11:49 Jiakai Xu
2026-07-16 11:49 ` Jiakai Xu
2026-07-16 12:04 ` Jiakai Xu
2026-07-16 12:04   ` Jiakai Xu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=599C4370-89C3-4C73-8296-5723D73DDBB8@linux.dev \
    --to=muchun.song@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=alex@ghiti.fr \
    --cc=aou@eecs.berkeley.edu \
    --cc=david@kernel.org \
    --cc=guoren@kernel.org \
    --cc=junhui.liu@pigmoral.tech \
    --cc=kas@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-riscv@lists.infradead.org \
    --cc=namcao@linutronix.de \
    --cc=palmer@dabbelt.com \
    --cc=pjw@kernel.org \
    --cc=rppt@kernel.org \
    --cc=vishal.moola@gmail.com \
    --cc=wangruikang@iscas.ac.cn \
    --cc=xujiakai2025@iscas.ac.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.