Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v5 0/3] mm: make persistent huge zero folio read-only
@ 2026-07-27 14:34 Xueyuan Chen
  2026-07-27 14:34 ` [PATCH v5 1/3] " Xueyuan Chen
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Xueyuan Chen @ 2026-07-27 14:34 UTC (permalink / raw)
  To: linux-mm, akpm, david, ljs
  Cc: linux-kernel, linux-arm-kernel, catalin.marinas, will, tglx,
	mingo, bp, dave.hansen, x86, hpa, luto, peterz, lance.yang,
	usama.arif, jannh, yang, rppt, ziy, baolin.wang, liam, npache,
	ryan.roberts, dev.jain, baohua

The persistent huge zero folio is shared globally and should stay zero
after initialization. As Jann Horn pointed out[1], kernel bugs have ended
up writing to pages that were meant to be read-only, including in
security-sensitive cases. Making the folio read-only in the direct map
turns such writes into faults instead of silent zero-page corruption.

This series adds set_direct_map_ro_noflush() so mm code can make a
direct-map range read-only, then uses it for the persistent huge zero
folio. The helper is direct-map specific, takes an address-based range as
discussed for set_direct_map* helpers[2], and leaves TLB invalidation to
the caller.

The folio is allocated and zeroed through the writable direct map before
thp_shrinker_init() changes its permissions. thp_shrinker_init() is called
from hugepage_init(), which is registered as a subsys_initcall and runs
after SMP initialization. Writable TLB entries may therefore already be
cached when the page-table permissions change. Patch 1 flushes the exact
direct-map range immediately after the noflush page-table update. Keeping
the flush at the call site preserves the helper's explicit noflush contract
and follows existing direct-map helper users such as secretmem and
hibernation.

This is the first non-RFC posting of the series. The direct-map helper
interface and its TLB invalidation contract have stabilized through the
RFC discussion, and the series is now intended for regular review and
possible merging.

Patches 2 and 3 add arm64 and x86 implementations.

Link: https://lore.kernel.org/linux-mm/20260508-ro-zeropage-v1-1-9808abc20b49@google.com/ [1]
Link: https://lore.kernel.org/linux-mm/0e5b23a6-4895-454a-9dfa-6dc21adc2991@kernel.org/ [2]
Link: https://lore.kernel.org/linux-mm/CAHbLzkrXXe7r3n3jXgDKtwZhRqj=jDx9E6dLOULohnhBguvi9A@mail.gmail.com/ [3]

RFC v4 -> v5:
- Drop the RFC tag.
- No code changes.
Link: https://lore.kernel.org/all/20260718095647.182592-1-xueyuan.chen21@gmail.com/

RFC v3 -> RFC v4:
- Patch #01: Flush the direct-map range after changing it read-only, since
  the folio was cleared through writable mappings after SMP initialization
  (per Usama, thanks!).
- Patch #01: Keep the flush in the caller to preserve the
  set_direct_map_ro_noflush() contract and make the flushed range explicit.
- Patch #01: Clarify the noflush API contract and the reason stale writable
  translations must be invalidated.
Link: https://lore.kernel.org/linux-mm/20260706130440.9295-1-xueyuan.chen21@gmail.com/

RFC v2 -> RFC v3:
- Patch #01: Replace arch_make_pages_readonly() with
  set_direct_map_ro_noflush() in the existing set_direct_map* family
  (per Mike and David, thanks!).
- Patch #01: Use a direct-map address and number of pages, and document the
  direct-map-only and no-TLB-flush semantics (per David, thanks!).
- Patch #02 and #03: Update the arm64 and x86 implementations for
  set_direct_map_ro_noflush().
Link: https://lore.kernel.org/linux-mm/20260609143801.7917-1-xueyuan.chen21@gmail.com/

RFC v1 -> RFC v2:
- Patch #01: Drop the READONLY_HUGE_ZERO_FOLIO Kconfig option
  (per Dave, thanks!).
- Patch #01: Replace the huge-zero-folio-specific hook with a generic
  page-range hook (per David, thanks!).
- Patch #02 and #03: Update the arm64 and x86 implementations for the new
  hook.
Link: https://lore.kernel.org/linux-mm/20260527035607.14919-1-xueyuan.chen21@gmail.com/

Xueyuan Chen (3):
  mm: make persistent huge zero folio read-only
  arm64/mm: add set_direct_map_ro_noflush()
  x86/mm: add set_direct_map_ro_noflush()

 arch/arm64/include/asm/set_memory.h |  2 ++
 arch/arm64/mm/pageattr.c            | 10 ++++++++++
 arch/x86/include/asm/set_memory.h   |  2 ++
 arch/x86/mm/pat/set_memory.c        | 15 +++++++++++++++
 include/linux/set_memory.h          | 29 +++++++++++++++++++++++++++++
 mm/huge_memory.c                    | 16 +++++++++++++++-
 6 files changed, 73 insertions(+), 1 deletion(-)

-- 
2.47.3


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v5 1/3] mm: make persistent huge zero folio read-only
  2026-07-27 14:34 [PATCH v5 0/3] mm: make persistent huge zero folio read-only Xueyuan Chen
@ 2026-07-27 14:34 ` Xueyuan Chen
  2026-07-27 14:34 ` [PATCH v5 2/3] arm64/mm: add set_direct_map_ro_noflush() Xueyuan Chen
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 5+ messages in thread
From: Xueyuan Chen @ 2026-07-27 14:34 UTC (permalink / raw)
  To: linux-mm, akpm, david, ljs
  Cc: linux-kernel, linux-arm-kernel, catalin.marinas, will, tglx,
	mingo, bp, dave.hansen, x86, hpa, luto, peterz, lance.yang,
	usama.arif, jannh, yang, rppt, ziy, baolin.wang, liam, npache,
	ryan.roberts, dev.jain, baohua

The persistent huge zero folio is shared globally and should stay zero
after initialization. As Jann Horn pointed out[1], kernel bugs have ended
up writing to pages that were meant to be read-only, including in
security-sensitive cases. Making the persistent huge zero folio read-only
in the direct map turns such writes into faults instead of silent zero-page
corruption.

Add set_direct_map_ro_noflush() so mm code can make a direct-map range
read-only. Use an address-based signature to match ongoing direct-map
helper work[2], where existing page-based helpers may move the same way.
The helper is direct-map specific and leaves TLB invalidation to its
caller. Architectures without direct-map permission support keep existing
behavior through the generic stub.

The folio is allocated and zeroed through the writable direct map before
thp_shrinker_init() changes its permissions. thp_shrinker_init() is called
from hugepage_init(), which is registered as a subsys_initcall and runs
after SMP initialization. Stale writable kernel TLB entries may therefore
exist. Flush the direct-map range immediately after the page-table update
so they cannot bypass the read-only mapping.

Treat the direct-map permission change as best-effort. Architectures that
do not implement the helper keep the existing behavior via the generic
stub.

Inspired by Jann Horn's read-only zero page work[1] and follow-up
discussion[3] with Yang Shi.

[1] https://lore.kernel.org/linux-mm/20260508-ro-zeropage-v1-1-9808abc20b49@google.com/
[2] https://lore.kernel.org/linux-mm/0e5b23a6-4895-454a-9dfa-6dc21adc2991@kernel.org/
[3] https://lore.kernel.org/linux-mm/CAHbLzkrXXe7r3n3jXgDKtwZhRqj=jDx9E6dLOULohnhBguvi9A@mail.gmail.com/

Suggested-by: David Hildenbrand <david@kernel.org>
Suggested-by: Usama Arif <usama.arif@linux.dev>
Co-developed-by: Lance Yang <lance.yang@linux.dev>
Signed-off-by: Lance Yang <lance.yang@linux.dev>
Signed-off-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
 include/linux/set_memory.h | 29 +++++++++++++++++++++++++++++
 mm/huge_memory.c           | 16 +++++++++++++++-
 2 files changed, 44 insertions(+), 1 deletion(-)

diff --git a/include/linux/set_memory.h b/include/linux/set_memory.h
index 3030d9245f5a..e83ced6a3827 100644
--- a/include/linux/set_memory.h
+++ b/include/linux/set_memory.h
@@ -40,6 +40,24 @@ static inline int set_direct_map_valid_noflush(struct page *page,
 	return 0;
 }
 
+/**
+ * set_direct_map_ro_noflush - make a direct-map range read-only
+ * @addr: start address in the direct map
+ * @nr_pages: number of pages starting at @addr
+ *
+ * Make the direct-map range starting at @addr read-only without invalidating
+ * TLBs. Callers must either ensure that no stale writable translations can
+ * be used, or treat the permission change as a best-effort hardening step.
+ *
+ * Return: 0 on success or when direct-map permission changes are unsupported,
+ * or a negative errno on failure.
+ */
+static inline int set_direct_map_ro_noflush(const void *addr,
+					    unsigned long nr_pages)
+{
+	return 0;
+}
+
 static inline bool kernel_page_present(struct page *page)
 {
 	return true;
@@ -56,6 +74,17 @@ static inline bool can_set_direct_map(void)
 }
 #define can_set_direct_map can_set_direct_map
 #endif
+
+#ifndef set_direct_map_ro_noflush
+/* See the comment above the generic fallback for the _noflush contract. */
+static inline int set_direct_map_ro_noflush(const void *addr,
+					    unsigned long nr_pages)
+{
+	return 0;
+}
+
+#define set_direct_map_ro_noflush set_direct_map_ro_noflush
+#endif
 #endif /* CONFIG_ARCH_HAS_SET_DIRECT_MAP */
 
 #ifdef CONFIG_X86_64
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 970e077019b7..4425ae5560cb 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -40,8 +40,10 @@
 #include <linux/pgalloc.h>
 #include <linux/pgalloc_tag.h>
 #include <linux/pagewalk.h>
+#include <linux/set_memory.h>
 
 #include <asm/tlb.h>
+#include <asm/tlbflush.h>
 #include "internal.h"
 #include "swap.h"
 
@@ -932,6 +934,8 @@ static int __init thp_shrinker_init(void)
 	shrinker_register(deferred_split_shrinker);
 
 	if (IS_ENABLED(CONFIG_PERSISTENT_HUGE_ZERO_FOLIO)) {
+		unsigned long addr;
+
 		/*
 		 * Bump the reference of the huge_zero_folio and do not
 		 * initialize the shrinker.
@@ -940,8 +944,18 @@ static int __init thp_shrinker_init(void)
 		 * that get_huge_zero_folio() will most likely not fail as
 		 * thp_shrinker_init() is invoked early on during boot.
 		 */
-		if (!get_huge_zero_folio())
+		if (!get_huge_zero_folio()) {
 			pr_warn("Allocating persistent huge zero folio failed\n");
+			return 0;
+		}
+
+		addr = (unsigned long)folio_address(huge_zero_folio);
+		/*
+		 * The folio was zeroed through the writable direct map. Flush
+		 * after the page-table update to invalidate stale translations.
+		 */
+		set_direct_map_ro_noflush((void *)addr, HPAGE_PMD_NR);
+		flush_tlb_kernel_range(addr, addr + HPAGE_PMD_SIZE);
 		return 0;
 	}
 
-- 
2.47.3



^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v5 2/3] arm64/mm: add set_direct_map_ro_noflush()
  2026-07-27 14:34 [PATCH v5 0/3] mm: make persistent huge zero folio read-only Xueyuan Chen
  2026-07-27 14:34 ` [PATCH v5 1/3] " Xueyuan Chen
@ 2026-07-27 14:34 ` Xueyuan Chen
  2026-07-27 14:34 ` [PATCH v5 3/3] x86/mm: " Xueyuan Chen
  2026-07-27 18:35 ` [PATCH v5 0/3] mm: make persistent huge zero folio read-only Andrew Morton
  3 siblings, 0 replies; 5+ messages in thread
From: Xueyuan Chen @ 2026-07-27 14:34 UTC (permalink / raw)
  To: linux-mm, akpm, david, ljs
  Cc: linux-kernel, linux-arm-kernel, catalin.marinas, will, tglx,
	mingo, bp, dave.hansen, x86, hpa, luto, peterz, lance.yang,
	usama.arif, jannh, yang, rppt, ziy, baolin.wang, liam, npache,
	ryan.roberts, dev.jain, baohua

Implement set_direct_map_ro_noflush() for arm64 with update_range_prot() on
the linear map, setting PTE_RDONLY and clearing PTE_WRITE. Keep the
existing can_set_direct_map() guard and leave TLB invalidation to the
caller.

Co-developed-by: Lance Yang <lance.yang@linux.dev>
Signed-off-by: Lance Yang <lance.yang@linux.dev>
Signed-off-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
 arch/arm64/include/asm/set_memory.h |  2 ++
 arch/arm64/mm/pageattr.c            | 10 ++++++++++
 2 files changed, 12 insertions(+)

diff --git a/arch/arm64/include/asm/set_memory.h b/arch/arm64/include/asm/set_memory.h
index 90f61b17275e..7083260303c3 100644
--- a/arch/arm64/include/asm/set_memory.h
+++ b/arch/arm64/include/asm/set_memory.h
@@ -14,6 +14,8 @@ int set_memory_valid(unsigned long addr, int numpages, int enable);
 int set_direct_map_invalid_noflush(struct page *page);
 int set_direct_map_default_noflush(struct page *page);
 int set_direct_map_valid_noflush(struct page *page, unsigned nr, bool valid);
+int set_direct_map_ro_noflush(const void *addr, unsigned long nr_pages);
+#define set_direct_map_ro_noflush set_direct_map_ro_noflush
 bool kernel_page_present(struct page *page);
 
 int set_memory_encrypted(unsigned long addr, int numpages);
diff --git a/arch/arm64/mm/pageattr.c b/arch/arm64/mm/pageattr.c
index ce035e1b4eaf..c51236b61651 100644
--- a/arch/arm64/mm/pageattr.c
+++ b/arch/arm64/mm/pageattr.c
@@ -365,6 +365,16 @@ int set_direct_map_valid_noflush(struct page *page, unsigned nr, bool valid)
 	return set_memory_valid(addr, nr, valid);
 }
 
+int set_direct_map_ro_noflush(const void *addr, unsigned long nr_pages)
+{
+	if (!can_set_direct_map())
+		return 0;
+
+	return update_range_prot((unsigned long)addr, PAGE_SIZE * nr_pages,
+				 __pgprot(PTE_RDONLY),
+				 __pgprot(PTE_WRITE));
+}
+
 #ifdef CONFIG_DEBUG_PAGEALLOC
 /*
  * This is - apart from the return value - doing the same
-- 
2.47.3



^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v5 3/3] x86/mm: add set_direct_map_ro_noflush()
  2026-07-27 14:34 [PATCH v5 0/3] mm: make persistent huge zero folio read-only Xueyuan Chen
  2026-07-27 14:34 ` [PATCH v5 1/3] " Xueyuan Chen
  2026-07-27 14:34 ` [PATCH v5 2/3] arm64/mm: add set_direct_map_ro_noflush() Xueyuan Chen
@ 2026-07-27 14:34 ` Xueyuan Chen
  2026-07-27 18:35 ` [PATCH v5 0/3] mm: make persistent huge zero folio read-only Andrew Morton
  3 siblings, 0 replies; 5+ messages in thread
From: Xueyuan Chen @ 2026-07-27 14:34 UTC (permalink / raw)
  To: linux-mm, akpm, david, ljs
  Cc: linux-kernel, linux-arm-kernel, catalin.marinas, will, tglx,
	mingo, bp, dave.hansen, x86, hpa, luto, peterz, lance.yang,
	usama.arif, jannh, yang, rppt, ziy, baolin.wang, liam, npache,
	ryan.roberts, dev.jain, baohua

Implement set_direct_map_ro_noflush() for x86 using CPA directly on the
passed direct-map address. Clear _PAGE_RW and _PAGE_DIRTY, keep alias
checks disabled like the existing direct-map _noflush helpers, and leave
TLB invalidation to the caller.

Co-developed-by: Lance Yang <lance.yang@linux.dev>
Signed-off-by: Lance Yang <lance.yang@linux.dev>
Signed-off-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
 arch/x86/include/asm/set_memory.h |  2 ++
 arch/x86/mm/pat/set_memory.c      | 15 +++++++++++++++
 2 files changed, 17 insertions(+)

diff --git a/arch/x86/include/asm/set_memory.h b/arch/x86/include/asm/set_memory.h
index 4362c26aa992..bd3817e06052 100644
--- a/arch/x86/include/asm/set_memory.h
+++ b/arch/x86/include/asm/set_memory.h
@@ -89,6 +89,8 @@ int set_pages_rw(struct page *page, int numpages);
 int set_direct_map_invalid_noflush(struct page *page);
 int set_direct_map_default_noflush(struct page *page);
 int set_direct_map_valid_noflush(struct page *page, unsigned nr, bool valid);
+int set_direct_map_ro_noflush(const void *addr, unsigned long nr_pages);
+#define set_direct_map_ro_noflush set_direct_map_ro_noflush
 bool kernel_page_present(struct page *page);
 
 extern int kernel_set_to_readonly;
diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c
index d023a40a1e03..5987f4c84f6f 100644
--- a/arch/x86/mm/pat/set_memory.c
+++ b/arch/x86/mm/pat/set_memory.c
@@ -2662,6 +2662,21 @@ int set_direct_map_valid_noflush(struct page *page, unsigned nr, bool valid)
 	return __set_pages_np(page, nr);
 }
 
+int set_direct_map_ro_noflush(const void *addr, unsigned long nr_pages)
+{
+	unsigned long tempaddr = (unsigned long)addr;
+	struct cpa_data cpa = {
+		.vaddr = &tempaddr,
+		.pgd = NULL,
+		.numpages = nr_pages,
+		.mask_set = __pgprot(0),
+		.mask_clr = __pgprot(_PAGE_RW | _PAGE_DIRTY),
+		.flags = CPA_NO_CHECK_ALIAS,
+	};
+
+	return __change_page_attr_set_clr(&cpa, 1);
+}
+
 #ifdef CONFIG_DEBUG_PAGEALLOC
 void __kernel_map_pages(struct page *page, int numpages, int enable)
 {
-- 
2.47.3



^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH v5 0/3] mm: make persistent huge zero folio read-only
  2026-07-27 14:34 [PATCH v5 0/3] mm: make persistent huge zero folio read-only Xueyuan Chen
                   ` (2 preceding siblings ...)
  2026-07-27 14:34 ` [PATCH v5 3/3] x86/mm: " Xueyuan Chen
@ 2026-07-27 18:35 ` Andrew Morton
  3 siblings, 0 replies; 5+ messages in thread
From: Andrew Morton @ 2026-07-27 18:35 UTC (permalink / raw)
  To: Xueyuan Chen
  Cc: linux-mm, david, ljs, linux-kernel, linux-arm-kernel,
	catalin.marinas, will, tglx, mingo, bp, dave.hansen, x86, hpa,
	luto, peterz, lance.yang, usama.arif, jannh, yang, rppt, ziy,
	baolin.wang, liam, npache, ryan.roberts, dev.jain, baohua

On Mon, 27 Jul 2026 22:34:23 +0800 Xueyuan Chen <xueyuan.chen21@gmail.com> wrote:

> The persistent huge zero folio is shared globally and should stay zero
> after initialization. As Jann Horn pointed out[1], kernel bugs have ended
> up writing to pages that were meant to be read-only, including in
> security-sensitive cases. Making the folio read-only in the direct map
> turns such writes into faults instead of silent zero-page corruption.
> 
> This series adds set_direct_map_ro_noflush() so mm code can make a
> direct-map range read-only, then uses it for the persistent huge zero
> folio. The helper is direct-map specific, takes an address-based range as
> discussed for set_direct_map* helpers[2], and leaves TLB invalidation to
> the caller.

Thanks.  AI review asked about a few things, some pre-existing.  Includes
a possible pre-existing ARM barrier issue in arch/arm64/mm/pageattr.c
	https://sashiko.dev/#/patchset/20260727143426.1077133-1-xueyuan.chen21@gmail.com


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-07-27 18:35 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-27 14:34 [PATCH v5 0/3] mm: make persistent huge zero folio read-only Xueyuan Chen
2026-07-27 14:34 ` [PATCH v5 1/3] " Xueyuan Chen
2026-07-27 14:34 ` [PATCH v5 2/3] arm64/mm: add set_direct_map_ro_noflush() Xueyuan Chen
2026-07-27 14:34 ` [PATCH v5 3/3] x86/mm: " Xueyuan Chen
2026-07-27 18:35 ` [PATCH v5 0/3] mm: make persistent huge zero folio read-only Andrew Morton

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox