All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v2] kexec: keep the next kernel off hardware-poisoned pages
@ 2026-07-30 15:55 Breno Leitao
  2026-07-30 16:18 ` Breno Leitao
  0 siblings, 1 reply; 2+ messages in thread
From: Breno Leitao @ 2026-07-30 15:55 UTC (permalink / raw)
  To: Andrew Morton, David Hildenbrand, Lorenzo Stoakes,
	Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
	Suren Baghdasaryan, Michal Hocko, Baoquan He, Pasha Tatashin,
	Pratyush Yadav, Miaohe Lin, Naoya Horiguchi
  Cc: linux-mm, linux-kernel, kexec, rmikey, riel, kernel-team,
	Kiryl Shutsemau, Breno Leitao

Memory failures (such as unrecoverable ECCs errors) are getting more and
more common. The kernel knows how to handle it while running, marking it
as poisoned (and SIGBUS user tasks).

Poisoned memory is removed from the buddy allocator, but, not from
other places. A current problem is that kexec will load new kernel
on top of a bad/poisoned memory, which is undesirable.

If the next kernel's image, initrd or purgatory lands on poisoned frame,
the relocation copy writes to the bad memory and the machine checks
during the kexec.

Skip hardware-poisoned frames when placing segments: check them in the
kexec_file hole finder so it lays the next kernel down on good memory,
and reject a poisoned destination in sanity_check_segment_list() for
the kexec_load path, which cannot relocate.

Suggested-by: Kiryl Shutsemau <kas@kernel.org>
Signed-off-by: Breno Leitao <leitao@debian.org>
---
Changes in v2:
- EDITME: describe what is new in this series revision.
- EDITME: use bulletpoints and terse descriptions.
- Link to v1: https://patch.msgid.link/20260728-kexec_posioned-v1-1-160c81d180fe@debian.org
---
 include/linux/mm.h  |  8 ++++++++
 kernel/kexec_core.c | 10 ++++++++++
 kernel/kexec_file.c | 19 +++++++++++++++++++
 mm/memory-failure.c | 28 ++++++++++++++++++++++++++++
 4 files changed, 65 insertions(+)

diff --git a/include/linux/mm.h b/include/linux/mm.h
index 7fabe6c66b4b7..48cad9a519d08 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -5192,6 +5192,8 @@ extern const struct attribute_group memory_failure_attr_group;
 extern void memory_failure_queue(unsigned long pfn, int flags);
 void num_poisoned_pages_inc(unsigned long pfn);
 void num_poisoned_pages_sub(unsigned long pfn, long i);
+bool range_contains_hwpoison(phys_addr_t start, unsigned long size,
+			     phys_addr_t *poison);
 #else
 static inline void memory_failure_queue(unsigned long pfn, int flags)
 {
@@ -5204,6 +5206,12 @@ static inline void num_poisoned_pages_inc(unsigned long pfn)
 static inline void num_poisoned_pages_sub(unsigned long pfn, long i)
 {
 }
+
+static inline bool range_contains_hwpoison(phys_addr_t start, unsigned long size,
+					   phys_addr_t *poison)
+{
+	return false;
+}
 #endif
 
 #if defined(CONFIG_MEMORY_FAILURE) && defined(CONFIG_MEMORY_HOTPLUG)
diff --git a/kernel/kexec_core.c b/kernel/kexec_core.c
index dc770b9a6d053..9f6ed3c04299b 100644
--- a/kernel/kexec_core.c
+++ b/kernel/kexec_core.c
@@ -212,6 +212,16 @@ int sanity_check_segment_list(struct kimage *image)
 	}
 #endif
 
+	/*
+	 * Reject destinations that land on hardware-poisoned memory: the
+	 * relocation copy would machine-check on the bad frame.
+	 */
+	for (i = 0; i < nr_segments; i++) {
+		if (range_contains_hwpoison(image->segment[i].mem,
+					    image->segment[i].memsz, NULL))
+			return -EADDRNOTAVAIL;
+	}
+
 	/*
 	 * The destination addresses are searched from system RAM rather than
 	 * being allocated from the buddy allocator, so they are not guaranteed
diff --git a/kernel/kexec_file.c b/kernel/kexec_file.c
index 59fb9d71e9d86..95cc981a557be 100644
--- a/kernel/kexec_file.c
+++ b/kernel/kexec_file.c
@@ -475,6 +475,7 @@ static int locate_mem_hole_top_down(unsigned long start, unsigned long end,
 {
 	struct kimage *image = kbuf->image;
 	unsigned long temp_start, temp_end;
+	phys_addr_t poisoned_addr;
 
 	temp_end = min(end, kbuf->buf_max);
 	temp_start = temp_end - kbuf->memsz + 1;
@@ -504,6 +505,14 @@ static int locate_mem_hole_top_down(unsigned long start, unsigned long end,
 			continue;
 		}
 
+		if (range_contains_hwpoison(temp_start, temp_end - temp_start + 1,
+					    &poisoned_addr)) {
+			if (poisoned_addr < kbuf->memsz)
+				return 0;
+			temp_start = poisoned_addr - kbuf->memsz;
+			continue;
+		}
+
 		/* We found a suitable memory range */
 		break;
 	} while (1);
@@ -520,6 +529,7 @@ static int locate_mem_hole_bottom_up(unsigned long start, unsigned long end,
 {
 	struct kimage *image = kbuf->image;
 	unsigned long temp_start, temp_end;
+	phys_addr_t poisoned_addr;
 
 	temp_start = max(start, kbuf->buf_min);
 
@@ -546,6 +556,15 @@ static int locate_mem_hole_bottom_up(unsigned long start, unsigned long end,
 			continue;
 		}
 
+		/*
+		 * Avoid placing the next kernel on hardware-poisoned memory.
+		 */
+		if (range_contains_hwpoison(temp_start, temp_end - temp_start + 1,
+					    &poisoned_addr)) {
+			temp_start = poisoned_addr + PAGE_SIZE;
+			continue;
+		}
+
 		/* We found a suitable memory range */
 		break;
 	} while (1);
diff --git a/mm/memory-failure.c b/mm/memory-failure.c
index a8b03e2920ba8..04a1883d51fda 100644
--- a/mm/memory-failure.c
+++ b/mm/memory-failure.c
@@ -96,6 +96,34 @@ void num_poisoned_pages_sub(unsigned long pfn, long i)
 		memblk_nr_poison_sub(pfn, i);
 }
 
+/*
+ * Return true if any online page in [start, start + size) is hardware
+ * poisoned.  On a hit, when @poison is not NULL, @poison is set to the
+ * address of the first poisoned page, which is ugly, but I cannot only
+ * return phys_addr_t, thus this "extra" parameter, instead of returning
+ * the hit.
+ */
+bool range_contains_hwpoison(phys_addr_t start, unsigned long size,
+			     phys_addr_t *poison)
+{
+	unsigned long pfn, end_pfn;
+
+	if (!size || !atomic_long_read(&num_poisoned_pages))
+		return false;
+
+	end_pfn = PHYS_PFN(start + size - 1);
+	for (pfn = PHYS_PFN(start); pfn <= end_pfn; pfn++) {
+		struct page *page = pfn_to_online_page(pfn);
+
+		if (page && PageHWPoison(page)) {
+			if (poison)
+				*poison = PFN_PHYS(pfn);
+			return true;
+		}
+	}
+	return false;
+}
+
 /**
  * MF_ATTR_RO - Create sysfs entry for each memory failure statistics.
  * @_name: name of the file in the per NUMA sysfs directory.

---
base-commit: c5e32e86ca02b003f86e095d379b38148999293d
change-id: 20260727-kexec_posioned-72bb0a4143a0

Best regards,
--  
Breno Leitao <leitao@debian.org>



^ permalink raw reply related	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-07-30 16:18 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-30 15:55 [PATCH v2] kexec: keep the next kernel off hardware-poisoned pages Breno Leitao
2026-07-30 16:18 ` Breno Leitao

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.