From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 813EA3783C4 for ; Mon, 10 Aug 2026 09:49:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786355366; cv=none; b=JLiz+oL5o+SsFwJXg5CT9eX8RTsHrMwUBJql6ds1971glYf+N73BHn46uYpom8SM0JmuoPCD4WpZpRMxp4tYaGfeP4rjbeS9S/WQT7m0jTQrLrnIPTfAJHMLdRuo/nsGoReqZb6jTAYAaksjj23nB8zFqgzHvAQ1mtb6HH7q1e4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786355366; c=relaxed/simple; bh=6Va+609HO3RYMyPS/yNDlj4dwjDSkXpdyKfndQ9lJDk=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=PHZnSlEmh7unSMTWRyGPBZ8+sxaISHZQPI5ZA0sVwgF3wVwSJKuShSfLAwcR/yXUnrEb3I0VzDYbSrzgNU5gRZvXalsMzEgVZnzToNm7Mvzig4leA6ZkKK0ccEsgofPqULTC7fLWiIijCRQXe+pBR/ERM3VjUkkbp5ZSJQxshpE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Wv/1X9ER; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Wv/1X9ER" Received: by smtp.kernel.org (Postfix) with ESMTPSA id B8B111F00A3A; Mon, 10 Aug 2026 09:49:23 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786355364; bh=+SkRVVFBnEWD5Uv1sCu0rIS7Pvrxt+IlmbSlVY+WN7Y=; h=Date:From:To:Cc:Subject:References:In-Reply-To; b=Wv/1X9ERDgtWjYDxjzZQQsH8DYErF0xSN/m1nsTerJ1Er+kroo9T/21BDk/j7AewG mb2IJDCbaMfAABiZxqGXtWe5JStNGY6g0/WzmFXrx554ub8uhd/ljk/R+OWXnj3cp/ 0tygfj2/ykqQIc+PBCczkCegW+aR7tcmzWLCT1SxdlBpBCcab2B4uHgg/Wvm3+ujcx y6Xghqh0QnekpoJg3KSLMBFUs8G2mZiWSS4VeVJvhj/hxr72sfn6l/Iyuz8A2Oi/XI X36EjV8hcVDySuD886VIh8l+yt1mrmxrYolrGduttG11EGbUV8UCoJuayodjM0Kfxw d1BGVKDTr82WQ== Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfauth.ams.internal (Postfix) with ESMTP id 295951980070; Mon, 10 Aug 2026 05:49:19 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Mon, 10 Aug 2026 05:49:22 -0400 X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFtndxpDvAd2Xc8GhLUujyEL1KRBU/KLtnlxPJZ7PwJKuizNvh0Z5nvw5UODqwllm V4MP/pzGB/+3yE3nyd/lTZP/DgxMaCzSp9LGMHqIIu5mW0q53R2kwdRWr/xnKVxR15lPSC 6b/OyYNV6nBONIIr/iqg5mahro16YCqwYKiUDNTuj1ocDqtFiwahujtpDXS8/QEl3XAZQC 56SNpjPa2MPTK7IwzS1MW+R5q4t+LW63JuusoKHkwFCrP7blr7f0yikhAdellsewTJzzVc AtFf/73lGdkx0+oyvTFwmDb2fmam7twddPl8c1I0F76DieNBdRAaxIqGNcDWsWNFhVYmk6 tFPJssgPaanMKR0g5JNUTeSZpap8AVWgDRH6m9ouaqcRUvZ5DAc4z6uBq3qwaFGaFy+j8b rDWdAj4xmKJHMDe2hc6z4PPA6xyZA9BM50A67IBlmXVlMYlZnc8T1W75xqs7t0zVhPqEmz NevDSjJ2NtfeSGWwVymgzkAj0Zq3aPHtl+TAmYFHPzP9z+lLetaHhmaC7OYQlMQL9NXzo6 5OlgE1tfBVtJmrb//XHvAY24EfEV1wlL1MrwWR4oI+8Be9WUdtOAwLoQx4SAIfo9cpncKU 8/SMJs8PW0OzgA8cvO95mlFUTED6FWBB2SZeRXr7kiWpyFH6xcJmYOp+QxEg X-ME-Proxy: Feedback-ID: i10464835:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Mon, 10 Aug 2026 05:49:17 -0400 (EDT) Date: Mon, 10 Aug 2026 10:49:16 +0100 From: Kiryl Shutsemau To: Breno Leitao Cc: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Baoquan He , Pasha Tatashin , Pratyush Yadav , Miaohe Lin , Naoya Horiguchi , linux-mm@kvack.org, linux-kernel@vger.kernel.org, kexec@lists.infradead.org, rmikey@meta.com, riel@surriel.com, kernel-team@meta.com Subject: Re: [PATCH v4] kexec: keep the next kernel off hardware-poisoned pages Message-ID: References: <20260807-kexec_posioned-v4-1-70d57f14625d@debian.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260807-kexec_posioned-v4-1-70d57f14625d@debian.org> On Fri, Aug 07, 2026 at 07:05:09AM -0700, Breno Leitao wrote: > diff --git a/kernel/kexec_core.c b/kernel/kexec_core.c > index dc770b9a6d053..e097e980b1439 100644 > --- a/kernel/kexec_core.c > +++ b/kernel/kexec_core.c > @@ -212,6 +212,16 @@ int sanity_check_segment_list(struct kimage *image) > } > #endif > > + /* > + * Reject destinations that land on hardware-poisoned memory: the > + * relocation copy would machine-check on the bad frame. > + */ > + for (i = 0; i < nr_segments; i++) { > + if (range_first_hwpoison(image->segment[i].mem, > + image->segment[i].memsz) != PHYS_ADDR_MAX) > + return -EADDRNOTAVAIL; Other -EADDRNOTAVAIL usage indicate error on user side. But this is not a user fault. Maybe -EHWPOISON instead. > + } > + > /* > * The destination addresses are searched from system RAM rather than > * being allocated from the buddy allocator, so they are not guaranteed ... > diff --git a/mm/memory-failure.c b/mm/memory-failure.c > index a8b03e2920ba8..c485e205fb633 100644 > --- a/mm/memory-failure.c > +++ b/mm/memory-failure.c > @@ -96,6 +96,78 @@ void num_poisoned_pages_sub(unsigned long pfn, long i) > memblk_nr_poison_sub(pfn, i); > } > > +/* > + * Return the first or the last hardware-poisoned online page in [start, > + * start + size), or PHYS_ADDR_MAX if the range is clean. > + */ > +static phys_addr_t range_hwpoison(phys_addr_t start, unsigned long size, > + bool first) > +{ > + phys_addr_t poison = PHYS_ADDR_MAX; > + unsigned long pfn, end_pfn; > + > + if (!size || !atomic_long_read(&num_poisoned_pages)) > + return poison; > + > + end_pfn = PHYS_PFN(start + size - 1); > + for (pfn = PHYS_PFN(start); pfn <= end_pfn; pfn++) { > + struct page *page = pfn_to_online_page(pfn); > + struct folio *folio; > + > + cond_resched(); > + > + if (!page) > + continue; > + > + folio = page_folio(page); > + if (folio_test_hugetlb(folio)) { > + /* > + * hugetlbfs is a bit special, given the poison > + * information is at the folio, not at the page > + */ > + unsigned long folio_end; > + > + /* > + * No hugetlb_lock: the scan is racy either way, a frame > + * can be poisoned right after it. Just don't let a folio > + * dissolved under us walk the scan backwards. > + */ > + folio_end = folio_pfn(folio) + folio_nr_pages(folio) - 1; > + folio_end = max(folio_end, pfn); > + > + if (folio_test_hwpoison(folio)) { > + if (first) > + return PFN_PHYS(pfn); > + poison = PFN_PHYS(min(folio_end, end_pfn)); > + } > + /* skip all the pfns that belong to hugetlb */ > + pfn = folio_end; > + continue; > + } > + > + if (!PageHWPoison(page)) > + /* page is good, let's go to the next one */ > + continue; If you don't care about re-using clean part of poisoned hugetlb folio, use is_page_hwpoison(page). This would do: if (!page || !is_page_hwpoison(page)) continue; You would spin a bit on the same folio, but shouldn't be a big deal. > + > + if (first) > + return PFN_PHYS(pfn); > + > + poison = PFN_PHYS(pfn); > + } > + > + return poison; > +} -- Kiryl Shutsemau / Kirill A. Shutemov