From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 68A32483803 for ; Wed, 29 Jul 2026 14:14:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785334447; cv=none; b=GseTMCxzERYHoJK+45GptxpE71mdXVQ1OqHgPFj7u67GruKpcYUvK1xGGKq59V0vP3O4O0lpAmaAwZlMerFBhUg9pMRMVV9t98Q/cqpLoTMEJCtrq7vfDDt63Q1XP+L7ln4lt7hw0XUJ6Hz3T+dyIdnZcI7yvE4L4PKcMAD6X38= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785334447; c=relaxed/simple; bh=D64Cq6NjEI1sLu5olSFwnWapKOryUWAA7MrhrVxWOks=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=RDaSWEoaErQY7nkxsnuAK7UGxVpneyEP6wceOsyMG1UEKnUkLkzUp9amVw35WQuQh0FFWkdalsLKXh39jJP4T7F03V4E6XsS8pEWG3GpqP5FoBzgBg+ujocKxIYr21ZqjVARgsEb+O/RSQ7TqYbRfy3WnYcrGGS0foqD0eSyYt4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=TFN9Qiak; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="TFN9Qiak" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DDB011F000E9; Wed, 29 Jul 2026 14:13:58 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785334446; bh=1S6/CzOMnNlnmutzVvvJHpOovw2TlW70jeCaiYJeN3w=; h=Date:From:To:Cc:Subject:References:In-Reply-To; b=TFN9QiakNYMFhzGLwDJxsm/HLxZnGEl7qMcgqyKLaRbgD05yigUzTEmydkRB8Y7/a ExO6Eq7ijak8TXrN1eSFdHlyPsmOuKcXkpcJm9vj64UY9Lkxp1k62npkUOcw5xGTI+ q8bHkZwSEpxppi2ooWuWB0ucfjuvUNDtSAtxPDD6XvqQUCn2GfOo97/CY+/C5WEMwY aHulRIisdd2uJ5wStNzj2J/SZ2WbBvEPq3mYFSjiXl3dXUv5d1ANKuQC0M8NeVfoMm Zi+MT7NgiUsX31HPK0V19Vd/IcTYaDBIb1GGrIGp/getp2jGtgSQ13ny0TFY7vA2GY N+Z/x+46hJISw== Date: Wed, 29 Jul 2026 17:13:55 +0300 From: Mike Rapoport To: Peter Zijlstra Cc: Dave Hansen , linux-kernel@vger.kernel.org, Andy Lutomirski , Borislav Petkov , David Hildenbrand , Ingo Molnar , Jason Gunthorpe , Juergen Gross , Kevin Tian , Kiryl Shutsemau , "Liam R. Howlett" , Lorenzo Stoakes , Lu Baolu , "H. Peter Anvin" , Shakeel Butt , Suren Baghdasaryan , Thomas Gleixner , Toshi Kani , Vlastimil Babka , Will Deacon , linux-mm@kvack.org, x86@kernel.org Subject: Re: [PATCH 3/3] x86/mm: Fix and document DEBUG_PAGEALLOC Message-ID: References: <20260729110807.797920433@infradead.org> <20260729111119.604452135@infradead.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260729111119.604452135@infradead.org> On Wed, Jul 29, 2026 at 01:08:10PM +0200, Peter Zijlstra wrote: > It turns out that commit 5fce67641a3e ("x86/mm/pat: Don't gate > cpa_lock on debug_pagealloc_enabled()") was a little too quick to > remove the debug_pagealloc exception for cpa_lock. > > Notably __kernel_map_pages() is used by the page-allocator from any > context the page-allocator itself is used, which violates the cpa_lock > rules. > > Re-instate the exception, except make it specific to the > __kernel_map_pages() such that any other cpa() usage is still fully > serialized by cpa_lock. Also note that since cpa() should not be used > on memory that isn't allocated, the page-allocator locking and cpa are > infact mutually exclusive and all cpa usage in fully serialized. > > Add a comment explaining this and other 'funnies' surrounding > DEBUG_PAGEALLOC, including how pgd_lock is not affected and the TLB > trickery. > > Fixes: 5fce67641a3e ("x86/mm/pat: Don't gate cpa_lock on debug_pagealloc_enabled()") > Signed-off-by: Peter Zijlstra (Intel) > > + /* > + * DEBUG_PAGEALLOC is special; it is called from any context the > + * page-allocator is, which violates the normal cpa_lock locking > + * rules. > + * > + * However, since it is part of the page-allocator, things are still > + * properly serialized by the page-allocator locking and the fact that > + * when a page is owned by the page-allocator, it isn't owned by > + * anybody else. That is, you *SHOULD* not be calling cpa() on memory *SHOULD NOT* ? > + * that isn't allocated. > + * > + * Additionally, DEBUG_PAGEALLOC ensures (per probe_page_size_mask()) > + * that the kernel mapping is 4k pages, therefore there are no large > + * pages to split/collapse. > + * > + * Furthermore, the page-allocator strictly manages pages that > + * *exist*, avoiding pgd_lock. > + * > + * Therefore, it is safe to not take cpa_lock. > + */ > + if (debug_pagealloc_enabled() && (cpa->flags & CPA_DEBUG_PAGEALLOC)) > + lock = false; > + > while (rempages) { > /* > * Store the remaining nr of pages for the large page > @@ -2008,9 +2033,12 @@ static int __change_page_attr_set_clr(st > if (cpa->flags & (CPA_ARRAY | CPA_PAGES_ARRAY)) > cpa->numpages = 1; > > - spin_lock(&cpa_lock); > - ret = __change_page_attr(cpa, primary); > - spin_unlock(&cpa_lock); > + if (lock) { > + guard(spinlock)(&cpa_lock); > + ret = __change_page_attr(cpa, primary); > + } else { > + ret = __change_page_attr(cpa, primary); > + } This does make DEBUG_PAGEALLOC exception more explicit *here*, but OTOH the spin_(un)lock(&cpa_lock) in split_large_page() becomes confusing. I like my version with your comments added there more as it localizes the DEBUG_PAGEALLOC exception in the lock wrappers. > if (ret) > goto out; > > @@ -2661,15 +2689,23 @@ void __kernel_map_pages(struct page *pag > * and hence no memory allocations during large page split. > */ I'd also return early and maybe even WARN if !debug_pagealloc_enabled(). > if (enable) > - __set_pages_p(page, numpages); > + __set_pages_p(page, numpages, CPA_DEBUG_PAGEALLOC); > else > - __set_pages_np(page, numpages); > + __set_pages_np(page, numpages, CPA_DEBUG_PAGEALLOC); -- Sincerely yours, Mike.