From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f179.google.com (mail-pl1-f179.google.com [209.85.214.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2737939C63E for ; Mon, 20 Jul 2026 20:01:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.179 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784577666; cv=none; b=JioddymV+twcosBQ5M6XEUqSWUO9EVKrEm7e1tsYu3uk1osXEi52vns6oRS09FLXvB4aFJr74/9Tq39gNH/znigRRf+Aw9V6P0Hg70I6+B3tAL4TZm5HAVP6iBIZojqmXV8wPJ9IQUohV1jPfohenAgTGLt44Ra14nEaqVSh3p4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784577666; c=relaxed/simple; bh=fh9vsor5ESw9tLp2QbVw66m9no3crF0ZgJPINwiuibM=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=ellQJmyFtTLNKWylETLzVNxDZakphlVBixPz9DJg3UK5f3ZclGkqaRFeFxCIrMwnMLZIINxkNCTPjt7lfhwasG5OJ+z/sIMSM44sJp4BX0dFj/vSpxOYNN3nVYLIXtH2p5Hpd3Yax4tzaDFk8EamUWg6pSvheE4aZh/+sLmXA00= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=fNGx2EFx; arc=none smtp.client-ip=209.85.214.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="fNGx2EFx" Received: by mail-pl1-f179.google.com with SMTP id d9443c01a7336-2ceb096e675so104437015ad.0 for ; Mon, 20 Jul 2026 13:01:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784577664; x=1785182464; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=nHyzRN69SIvrGhU2BxqAEaCwqvk6nXUkdLyvQIyO8KI=; b=fNGx2EFx5VYsI6m9bznuyz7JMhaK/Au0F5/obxysK7ICslrtdkC44k5C1k3E+B/ylS MwQc/BTnhJftBzU6NLIGD3obrLUQq3PjhLVayYMRlqCbYvL87YRb8WRGAQ4UHKN9TcOf GddUrVsZa+DMJFG0xGsDiJ3U+9x+cQs5wA+KIjySAaTUWZXmqyed0zC5DbNWnEt2tjAR n6EG0zXj5Svs7rTm8H8Rlf7zwO83h1qSxdCd2vnP6C6C4sgORf2FJwmQrPV2QmyPgSg2 zSHbW42cYZjFbWPNVE+T5X77IajV843s+C12VFHvnm3PUwIDlHtuHMlUh8c3zhMT/WCd 1Mtw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784577664; x=1785182464; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=nHyzRN69SIvrGhU2BxqAEaCwqvk6nXUkdLyvQIyO8KI=; b=ZHlUjbPzqhdAGeKvmSTya2xzApDzmCdAGilodFY90qCTk7hd/Y5VXTk8U2wPUwvb/e hqXwmKrv4BpjkUA8NXxQ27ZH3T0F2JyVztUx+Dt9Omg9m1rq8mZsArGgklNw+hCANkD3 hcdsD0ERNnl0fvCVcEKow2BTl/1Dr0mBChBScC9qz9eGhyn5gqlBEQEI20+/GgkqY4uw NJue7pIqv+U57uOIQ5mY/2Y5S3veb/eKHr0mqntJ+0GC22vqT5WH+N77dDfUvERDM33h a0VVanrBRU8ynl068kH6t/whdNDdmov/X7lEFamKkz93UUsN8tClgP2lSclunNC0EufD zc+g== X-Forwarded-Encrypted: i=1; AHgh+Rrrp3ETLeqhDZvRajk0N6NoO7z34nLayl0jc/Nddz4FLhwBhJUDVnKC8V7APSKyUK051ImHAlOBiMIbteg=@vger.kernel.org X-Gm-Message-State: AOJu0Yx7E4FRTO4Iqwh5oELHsfAxN7G3oV1KKxOJyVMN8XQ/UNqBsAu0 murh5A931VRHDwBykRcwmu8AbemJvfVWKaT6Up+9c6MpyQXhISkMwIJx X-Gm-Gg: AR+sD10VPSkKPmCYFYZTCT8S7tmyDftDH1qx39dtnx5xo+YxwoWLawAaJ9kiF61gbjj CVZFaVJrhC/9Q3CCuqRpgCNKaUnORRbwJmIzxWiN98pz+zO9uIk6SIC0BkUfJBLpHfD7x6EZRK9 HGJCzsc6YJ9x64CRYew6J1fHE5xd0chKwcohRsxgu4EWGdcVFJwfpfc0b0BSY07xL50q5GFXI4A XK+82faDIFIY37cuSZDVY2+6KDlShLWVkeBPJe1ry3b6mra43lM4/zbG0zws4whbeyF+CBrdTau 2uIVM5bDpZRlKmJAwTIi0LRE40wYRbxk4OklmoodyC/gUXR1MKYF4r1q+Prjw8p07tB9oq4fvkH RYRLgRzYDFPPYSwpPdMWulw2h8TL2c5G4GrG7ixoZwGPhxUhO/cQOTabjCErbvbpzp8LDwlh/H5 KU6RFV9ZeN3vahypoj1NvWPYPAQZMEoNZQzw== X-Received: by 2002:a17:902:eccb:b0:2c9:ff83:41fa with SMTP id d9443c01a7336-2cf3496204fmr169889025ad.24.1784577664139; Mon, 20 Jul 2026 13:01:04 -0700 (PDT) Received: from fedora ([2601:644:937c:6c90:6d4e:7b2d:4a39:fb0c]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2cf3476f70esm61611825ad.77.2026.07.20.13.01.02 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 20 Jul 2026 13:01:03 -0700 (PDT) Date: Mon, 20 Jul 2026 13:01:00 -0700 From: Vishal Moola To: "Lorenzo Stoakes (ARM)" Cc: Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , "Mike Rapoport (Microsoft)" , Jason Gunthorpe , Lu Baolu , Andrew Morton , David Hildenbrand , linux-kernel@vger.kernel.org, linux-mm@kvack.org, Kiryl Shutsemau , iommu@lists.linux.dev, Kevin Tian , stable@vger.kernel.org Subject: Re: [PATCH] x86/mm/pat: allocate split page tables as kernel page tables Message-ID: References: <20260720-fix-cpa-kernel-pagetables-v1-1-0766e782cefe@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260720-fix-cpa-kernel-pagetables-v1-1-0766e782cefe@kernel.org> On Mon, Jul 20, 2026 at 10:27:29AM +0100, Lorenzo Stoakes (ARM) wrote: > When splitting a large page in CPA in __split_large_page() we allocate a > PTE directly without going through the standard page table allocation > routines such as pte_alloc_one_kernel(). > > This means the page table constructor is never called nor is the page table > marked as a kernel page table. > > The former results in the folio associated with the page table not being > marked as a page table (__pagetable_ctor() is never called thus neither is > __folio_set_pgtable()) nor are statistics updated to reflect > it (lruvec_stat_add_folio() is never called). > > The latter issue of failing to mark the page table as a kernel page > table (ptdesc_set_kernel() is never called) is far more problematic. > > Since commit 5ba2f0a15564 ("mm: introduce deferred freeing for kernel page > tables") kernel page table freeing has been batched and since the > subsequent commit e37d5a2d60a3 ("iommu/sva: invalidate stale IOTLB entries > for kernel address space") IOTLB cache entries for kernel page tables have > been invalidated upon being freed. > > Since split page tables are freed without this invalidation, the IOTLB can > contain stale entries for them. > > Resolve the issue by using the ordinary PTE allocation API at split time. > > This results in these kernel page tables invoking a page table constructor, > and thus requires a page table destructor. > > Since we cannot assume one is always present (early allocated direct map > page tables are not marked as such), we conditionally call > pagetable_dtor_free() if the PG_table folio flag for the ptdesc is set, > otherwise we free the page table via pagetable_free(). > > Regardless of which path is taken page tables marked as kernel page tables, > which now includes split page tables, take the correct route through > pagetable_free_kernel(). > > There is a user-visible side effect in that split page tables will appear > in nr_page_table_pages in /proc/vmstat (as do other kernel page tables > allocated after early boot), however this is a positive change. > > This issue started being markedly problematic after commit > 5ba2f0a15564 ("mm: introduce deferred freeing for kernel page tables") so > choose this as the Fixes target. > > Fixes: 5ba2f0a15564 ("mm: introduce deferred freeing for kernel page tables") > Cc: stable@vger.kernel.org > Signed-off-by: Lorenzo Stoakes (ARM) > --- > arch/x86/mm/pat/set_memory.c | 21 ++++++++++++--------- > 1 file changed, 12 insertions(+), 9 deletions(-) > > diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c > index 301fb9e77d91..a67ca33b9dd1 100644 > --- a/arch/x86/mm/pat/set_memory.c > +++ b/arch/x86/mm/pat/set_memory.c > @@ -439,7 +439,11 @@ static void __cpa_collapse_large_pages(struct cpa_data *cpa) > > list_for_each_entry_safe(ptdesc, tmp, &pgtables, pt_list) { > list_del(&ptdesc->pt_list); > - pagetable_free(ptdesc); > + > + if (folio_test_pgtable(ptdesc_folio(ptdesc))) > + pagetable_dtor_free(ptdesc); > + else > + pagetable_free(ptdesc); Lets not introduce more folio-ptdesc crossovers, we're trying to get rid of them :) I believe pagetable_dtor_free() should do what you're looking for on its own anyway.