From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f175.google.com (mail-pf1-f175.google.com [209.85.210.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BCBD63B2D0D for ; Mon, 20 Jul 2026 20:04:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.175 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784577843; cv=none; b=M0T04wtN2JrljASw5Z171XsdUbWtSXVMFtF890N+m4DS1yWjFdK2QsjLg/woORZFBeNqve6RB/LbjfSU1vHUx8ECu/X4ZQU1WfS2mko0pDWg7LSqQopmZZkfuRhj79STOP+zCI2XgdLkZ7+cufmXi3vTD2fU8fJ5AQvf7hPrerA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784577843; c=relaxed/simple; bh=S6sTPTSALV5TVsWCXeAxXPZhY5YPkKcV0So5jf/Xvl0=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=ufDA5DLbCAmZ3BpHcStFA0ahs7m7zwy6RhKPELG+abiKRSF2hRUocZRBVnjq/73iMRK51CqFawU4d61r8QeYCBkEXIzXhTr0WSkWsQvmOM36B/AHX3W9XWvIbhX7l9ZlNxLxMeEtXEQSYoeexGRkdNcMEPSfwuwqfA/vGiQ47b4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=sEaHrRGf; arc=none smtp.client-ip=209.85.210.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="sEaHrRGf" Received: by mail-pf1-f175.google.com with SMTP id d2e1a72fcca58-8487b7b3fc7so4164567b3a.0 for ; Mon, 20 Jul 2026 13:04:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784577841; x=1785182641; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=0B7Ue8sunnZ4lhynaJhVBIqW08g9nfoJaXMan4+HMrw=; b=sEaHrRGfT6Xf2Xo9APtBOma8lK2YOyEaUkrBU98PwiGaXaK+OkELDOxwveFYmXozfX hnITH854K8hGrVh96EuThTEOuH83IkqsXZjt4iP19xCyiVx3EKQNPV2Ob9ye8LhkG3/N W2FJ+dFF3OZngI68LzM4Gliec3CUoDmBiazEzd4w2jJOKTzymgIntZpBzUwfjDwNum1G L7UUFK/a5ReqUpDWr2l4X14zsI5/WJV+6hU1qfQ5HEhrlkwkt5rv6XIgDvK2DxWa37M9 HlE7o+Tf4m12kb1aeAkCV2A0pIdN2S/6dRI/XQrSDqEq4d6RrGjAiPWH0boIcSK+g4M6 fICw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784577841; x=1785182641; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=0B7Ue8sunnZ4lhynaJhVBIqW08g9nfoJaXMan4+HMrw=; b=COsjGet7PGUv1m1dC2+fdm591gPSNS+xtcwaUA9P5vRuVSdOOxrJ0aLn4GnfoXi/rg vK0wGHnzpn1byB5mVVGUWNfpHuCPKFVkahFRV/zX3XYUUwkfBgdv3TZa5hFPR9aF0clL WtXkLpibHYOxx+qt4SplH5PiZy80/EsjpTIop6p9IbdOF7n+bBPHvEKv2OjNYpnmln1N y2IwTnvTz875dJBZfAXPx+pWBbdVUMwC7OW4cKfgj+FqkmUqXvbdpVhrWV3iH7daI8z4 21v2Rd/9cnXe1aTEpkANRXxKGasLgEWm9fq7dvBthDwr3nzF4mgmrYNsTQvV4QOdxkct 2CwA== X-Forwarded-Encrypted: i=1; AHgh+Rodj68VZ5e1mLNhv/3Rt/Iyf8S+hI5YXUsUMLEybgBjjlIII/tS6aMSFVRqbOOh7qcy4IRUg1lcvNeAJQQ=@vger.kernel.org X-Gm-Message-State: AOJu0YwOL8LtAfgPQBH4meXgAE1fDGUHk7jhSQ63ZuCyRjXcu8h7/Amp 7NVw/H8l+ovft0m4g/CriEIosjo6qXD9OhWcqYRJjCXvS43eGx3Fz5lP X-Gm-Gg: AfdE7ckXH3/mAQvMTS8+kf/7qyaQrOUxB4Pr47aO3/DEzFpp5CAURfZg7ghICCLG6JL z2hFMnojb5bNarPQwH9YySfBB7c0hLxgRw3YTeOIc3S6txdUhuIvH6J5ncsSlmjlyp09kwcu6gx Qx3TaEJ9RrFfOUqBNZgDPsBCcx3nU+s2ezHrUKMXTiz+UYnhz114Zur6z+mpkVj+rLs4JxeJfCV bZUmqpStoieuCor36KQgHnElA2Uwli7lx3S3+bgC366zN1xDjv1NQkrwXvWzFB26KgA41x/jFve Rp3o6arjk9jUBQYlYIBl/NUFj4LfeFihLb3Lt/0Do8mf6Zr8lJ+5eoJJNZZmCt5/d30nojLb/nC FoV/YolZSDoYlmdUEb0T03jWlYUtwaRnMDkYMOYc+4bXEVy91N709t3oA7Z/wJ1MT4RhJ2NZJwW EcNn7WbL2Aa8DzyyiNQp28hHrySpflj1FXDA== X-Received: by 2002:a05:6a00:92aa:b0:848:3dae:66e2 with SMTP id d2e1a72fcca58-84c2933dfc2mr15708445b3a.26.1784577840929; Mon, 20 Jul 2026 13:04:00 -0700 (PDT) Received: from fedora ([2601:644:937c:6c90:6d4e:7b2d:4a39:fb0c]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84c2af9af7esm5860115b3a.52.2026.07.20.13.03.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 20 Jul 2026 13:04:00 -0700 (PDT) Date: Mon, 20 Jul 2026 13:03:57 -0700 From: Vishal Moola To: "Lorenzo Stoakes (ARM)" Cc: Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , "Mike Rapoport (Microsoft)" , Jason Gunthorpe , Lu Baolu , Andrew Morton , David Hildenbrand , linux-kernel@vger.kernel.org, linux-mm@kvack.org, Kiryl Shutsemau , iommu@lists.linux.dev, Kevin Tian , stable@vger.kernel.org Subject: Re: [PATCH] x86/mm/pat: allocate split page tables as kernel page tables Message-ID: References: <20260720-fix-cpa-kernel-pagetables-v1-1-0766e782cefe@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Mon, Jul 20, 2026 at 01:01:00PM -0700, Vishal Moola wrote: > On Mon, Jul 20, 2026 at 10:27:29AM +0100, Lorenzo Stoakes (ARM) wrote: > > When splitting a large page in CPA in __split_large_page() we allocate a > > PTE directly without going through the standard page table allocation > > routines such as pte_alloc_one_kernel(). > > > > This means the page table constructor is never called nor is the page table > > marked as a kernel page table. > > > > The former results in the folio associated with the page table not being > > marked as a page table (__pagetable_ctor() is never called thus neither is > > __folio_set_pgtable()) nor are statistics updated to reflect > > it (lruvec_stat_add_folio() is never called). > > > > The latter issue of failing to mark the page table as a kernel page > > table (ptdesc_set_kernel() is never called) is far more problematic. > > > > Since commit 5ba2f0a15564 ("mm: introduce deferred freeing for kernel page > > tables") kernel page table freeing has been batched and since the > > subsequent commit e37d5a2d60a3 ("iommu/sva: invalidate stale IOTLB entries > > for kernel address space") IOTLB cache entries for kernel page tables have > > been invalidated upon being freed. > > > > Since split page tables are freed without this invalidation, the IOTLB can > > contain stale entries for them. > > > > Resolve the issue by using the ordinary PTE allocation API at split time. > > > > This results in these kernel page tables invoking a page table constructor, > > and thus requires a page table destructor. > > > > Since we cannot assume one is always present (early allocated direct map > > page tables are not marked as such), we conditionally call > > pagetable_dtor_free() if the PG_table folio flag for the ptdesc is set, > > otherwise we free the page table via pagetable_free(). > > > > Regardless of which path is taken page tables marked as kernel page tables, > > which now includes split page tables, take the correct route through > > pagetable_free_kernel(). > > > > There is a user-visible side effect in that split page tables will appear > > in nr_page_table_pages in /proc/vmstat (as do other kernel page tables > > allocated after early boot), however this is a positive change. > > > > This issue started being markedly problematic after commit > > 5ba2f0a15564 ("mm: introduce deferred freeing for kernel page tables") so > > choose this as the Fixes target. > > > > Fixes: 5ba2f0a15564 ("mm: introduce deferred freeing for kernel page tables") > > Cc: stable@vger.kernel.org > > Signed-off-by: Lorenzo Stoakes (ARM) > > --- > > arch/x86/mm/pat/set_memory.c | 21 ++++++++++++--------- > > 1 file changed, 12 insertions(+), 9 deletions(-) > > > > diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c > > index 301fb9e77d91..a67ca33b9dd1 100644 > > --- a/arch/x86/mm/pat/set_memory.c > > +++ b/arch/x86/mm/pat/set_memory.c > > @@ -439,7 +439,11 @@ static void __cpa_collapse_large_pages(struct cpa_data *cpa) > > > > list_for_each_entry_safe(ptdesc, tmp, &pgtables, pt_list) { > > list_del(&ptdesc->pt_list); > > - pagetable_free(ptdesc); > > + > > + if (folio_test_pgtable(ptdesc_folio(ptdesc))) > > + pagetable_dtor_free(ptdesc); > > + else > > + pagetable_free(ptdesc); > > Lets not introduce more folio-ptdesc crossovers, we're trying to get > rid of them :) > > I believe pagetable_dtor_free() should do what you're looking for on its > own anyway. Actually, looking at it closer, maybe not because of the conditional portion? But that makes me think it might be better to just replace the ptdesc_clear_kernel() with pagetable_dtor() in the free function...