From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f53.google.com (mail-pj1-f53.google.com [209.85.216.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3CF3C3E0C55 for ; Tue, 21 Jul 2026 17:06:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.53 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784653569; cv=none; b=kgZUsz2/DHuYnemS2CYHuPfDmny1pnYr68lrlspzCr+QO1vlZsOYSGsJ2wHYhRGWO8HogvgFgllaJXA0arxeAhNmds0k98kd4ySBPqprcG4sZ50Bza7Isj7eXop+Oq5axL5yji899SVsLSqKeoJs//yS7Pa/RhkmOJIyWMBzMUI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784653569; c=relaxed/simple; bh=4IrHsiAS5RlY/yZO5MU8rItGwplYzpm+UIMFHUYdnGI=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Z1YVmHycVU6B42UPDlLBWhZxA7WD3YiZTY9XifQqoViUeuP4XKrh1edw7ca3FR0nWsKBWtKrIgsrSNgX4TgvI7vO1R/5A3Oy4LtF8KY4AaqMn5fVL4eneh9M/OL4sT+edb4Lxl2qUOO0OOd/xpq0zorcW2KBMDnmzDmkERI8v44= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=d9FK3e8F; arc=none smtp.client-ip=209.85.216.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="d9FK3e8F" Received: by mail-pj1-f53.google.com with SMTP id 98e67ed59e1d1-3811f512167so11250371a91.3 for ; Tue, 21 Jul 2026 10:06:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784653567; x=1785258367; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=UaZs0dUksz9lWWombvtYA7KfGwaIYElA59FlH0y/uOc=; b=d9FK3e8FyUAgFPXiYlNefAbGWFW1kC9ZsLQnIwzOjVewPtv32w647EgevRHL//w36o FEO24GZsOahUTAud+At9XUXnDxvq01EBLtrGfq7pPQK1GuG6pnyyOOgMR0Cjb6SPCBhy 3m3cUMQh7F7DULAt7hbPNJ/wx0dVktEj+vIiLYT6rkGqLuAk/vmZ3SV8wuCnS9hNtDm5 S4sggOzTL05KTrGVV9FOmPI1n0bkyr/La/a5uBPKyIyKGmUn93q31QLlH6a4SSqsSsWm vCtKyGcx1w/Wwg3YmQa0kQ0MAdUavj/2Kb3Kkuh54lXvLrxG8qVS9RnUTNJ6giZzX1Rs EcBQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784653567; x=1785258367; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=UaZs0dUksz9lWWombvtYA7KfGwaIYElA59FlH0y/uOc=; b=bK1rcaqEApISt9D2HoBVDF8DiWjoQHnCJZbB0vDSrZ/H0yd6s2yE85CDdNOT4ewZeD 4Mp9P4YYDaj5e6J8PYSJPiAK3tz/2xmxdijfj9156/hQeo0zBrkFHpeXwEpeOOUXw9aS /uaX1R7t/SXNOfwrGEj+Y/v6C2agXUnEErcgopFyVmLM/meoEU5iaapcvHD1WsTPbJQj CPlxgnA8KnGxpMzvtyr0BRjOTBz/AAa3z243O3G7/zYm34wYLsScd5rnifmFmLQfhdaS ZAfVbhU+A2SbsNS57eoUAI+pNDwm6yrH9+O0CQQY9PeoHzbJTzIrsqh3QLi3HtN90Lm0 G99A== X-Forwarded-Encrypted: i=1; AHgh+Rrs1/+yNXx7xPfi7js1ype4g6/xleGD7SxRSklhHF1IaixAyujuKjanhU37dwd/6DfFwut6Kgm8YbRHz5g=@vger.kernel.org X-Gm-Message-State: AOJu0Yz4eSO2lR+VFd0h6s3WAubvP1J8qubQwKYvj5qdxvTJOnxEL2DF nUziI7Va39gBSj61ZlexZC4OWvq7cWFBPqy4v7IGQWsAA3DxZhPgJWdy X-Gm-Gg: AR+sD12sxSEmf/MsyV2xR/IBmvsm4oyPieh2kddOyC1+YXWD5fqeOP/V0EGHFASW+tC X4kR6vR60AMHSjBFMwTmrI0mtFwZDkyvmIFQ0xZXJ5JnKFEERRb225H4y06b0aGzoYCNdzOrrlQ SMS+sqQiz6N7nNdzkM/0DCkctSZZ9ZQI06tF5kxmcXW4hMGu/kOwZDKofkP3nMoim6Pw4XcST6B PpulqLhx9jwx6qgUWGHjWgjz09vcTJKSxBn2WrTh9NNOrBYfRkyGUr20fA2m33lL4mZ0/AAYDT7 wE+GQD2hG+lLe3FWi9feRuk2CIVfL6OMP2/QO9zP62rsWvydP1vbrfeN2IR3BUJ0hTaPcCkorhL b/Juy+wT+eH1KPxwUnh56BFFa/KSXhyGh2gy6V10EnlpFOvgutYD+xpuHbrFor1rQAI5UBRp+Ij 2rmqf5UVRCwKs5ZjXtP6ptKKoNSn8+tAB0oQ== X-Received: by 2002:a17:90b:17cc:b0:381:a766:efcb with SMTP id 98e67ed59e1d1-38e4b3dc089mr20506488a91.4.1784653567393; Tue, 21 Jul 2026 10:06:07 -0700 (PDT) Received: from fedora ([2601:644:937c:6c90:6d4e:7b2d:4a39:fb0c]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-13ce2cbb6e8sm39508170c88.9.2026.07.21.10.06.05 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 Jul 2026 10:06:06 -0700 (PDT) Date: Tue, 21 Jul 2026 10:06:03 -0700 From: Vishal Moola To: Mike Rapoport Cc: "Lorenzo Stoakes (ARM)" , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Jason Gunthorpe , Lu Baolu , Andrew Morton , David Hildenbrand , linux-kernel@vger.kernel.org, linux-mm@kvack.org, Kiryl Shutsemau , iommu@lists.linux.dev, Kevin Tian , stable@vger.kernel.org Subject: Re: [PATCH] x86/mm/pat: allocate split page tables as kernel page tables Message-ID: References: <20260720-fix-cpa-kernel-pagetables-v1-1-0766e782cefe@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Tue, Jul 21, 2026 at 04:51:31PM +0300, Mike Rapoport wrote: > On Tue, Jul 21, 2026 at 01:09:27PM +0100, Lorenzo Stoakes (ARM) wrote: > > On Tue, Jul 21, 2026 at 01:00:53PM +0100, Lorenzo Stoakes (ARM) wrote: > > > On Tue, Jul 21, 2026 at 01:32:44PM +0300, Mike Rapoport wrote: > > > > On Tue, Jul 21, 2026 at 10:58:50AM +0100, Lorenzo Stoakes (ARM) wrote: > > > > > On Tue, Jul 21, 2026 at 02:45:43AM -0700, Vishal Moola wrote: > > > > > > > > > > > > > > Well some kernel page tables are still allocated without ctor (early allocated > > > > > > > direct map for isntance), and if you did pagetable_dtor_free() it > > > > > > > unconditionally calls pagetable_dtor(). > > > > > > > > TBH, I cannot think of a scenario when page tables allocated at boot would > > > > be collapsed. But surely, checking the page type is safer just in case. > > > > > > Yeah nor can to be honest, anything that could be made large in the direct map > > > would already be large right? > > > > > > But it's 'just in case' somebody did something dumb :) Later can maybe make it a > > > WARN_ON(). But just to fix the proximate issue for now. > > > > > > > > > > > > > > The ptlock_free() and __folio_clear_pgtable() there would be harmelss (no locks > > > > > > > assigned for kernel page table, and if PG_table never set clearing it is a noop) > > > > > > > but the lruvec_stat_sub_folio() would cause an unbalanced decrement of > > > > > > > nr_page_table_pages. > > > > > > > > > > > > Gotcha, thanks for the explanation :) > > > > > > > > > > No worries, this is subtle stuff with lots of weird gotchas and stuff we need to > > > > > improve... I seem to have fallen down an unexpected rabbit hole with these fixes > > > > > :) > > > > > > > > > > > > > > > > > > It sucks, but until everything is updated to call the ctor we have to do it this > > > > > > > way :>) > > > > > > > > > > > > Yeah that makes sense. Although I'd rather see the condition as: > > > > > > if(PageTable(ptdesc_page(...))) > > > > > > > > > > > > We really shouldn't be calling ptdesc_folio() anywhere anymore. > > > > > > > > > > I think better for a follow up since the code already uses ptdesc all over the > > > > > place (fundamental to the approach really, keeping a list of page tables etc.) > > > > > and this is a fix that needs backporting. > > > > > > > > I agree with Vishal that it's better to use page type rather than folio > > > > type. And it's the same for backporting ;-) > > > > > > Ah sorry misunderstood, you mean straight up PageTable(ptdesc_page()), I thought > > > Vishal was saying we shouldn't be directly referencing ptdesc's at all (which > > > would be the rework). > > > > > > I guess definitionally page tables are never folios. I lazily went with what I > > > saw elsewhere, my bad :) > > > > Ah yeah I remember now, i saw __pagetable_ctor() dealt with folios: > > > > static inline void __pagetable_ctor(struct ptdesc *ptdesc) > > { > > struct folio *folio = ptdesc_folio(ptdesc); > > > > __folio_set_pgtable(folio); > > lruvec_stat_add_folio(folio, NR_PAGETABLE); > > } > > > > > > And was like 'huh?' (surely definitionally they're _not_ folios) but went with > > that on that basis. > > I wonder why setting the type is even in ctor rather than in allocation. Yeah, you're not the only one wondering that ;) Kevin is actively looking at moving those to the allocation/free site instead[1]! > > Another place to clean up I guess? (that one _definitely_ is a follow up though > > ;) > > Yep :) Yup. It makes most sense to clean those up when we add a per-memdesc api (i.e. for memcg in this case). I haven't really had the time to work on those though :/ > > Cheers, Lorenzo > > -- > Sincerely yours, > Mike. [1] https://lore.kernel.org/linux-mm/20260714-remove_pgtable_cdtor-v1-14-44be8a7685d7@arm.com/