From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 83DD4C61DD3 for ; Thu, 3 Sep 2026 16:48:13 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:Content-Transfer-Encoding: Content-Type:In-Reply-To:From:References:Cc:To:Subject:MIME-Version:Date: Message-ID:Reply-To:Content-ID:Content-Description:Resent-Date:Resent-From: Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=ZpIyHBMuPpbjAnx97BPXthrcY+hpLbMXJmomqpqPlsc=; b=QIvPu/J/ESNnAKSga8Fx0AJSIF 8W5KrcrQ85kx4QeJkO6G6ipAIimGVNDgLwU35gloPGKuKLY8yroL6dY9drt3jRiMsdAsHXiFus3Mr RWHULujZQENDrhicRvhEM6XnMIk3X1iX5KEizn3eWX2NkYuutxWBP4yUe7mJOvDKJUtCGsERA9gIj Fk1c5fH8QHnTFk6WurPGJbILqNNp0vvlniH8dHiKS5IsoujrftWATkMHEMaTQIDeSa/2Tpwq3fNgQ LnJwsq/dY6jbTTWoBS9+WzC/Rh8pDZoRy+pbDTV8MpJ9iYPhOrcWUjuhB1Fs9glSeC+6A5Ac4Ylse WXbPb7cQ==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1x2Abd-00000000D7J-2u00; Thu, 03 Sep 2026 16:48:06 +0000 Received: from foss.arm.com ([217.140.110.172]) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1x2Abb-00000000D6Z-1ZGm for linux-arm-kernel@lists.infradead.org; Thu, 03 Sep 2026 16:48:04 +0000 Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id BC3F71596; Thu, 3 Sep 2026 09:47:57 -0700 (PDT) Received: from [10.57.6.2] (unknown [10.57.6.2]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id 4A3283F673; Thu, 3 Sep 2026 09:47:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=simple/simple; d=arm.com; s=foss; t=1788454081; bh=FmovFpZo43JwPnG5o8GR9kxo0GqzPlOEUnziiqr2YBY=; h=Date:Subject:To:Cc:References:From:In-Reply-To:From; b=J5m3OkD+YOZz0+n5qB7t/n+c86nO4npM4vtmrEMzr87cSyilKLjH5mbi7+22sJcZA zlxUnGZtCYsyqCIxOiJ1WsTc9Ygg9uVHYIAgX66ev7T2DUrtAWAfRswWZoClkfVYzR XJb1wBn1DvE4OnD5l/iDI4yMICViiKvNhb+ETGec= Message-ID: <43382882-0e7b-4255-9c34-99cec5b0bbbc@arm.com> Date: Thu, 3 Sep 2026 18:47:50 +0200 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH RFC v9 00/25] pkeys-based page table hardening To: Linu Cherian Cc: linux-hardening@vger.kernel.org, Andrew Morton , Andy Lutomirski , Catalin Marinas , Dave Hansen , "David Hildenbrand (Arm)" , Jann Horn , Jeff Xu , Joey Gouly , Kees Cook , Linus Walleij , Marc Zyngier , Mark Brown , Matthew Wilcox , Maxwell Bland , "Mike Rapoport (IBM)" , Peter Zijlstra , Pierre Langlois , =?UTF-8?Q?Pierre-Cl=C3=A9ment_Tosi?= , Quentin Perret , Rick Edgecombe , Ryan Roberts , Vlastimil Babka , Will Deacon , Yang Shi , Yeoreum Yun , linux-arm-kernel@lists.infradead.org, linux-mm@kvack.org, x86@kernel.org, Ira Weiny , Lorenzo Stoakes , Thomas Gleixner References: <20260818-kpkeys-v9-0-743ad31b2c8f@arm.com> From: Kevin Brodsky Content-Language: en-GB In-Reply-To: Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.9.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20260903_094803_502953_A66970C0 X-CRM114-Status: GOOD ( 42.04 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org On 01/09/2026 16:24, Linu Cherian wrote: > Hi Kevin, > > On Tue, Aug 18, 2026 at 03:08:42PM +0100, Kevin Brodsky wrote: >> [Sending during the merge window in case reviewers have spare >> cycles; I'm not aiming to have this series merged in v7.3.] >> >> This is a proposal to leverage protection keys (pkeys) to harden >> critical kernel data, by making it mostly read-only. The series includes >> a simple framework called "kpkeys" to manipulate pkeys for in-kernel use, >> as well as a page table hardening feature based on that framework, >> "kpkeys_hardened_pgtables". Both are implemented on arm64 as a proof of >> concept, but they are designed to be compatible with any architecture >> that supports pkeys. >> >> The proposed approach is a typical use of pkeys: the data to protect is >> mapped with a given pkey P, and the pkey register is initially >> configured to grant read-only access to P. Where the protected data >> needs to be written to, the pkey register is temporarily switched to >> grant write access to P on the current CPU. >> >> The key fact this approach relies on is that the target data is >> only written to via a limited and well-defined API. This makes it >> possible to explicitly switch the pkey register where needed, without >> introducing excessively invasive changes, and only for a small amount of >> trusted code. >> >> Page tables are chosen as an initial target because of their especially >> critical nature - a single write may result in arbitrary pages becoming >> accessible to any context (including userspace). In order to keep the >> series digestible for reviewers, this version focuses on functionality >> rather than performance, making it most suitable as a debug feature. The >> key trade-off is the requirement to PTE-map the linear map - see section >> "Protected page table allocation" for details. >> >> This series has similarities with the "PKS write protected page tables" >> series posted by Rick Edgecombe a few years ago [1] but it is not >> specific to x86/PKS - the approach is meant to be generic. >> >> This proposal (as of RFC v5) was presented at Linux Security Summit >> Europe 2025 [2]. >> >> [Table of contents] >> >> * kpkeys >> - pkey register management >> >> * kpkeys_hardened_pgtables >> - Protected page table allocation >> - kpkeys context switching >> - Performance >> - Limitations >> >> * This series >> - Branches >> >> * Threat model >> >> * Further use-cases >> >> * Open questions >> >> kpkeys >> ====== >> >> The use of pkeys involves two separate mechanisms: assigning a pkey to >> pages, and defining the pkeys -> permissions mapping via the pkey >> register. This is implemented through the following interface: >> >> - Pages are assigned a pkey in the linear map using set_memory_pkey(). >> This is sufficient for this series, but it is also plausible for >> higher-level allocators to support marking allocations with a given >> pkey. >> >> - The pkey register is configured based on a *kpkeys context*. kpkeys >> contexts are represented as simple integers that correspond to a given >> configuration, for instance: >> >> KPKEYS_CTX_DEFAULT: >> RW access to KPKEYS_PKEY_DEFAULT >> RO access to any other KPKEYS_PKEY_* >> >> KPKEYS_CTX_: >> RW access to KPKEYS_PKEY_DEFAULT >> RW access to KPKEYS_PKEY_ >> RO access to any other KPKEYS_PKEY_* >> >> Only pkeys that are managed by the kpkeys framework are impacted; >> permissions for other pkeys are left unchanged (this allows for other >> schemes using pkeys to be used in parallel, and arch-specific use of >> certain pkeys). > > - Adding some basic details on what a scheme and context is quite helpful. > > - Giving some hints (may be an example) on how multiple schemes and multiple contexts > play together would be quite helpful. "scheme" doesn't mean anything precise, it's only the notion that pkeys that aren't reserved for kpkeys (i.e. anything but 0 or 1 in this series) may be used for other purposes. Happy to reword if you have a suggestion. "kpkeys context" is what is described above this paragraph, it's really just a set of permissions for the managed pkeys. Transitioning between context is described below. > Adding a documentation that covers these aspects would be much > appreciated. For sure, I am planning to have a documentation patch in a subsequent version. > My understanding is that pkeys are being partitioned across different > contexts. But then the introduction of the term "scheme" looks bit confusing to me. I wouldn't say pkeys are partitioned across contexts. Every context has a set of permissions for all the pkeys managed by kpkeys. Any other pkey is ignored (permissions left unchanged) by this framework. >> The current kpkeys context is changed by calling >> kpkeys_enter_context(), which will set the pkey register >> accordingly and return the original state. A >> subsequent call to kpkeys_leave_context() restores the original >> state (and thus the original kpkeys context). The numeric value of >> KPKEYS_CTX_* (kpkeys context) is purely symbolic and thus generic, >> however each architecture is free to define non-default pkeys >> values (KPKEYS_PKEY_*). >> > ..snip > >> Open questions >> ============== >> >> A few aspects in this RFC that are debatable and/or worth discussing: >> >> - There is currently no restriction on how kpkeys contexts map to pkeys >> permissions. A typical approach is to allocate one pkey per context and >> make it writable in that context only. As the number of contexts > Probably to avoid the assumption, may be we can we have something like > below > > For a pkey P, we could define > PKEY_P_PERM_CTXT_OTHERS //permission for pkey p in other contexts > PKEY_P_PERM_CTXT_SELF //permission for pkey p in self context > > With the assumption of one pkey mapped for every context, > the permission for the default context would look something like, > > PKEY_DEF_PERM_CTXT_SELF << PKEY_DEF_PKEY_SHIFT | > PKEY_CT0_PERM_CTXT_OTHERS << PKEY_CT0_PKEY_SHIFT | > PKEY_CT1_PERM_CTXT_OTHERS << PKEY_CT1_PKEY_SHIFT | > ...(for all valid contexts) > > where, > Permission key, PKEY_DEF is associated with context DEFAULT, > Permission key, PKEY_CT0 is associated with context CT0, > Permission key, PKEY_CT1 is associated with context CT1 This adds assumptions rather than avoiding them. *Typically* when adding a context you'd allocate a pkey that's only writable by this context, but it doesn't have to be this way. The configuration space is more easily understood by considering the other use-cases we've investigated (struct cred protection and eBPF isolation, linked further down). For instance, for cred protection, we had KPKEYS_LVL_UNRESTRICTED with write access to all pkeys, and for eBPF isolation, we need a level that is less privileged and therefore does *not* have write access to pkey 0. >> increases, we may however run out of pkeys, especially on arm64 (just >> 8 pkeys with POE). Depending on the use-cases, it may be acceptable to >> use the same pkey for the data associated to multiple contexts. > Lets say two contexts A and B, use the same pkey P as their permission matches. > But then, when we enter context A, permission for pkey P gets > relaxed, then that would relax permission for pages associated with > context B as well which is unintended ? That may be exactly what is intended, it all depends on the use-case. C1 may have a private pkey P1, and C2 P2, and then P3 that is shared by C1 and C2 (writable by both)   > As the hardware supports 16 pkeys, should we consider removing the limit > of 8 pkeys so that we can have unique pkeys for each context ? FEAT_S1POE only supports 4-bit pkeys when using 128-bit page tables. - Kevin