Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Kevin Brodsky <kevin.brodsky@arm.com>
To: Linu Cherian <linu.cherian@arm.com>
Cc: linux-hardening@vger.kernel.org,
	"Andrew Morton" <akpm@linux-foundation.org>,
	"Andy Lutomirski" <luto@kernel.org>,
	"Catalin Marinas" <catalin.marinas@arm.com>,
	"Dave Hansen" <dave.hansen@linux.intel.com>,
	"David Hildenbrand (Arm)" <david@kernel.org>,
	"Jann Horn" <jannh@google.com>, "Jeff Xu" <jeffxu@chromium.org>,
	"Joey Gouly" <joey.gouly@arm.com>, "Kees Cook" <kees@kernel.org>,
	"Linus Walleij" <linusw@kernel.org>,
	"Marc Zyngier" <maz@kernel.org>,
	"Mark Brown" <broonie@kernel.org>,
	"Matthew Wilcox" <willy@infradead.org>,
	"Maxwell Bland" <mbland@motorola.com>,
	"Mike Rapoport (IBM)" <rppt@kernel.org>,
	"Peter Zijlstra" <peterz@infradead.org>,
	"Pierre Langlois" <pierre.langlois@arm.com>,
	"Pierre-Clément Tosi" <ptosi@google.com>,
	"Quentin Perret" <qperret@google.com>,
	"Rick Edgecombe" <rick.p.edgecombe@intel.com>,
	"Ryan Roberts" <ryan.roberts@arm.com>,
	"Vlastimil Babka" <vbabka@kernel.org>,
	"Will Deacon" <will@kernel.org>,
	"Yang Shi" <yang@os.amperecomputing.com>,
	"Yeoreum Yun" <yeoreum.yun@arm.com>,
	linux-arm-kernel@lists.infradead.org, linux-mm@kvack.org,
	x86@kernel.org, "Ira Weiny" <iweiny@kernel.org>,
	"Lorenzo Stoakes" <ljs@kernel.org>,
	"Thomas Gleixner" <tglx@kernel.org>
Subject: Re: [PATCH RFC v9 00/25] pkeys-based page table hardening
Date: Thu, 3 Sep 2026 18:47:50 +0200	[thread overview]
Message-ID: <43382882-0e7b-4255-9c34-99cec5b0bbbc@arm.com> (raw)
In-Reply-To: <apbgDNjQPLBdsUe0@a079125.arm.com>

On 01/09/2026 16:24, Linu Cherian wrote:
> Hi Kevin,
>
> On Tue, Aug 18, 2026 at 03:08:42PM +0100, Kevin Brodsky wrote:
>> [Sending during the merge window in case reviewers have spare
>> cycles; I'm not aiming to have this series merged in v7.3.]
>>
>> This is a proposal to leverage protection keys (pkeys) to harden
>> critical kernel data, by making it mostly read-only. The series includes
>> a simple framework called "kpkeys" to manipulate pkeys for in-kernel use,
>> as well as a page table hardening feature based on that framework,
>> "kpkeys_hardened_pgtables". Both are implemented on arm64 as a proof of
>> concept, but they are designed to be compatible with any architecture
>> that supports pkeys.
>>
>> The proposed approach is a typical use of pkeys: the data to protect is
>> mapped with a given pkey P, and the pkey register is initially
>> configured to grant read-only access to P. Where the protected data
>> needs to be written to, the pkey register is temporarily switched to
>> grant write access to P on the current CPU.
>>
>> The key fact this approach relies on is that the target data is
>> only written to via a limited and well-defined API. This makes it
>> possible to explicitly switch the pkey register where needed, without
>> introducing excessively invasive changes, and only for a small amount of
>> trusted code.
>>
>> Page tables are chosen as an initial target because of their especially
>> critical nature - a single write may result in arbitrary pages becoming
>> accessible to any context (including userspace). In order to keep the
>> series digestible for reviewers, this version focuses on functionality
>> rather than performance, making it most suitable as a debug feature. The
>> key trade-off is the requirement to PTE-map the linear map - see section
>> "Protected page table allocation" for details.
>>
>> This series has similarities with the "PKS write protected page tables"
>> series posted by Rick Edgecombe a few years ago [1] but it is not
>> specific to x86/PKS - the approach is meant to be generic.
>>
>> This proposal (as of RFC v5) was presented at Linux Security Summit
>> Europe 2025 [2].
>>
>> [Table of contents]
>>
>> * kpkeys
>>   - pkey register management
>>
>> * kpkeys_hardened_pgtables
>>   - Protected page table allocation
>>   - kpkeys context switching
>>   - Performance
>>   - Limitations
>>
>> * This series
>>   - Branches
>>
>> * Threat model
>>
>> * Further use-cases
>>
>> * Open questions
>>
>> kpkeys
>> ======
>>
>> The use of pkeys involves two separate mechanisms: assigning a pkey to
>> pages, and defining the pkeys -> permissions mapping via the pkey
>> register. This is implemented through the following interface:
>>
>> - Pages are assigned a pkey in the linear map using set_memory_pkey().
>>   This is sufficient for this series, but it is also plausible for
>>   higher-level allocators to support marking allocations with a given
>>   pkey.
>>
>> - The pkey register is configured based on a *kpkeys context*. kpkeys
>>   contexts are represented as simple integers that correspond to a given
>>   configuration, for instance:
>>
>>   KPKEYS_CTX_DEFAULT:
>>         RW access to KPKEYS_PKEY_DEFAULT
>>         RO access to any other KPKEYS_PKEY_*
>>
>>   KPKEYS_CTX_<FEAT>:
>>         RW access to KPKEYS_PKEY_DEFAULT
>>         RW access to KPKEYS_PKEY_<FEAT>
>>         RO access to any other KPKEYS_PKEY_*
>>
>>   Only pkeys that are managed by the kpkeys framework are impacted;
>>   permissions for other pkeys are left unchanged (this allows for other
>>   schemes using pkeys to be used in parallel, and arch-specific use of
>>   certain pkeys).
>
> - Adding some basic details on what a scheme and context is quite helpful.
>
>  - Giving some hints (may be an example) on how multiple schemes and multiple contexts
>   play together would be quite helpful.

"scheme" doesn't mean anything precise, it's only the notion that pkeys
that aren't reserved for kpkeys (i.e. anything but 0 or 1 in this
series) may be used for other purposes. Happy to reword if you have a
suggestion.

"kpkeys context" is what is described above this paragraph, it's really
just a set of permissions for the managed pkeys. Transitioning between
context is described below.

> Adding a documentation that covers these aspects would be much
> appreciated.

For sure, I am planning to have a documentation patch in a subsequent
version.

> My understanding is that pkeys are being partitioned across different
> contexts. But then the introduction of the term "scheme" looks bit confusing to me.

I wouldn't say pkeys are partitioned across contexts. Every context has
a set of permissions for all the pkeys managed by kpkeys. Any other pkey
is ignored (permissions left unchanged) by this framework.

>>   The current kpkeys context is changed by calling
>>   kpkeys_enter_context(), which will set the pkey register
>>   accordingly and return the original state. A
>>   subsequent call to kpkeys_leave_context() restores the original
>>   state (and thus the original kpkeys context). The numeric value of
>>   KPKEYS_CTX_* (kpkeys context) is purely symbolic and thus generic,
>>   however each architecture is free to define non-default pkeys
>>   values (KPKEYS_PKEY_*).
>>
> ..snip
>
>> Open questions
>> ==============
>>
>> A few aspects in this RFC that are debatable and/or worth discussing:
>>
>> - There is currently no restriction on how kpkeys contexts map to pkeys
>>   permissions. A typical approach is to allocate one pkey per context and
>>   make it writable in that context only. As the number of contexts
> Probably to avoid the assumption, may be we can we have something like
> below 
>
> For a pkey P, we could define
> PKEY_P_PERM_CTXT_OTHERS	 //permission for pkey p in other contexts
> PKEY_P_PERM_CTXT_SELF	 //permission for pkey p in self context
>
> With the assumption of one pkey mapped for every context,
> the permission for the default context would look something like,
>
> PKEY_DEF_PERM_CTXT_SELF << PKEY_DEF_PKEY_SHIFT |
> PKEY_CT0_PERM_CTXT_OTHERS << PKEY_CT0_PKEY_SHIFT | 
> PKEY_CT1_PERM_CTXT_OTHERS << PKEY_CT1_PKEY_SHIFT |
> ...(for all valid contexts)
>
> where,
> Permission key, PKEY_DEF is associated with context DEFAULT,
> Permission key, PKEY_CT0 is associated with context CT0,
> Permission key, PKEY_CT1 is associated with context CT1

This adds assumptions rather than avoiding them. *Typically* when adding
a context you'd allocate a pkey that's only writable by this context,
but it doesn't have to be this way.

The configuration space is more easily understood by considering the
other use-cases we've investigated (struct cred protection and eBPF
isolation, linked further down). For instance, for cred protection, we
had KPKEYS_LVL_UNRESTRICTED with write access to all pkeys, and for eBPF
isolation, we need a level that is less privileged and therefore does
*not* have write access to pkey 0.

>>   increases, we may however run out of pkeys, especially on arm64 (just
>>   8 pkeys with POE). Depending on the use-cases, it may be acceptable to
>>   use the same pkey for the data associated to multiple contexts.
> Lets say two contexts A and B, use the same pkey P as their permission matches.
> But then, when we enter context A, permission for pkey P gets
> relaxed, then that would relax permission for pages associated with
> context B as well which is unintended ?

That may be exactly what is intended, it all depends on the use-case. C1
may have a private pkey P1, and C2 P2, and then P3 that is shared by C1
and C2 (writable by both)
 
> As the hardware supports 16 pkeys, should we consider removing the limit
> of 8 pkeys so that we can have unique pkeys for each context ?

FEAT_S1POE only supports 4-bit pkeys when using 128-bit page tables.

- Kevin


  reply	other threads:[~2026-09-03 16:48 UTC|newest]

Thread overview: 53+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-18 14:08 [PATCH RFC v9 00/25] pkeys-based page table hardening Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 01/25] mm: Introduce kpkeys Kevin Brodsky
2026-08-27 18:00   ` David Hildenbrand (Arm)
2026-08-31 15:25     ` Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 02/25] set_memory: Introduce set_memory_pkey() stub Kevin Brodsky
2026-09-01 14:33   ` Linu Cherian
2026-09-03 16:41     ` Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 03/25] arm64: mm: Enable overlays for all EL1 indirect permissions Kevin Brodsky
2026-09-01 14:39   ` Linu Cherian
2026-09-03 16:41     ` Kevin Brodsky
2026-09-01 14:41   ` Linu Cherian
2026-09-03 16:43     ` Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 04/25] arm64: Introduce por_elx_set_pkey_perms() helper Kevin Brodsky
2026-09-01 14:42   ` Linu Cherian
2026-09-03 16:43     ` Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 05/25] arm64: Implement asm/kpkeys.h using POE Kevin Brodsky
2026-09-01 14:48   ` Linu Cherian
2026-09-03 16:44     ` Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 06/25] arm64: set_memory: Implement set_memory_pkey() Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 07/25] arm64: Context-switch POR_EL1 Kevin Brodsky
2026-09-03  5:58   ` Linu Cherian
2026-08-18 14:08 ` [PATCH RFC v9 08/25] arm64: Initialize POR_EL1 register on cpu_resume() Kevin Brodsky
2026-09-03  5:56   ` Linu Cherian
2026-09-03 16:45     ` Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 09/25] arm64: Enable kpkeys Kevin Brodsky
2026-09-03  9:14   ` Linu Cherian
2026-08-18 14:08 ` [PATCH RFC v9 10/25] memblock: Move INIT_MEMBLOCK_* macros to header Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 11/25] mm: kpkeys: Introduce kpkeys_hardened_pgtables feature Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 12/25] mm: kpkeys: Protect regular page tables Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 13/25] mm: kpkeys: Introduce early page table allocator Kevin Brodsky
2026-08-27 18:08   ` David Hildenbrand (Arm)
2026-08-31 15:28     ` Kevin Brodsky
2026-08-27 18:17   ` Dave Hansen
2026-08-31 15:30     ` Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 14/25] mm: kpkeys: Protect vmemmap page tables Kevin Brodsky
2026-08-31 15:33   ` Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 15/25] mm: kpkeys: Introduce hook for protecting static " Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 16/25] arm64: kpkeys: Implement arch_supports_kpkeys_early() Kevin Brodsky
2026-08-18 14:08 ` [PATCH RFC v9 17/25] arm64: kpkeys: Support KPKEYS_CTX_PGTABLES Kevin Brodsky
2026-08-18 14:09 ` [PATCH RFC v9 18/25] arm64: kpkeys: Ensure the linear map can be modified Kevin Brodsky
2026-08-18 14:09 ` [PATCH RFC v9 19/25] arm64: kpkeys: Protect early page tables Kevin Brodsky
2026-08-18 14:09 ` [PATCH RFC v9 20/25] arm64: mm: Map kernel image alias of init_pg_dir read-only Kevin Brodsky
2026-08-18 14:09 ` [PATCH RFC v9 21/25] arm64: kpkeys: Protect init_pg_dir Kevin Brodsky
2026-08-18 14:09 ` [PATCH RFC v9 22/25] arm64: kpkeys: Guard page table writes Kevin Brodsky
2026-08-18 14:09 ` [PATCH RFC v9 23/25] arm64: kpkeys: Batch KPKEYS_CTX_PGTABLES switches Kevin Brodsky
2026-08-18 14:09 ` [PATCH RFC v9 24/25] arm64: kpkeys: Enable kpkeys_hardened_pgtables support Kevin Brodsky
2026-08-18 14:09 ` [PATCH RFC v9 25/25] mm: Add basic tests for kpkeys_hardened_pgtables Kevin Brodsky
2026-08-27 17:24 ` [PATCH RFC v9 00/25] pkeys-based page table hardening Yeoreum Yun
2026-08-31 15:35   ` Kevin Brodsky
2026-09-01 14:24 ` Linu Cherian
2026-09-03 16:47   ` Kevin Brodsky [this message]
2026-09-01 15:02 ` Linu Cherian
2026-09-01 15:13   ` Linu Cherian

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=43382882-0e7b-4255-9c34-99cec5b0bbbc@arm.com \
    --to=kevin.brodsky@arm.com \
    --cc=akpm@linux-foundation.org \
    --cc=broonie@kernel.org \
    --cc=catalin.marinas@arm.com \
    --cc=dave.hansen@linux.intel.com \
    --cc=david@kernel.org \
    --cc=iweiny@kernel.org \
    --cc=jannh@google.com \
    --cc=jeffxu@chromium.org \
    --cc=joey.gouly@arm.com \
    --cc=kees@kernel.org \
    --cc=linu.cherian@arm.com \
    --cc=linusw@kernel.org \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-hardening@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=luto@kernel.org \
    --cc=maz@kernel.org \
    --cc=mbland@motorola.com \
    --cc=peterz@infradead.org \
    --cc=pierre.langlois@arm.com \
    --cc=ptosi@google.com \
    --cc=qperret@google.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=tglx@kernel.org \
    --cc=vbabka@kernel.org \
    --cc=will@kernel.org \
    --cc=willy@infradead.org \
    --cc=x86@kernel.org \
    --cc=yang@os.amperecomputing.com \
    --cc=yeoreum.yun@arm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox