From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wr1-f73.google.com (mail-wr1-f73.google.com [209.85.221.73]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 71116375ABD for ; Tue, 31 Mar 2026 14:40:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.73 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1774968056; cv=none; b=ryGJF+2moluu8Gok0XoYw/Td/R/sEfSFNt1nJ+OftT61XNoKJvqpeIjihRWKaWaPSdMrRSliTdaefYBjww1luwL5JfLy1G2PSjNNFp1YXuqez6pKDKwf59gsS5gDKqnwysQxC20zMhGGC9fem05IaPeGDBBImu4fUZ36DbAR9Gw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1774968056; c=relaxed/simple; bh=dKtEBSNPsz4SC85+BqreMD82Vc0GQPF7dyhNajgebqg=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=kk9Sdmdn+g5eDAB9S66OELWozvouo3tgcRKj80eHUqpveCcneAKWVfo9kBqCaml7AbCfD4Drt+avxpIcVDIsxnbyyqmgLIGb40nAWHjnFDtrE7jE/f1pBC5pombz5qkeMXwboJmXhvUWemhkfZhmBAbzr8NkL/d7kXsB0DvtzDk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jackmanb.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=AbOc9RvK; arc=none smtp.client-ip=209.85.221.73 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jackmanb.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="AbOc9RvK" Received: by mail-wr1-f73.google.com with SMTP id ffacd0b85a97d-43cf906afb8so2566272f8f.3 for ; Tue, 31 Mar 2026 07:40:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1774968053; x=1775572853; darn=vger.kernel.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=uZfpcLfgq+5yt+rcx10V2YKKXlDGG+s50zObM35odx8=; b=AbOc9RvKHrjVtBnTUHeJGI4g6MVbKqRWTdujvaraYiSaiRaYEuflavZ4zxCBzizlmn JppANIUdBcd3Arrv+VKcjcnSdBXkShcpdv6OfNvOqNVsv7fxAr1TZmCK+F+RlSiZrTlk Mjy48KrcpznXAjQvCAelLAPkqyP4kFkMB3TFR40dqCe2os0wb/4iUrS8/mZVc3ETusQ3 amrojxGCoxLJuLWXQKp3QYXOHoqsp9u28kMsDWh9W4HOnqdIswu7p99+AzhdAiwAM64j s2JEFRCzfb2TZ1x7Heool/BW7xS6kD5p91nXo5E0GZJRr6ufP9U290fkV/9wf+azOlsV w/HA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1774968053; x=1775572853; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=uZfpcLfgq+5yt+rcx10V2YKKXlDGG+s50zObM35odx8=; b=DqeVHsB3XbM+/BvScIs0cD2bleQ3V5ryEHk3B4TU4OcPDMLKVX7cQ0IRsgDc/nEO3l bh3z9nFZ5hKPy43OzGHcqsK3/mWGvpyGo0c/fbGpU54ee8dz9mghHHGAohqrHeygBINn f5/XYnn8IjA1YfR8xQ8l1riKp5mgtDqNVnD9c7fz4Ly6FM6gf7B2YidpgkoRZXyCyGpN XxFz0aLTZc58VmRKNxe7VHBsE/lofYGESolVX1vsn8UxXvbY8nG3jnQ1rmxJQRLyxWfh e9ddnOnGqiWLkOEJLprYseBr/PmHydae+joCfLCZmRkKvcCjydJOoV0qRw5eEsJRvNgp EYzQ== X-Forwarded-Encrypted: i=1; AJvYcCWN5WvfB8ZHYKpwoyPPhcow8TjSGydif0LCQ5/7F7XpOfP/vLpZu5ibMZK89TeriQi8yknNaD3xV2JMYQw=@vger.kernel.org X-Gm-Message-State: AOJu0YyTNIbeW8xGw1zQxWKO90MEBZ1kn2viES3Poo/cEuONey/k8/mi qxVxqUhIWF1w6jO29Tgph6o4MUORAZ3ALhTQLK4MGoOFf6haduH0N4yjm+GAlrPp/K7FrNiYFtF ThpAXKxpQk1/Olw== X-Received: from wrwj2.prod.google.com ([2002:a5d:5642:0:b0:43c:fde3:c2bf]) (user=jackmanb job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6000:2f85:b0:43c:f66e:f2d with SMTP id ffacd0b85a97d-43cf66e113cmr20228191f8f.27.1774968052063; Tue, 31 Mar 2026 07:40:52 -0700 (PDT) Date: Tue, 31 Mar 2026 14:40:51 +0000 In-Reply-To: <20260320-page_alloc-unmapped-v2-22-28bf1bd54f41@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260320-page_alloc-unmapped-v2-0-28bf1bd54f41@google.com> <20260320-page_alloc-unmapped-v2-22-28bf1bd54f41@google.com> X-Mailer: aerc 0.21.0 Message-ID: Subject: Re: [PATCH v2 22/22] mm/secretmem: Use __GFP_UNMAPPED when available From: Brendan Jackman To: Brendan Jackman , Borislav Petkov , Dave Hansen , Peter Zijlstra , Andrew Morton , David Hildenbrand , Vlastimil Babka , Wei Xu , Johannes Weiner , Zi Yan , Lorenzo Stoakes Cc: , , , , Sumit Garg , , , Will Deacon , , "Kalyazin, Nikita" , , "Itazuri, Takahiro" , Andy Lutomirski , David Kaplan , Thomas Gleixner , Yosry Ahmed Content-Type: text/plain; charset="UTF-8" I skipped on posting most of the intervening AI reviews coz there's just so much stuff and most of it is pretty boring, but this one is interesting. https://sashiko.dev/#/patchset/20260320-page_alloc-unmapped-v2-0-28bf1bd54f41%40google.com On Fri Mar 20, 2026 at 6:23 PM UTC, Brendan Jackman wrote: > This is the simplest possible way to adopt __GFP_UNMAPPED. Use it to > allocate pages when it's available, meaning the > set_direct_map_invalid_noflush() call is no longer needed. > > Signed-off-by: Brendan Jackman > --- > mm/secretmem.c | 87 +++++++++++++++++++++++++++++++++++++++++++++++++--------- > 1 file changed, 74 insertions(+), 13 deletions(-) > > diff --git a/mm/secretmem.c b/mm/secretmem.c > index 5f57ac4720d32..9fef91237358a 100644 > --- a/mm/secretmem.c > +++ b/mm/secretmem.c > @@ -6,6 +6,7 @@ > */ > > #include > +#include > #include > #include > #include > @@ -47,13 +48,78 @@ bool secretmem_active(void) > return !!atomic_read(&secretmem_users); > } > > +/* > + * If it's supported, allocate using __GFP_UNMAPPED. This lets the page > + * allocator amortize TLB flushes and avoids direct map fragmentation. > + */ > +#ifdef CONFIG_PAGE_ALLOC_UNMAPPED > +static inline struct folio *secretmem_folio_alloc(gfp_t gfp, unsigned int order) > +{ > + int err; > + > + /* Required for __GFP_UNMAPPED|__GFP_ZERO. */ > + err = mermap_mm_prepare(current->mm); > + if (err) > + return ERR_PTR(err); Sashiko: > In remote access paths such as process_vm_readv or io_uring worker threads, > current->mm might be NULL or point to a different address space than the > faulted VMA. Should this use vmf->vma->vm_mm instead? This can't happen, right? I think the assumption is good; doing a secretmem fault in a kthread or any process that doesn't have the file mmap()'d would be a bug? But this definitely feels like something I could be wrong about. AI slop to empirically check the two examples mentioned by the AI fail early (slop for slop, it's slop all the way down...): - process_vm_readv(): https://paste.debian.net/hidden/5625ef2e - io_uring: https://paste.debian.net/hidden/e7763ad2 [...] > +static inline void secretmem_folio_restore(struct folio *folio) > +{ > + set_direct_map_default_noflush(folio_page(folio, 0)); > +} I defined a lovely helper here but neglected to actually call it. Sashiko says "this isn't a bug" but that's wrong, calling set_direct_map_default_noflush() before freeing a __GFP_UNMAPPED page is not OK. (Should it be? I think no. We _could_ define "default" so that it checks the pageblock flags and does the right thing for you. But then we'd be baking in the assumption that the page allocator can efficiently look that up. This would be rather tricky if we decided we need to mix mapped an unmapped pages in the same block). And this was hiding the more important bug which was that I forgot to do the mermap dance for zeroing the page. > +static inline void secretmem_folio_flush(struct folio *folio) > +{ > + unsigned long addr = (unsigned long)folio_address(folio); > + > + flush_tlb_kernel_range(addr, addr + PAGE_SIZE); > +} > +#endif > + > static vm_fault_t secretmem_fault(struct vm_fault *vmf) > { > struct address_space *mapping = vmf->vma->vm_file->f_mapping; > struct inode *inode = file_inode(vmf->vma->vm_file); > pgoff_t offset = vmf->pgoff; > gfp_t gfp = vmf->gfp_mask; > - unsigned long addr; > struct folio *folio; > vm_fault_t ret; > int err; > @@ -66,16 +132,9 @@ static vm_fault_t secretmem_fault(struct vm_fault *vmf) > retry: > folio = filemap_lock_folio(mapping, offset); > if (IS_ERR(folio)) { > - folio = folio_alloc(gfp | __GFP_ZERO, 0); > - if (!folio) { > - ret = VM_FAULT_OOM; > - goto out; > - } > - > - err = set_direct_map_invalid_noflush(folio_page(folio, 0)); > - if (err) { > - folio_put(folio); > - ret = vmf_error(err); > + folio = secretmem_folio_alloc(gfp | __GFP_ZERO, 0); > + if (IS_ERR_OR_NULL(folio)) { > + ret = folio ? vmf_error(PTR_ERR(folio)) : VM_FAULT_OOM; > goto out; > } > > err = filemap_add_folio(mapping, folio, offset, gfp); > if (unlikely(err)) { > /* > * If a split of large page was required, it > * already happened when we marked the page invalid > * which guarantees that this call won't fail > */ > set_direct_map_default_noflush(folio_page(folio, 0)); > folio_put(folio); > if (err == -EEXIST) > goto retry; This failure path leaks mermap TLB entries on ther CPUs. So if an attacker can trigger this path, and cause another CPU to populate its TLB for the mermap region, and then cause a victim to allocate the page they just freed, they can use those entries for a sidechannel attack. I did call out in the cover letter that there's some jank with the "if you use __GFP_UNMAPPED|__GFP_ZERO then you need to think about the TLB" thing. But, this shows it's worse than I thought. I need to think about how to mitigate this.