From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f200.google.com (mail-pl1-f200.google.com [209.85.214.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 74E0742E41C for ; Thu, 27 Aug 2026 19:08:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.200 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787857689; cv=none; b=dkkUqoZED8N+/BXYW+K8TKHd7fp6dCnF+MTB1yaJJjthr116nYdjLHIqpElrSgwKmxIL4WW9DzXAJjuRwAbVIs/12kt/RlyjH4cMw4S5R8saJLkRIC8vuwQMvn2lNRIF0WWkYi7rgZJtWPcmWpldDcC+1Nr2gRj+uDfj/ytiwcs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787857689; c=relaxed/simple; bh=YAu1i5AWtB/mdPrcibBFB5fs0b8x5YEB+6wjvbVTug4=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=G5au5uAJvRhwx+hQujVFGS9xr37kLMOrcW3Xx6C4f+dlw1rMVpt37SlzkadYM8wNLy97HIazi0gljvCVk40+C8JurmgwvqIuw4Uw2AF8fu8UoeHNz6keCpjF+Zq+MRSvjsYU9EoC/PRtM3xWa8SebBNCjYbWjOoAPZSIS7ryTpE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=VJmxmaO3; arc=none smtp.client-ip=209.85.214.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="VJmxmaO3" Received: by mail-pl1-f200.google.com with SMTP id d9443c01a7336-2cee1ec30f2so1639875ad.3 for ; Thu, 27 Aug 2026 12:08:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787857687; x=1788462487; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=JWGZwnfyNQRYCnM1UqJIynPGxtHIQDoEVfaq/bKW7II=; b=VJmxmaO31r7QLA8hqwiQw6U4P4CwtVzNizJ3PZzdr5ThgK/gxjKEC3mTiUvUCottSv oah9VlHb0ddZ2PEAAUkx8914qBHIH/Kx6vyLtwqflipv1dGedMiZZsVwzgrt0+x7a0Zu x/VS8XZNJySa9B1UW3ax7qFGwbQ6yABtFR4v1IulsDfsvjdBJ5y9s1U6j9fiS2eVW6fA cccq98BRoExqMMtnC83N50B5qUqeXYzAIALkGGvL/Z2SpfGkoAkcyjitP9S5ZqXZr/lI sWpzHFmbhI/RZ63ony6FUltQ8Clo6ZzsL28salyR+gJxS3vYv/tS/eVFW3hDiLANmwuX ZTVA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787857687; x=1788462487; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=JWGZwnfyNQRYCnM1UqJIynPGxtHIQDoEVfaq/bKW7II=; b=IiTs52X5Vingj2az9PJFgCBiSb2bYhQc7NI6P/y+K01HrhRzqLR4xu+q5sF9+Xvbo9 ItpKP/3Tj8hIHh+gl1eA2grBnGD5yRGBiHAehCIH8p1cZ2A9wRkIVHZu3i5D1u6M9Cdx 89ZkBJFPdcDFojKLpE7lgGaovyLM3YEJa0Vb+owgzPcc761lH8rKqhvfDfEWnJ21hMhq RtOujDNUFXoP/XEC7GEocCxAb+K7hlBZ4RUPUzMr/lVy3zNhaDCnDzSW9rmJHy2Hji1D A9PXAxeSqmE/R03vwHiABsOPT+DD5Lefivw4Ue/wkFGe8IHAxNmmGp71D0tbx2hhWP2y nThQ== X-Forwarded-Encrypted: i=1; AHgh+RotoV59o03CRVMN213ItTowo2kuO65DrdsvFg1SiwKOi2GBbkDSUsFLGw0wqykWySHFqrW1EA4zBFD/3hut4xYEFH8=@vger.kernel.org X-Gm-Message-State: AFuF++k35i4t1X5jZ/ABTtDgoDINgKmeHBtx/lLcXdPwc3Qhge/I8adQ Ou+ZHE63iiQWXY5SZk6aok4e/nDK1d0GvOrFWLajA4AIUxkKnpxAtPnParvoKWbCu8+dL++fBFO 17Fh9ig== X-Received: from plblq15.prod.google.com ([2002:a17:903:144f:b0:2d6:e58f:8e81]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:15cf:b0:2d7:1aeb:739 with SMTP id d9443c01a7336-2d74e1295c2mr19606415ad.13.1787857686248; Thu, 27 Aug 2026 12:08:06 -0700 (PDT) Date: Thu, 27 Aug 2026 12:08:05 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-trace-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <7eba513d-cdae-4832-abac-038dcdd88517@intel.com> Message-ID: Subject: Re: [PATCH v10 09/41] KVM: guest_memfd: Filter both shared and private when invalidating From: Sean Christopherson To: Xiaoyao Li Cc: Ackerley Tng , Michael Roth , Suzuki K Poulose , "David Hildenbrand (Arm)" , aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, jmattson@google.com, jthoughton@google.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, tabba@google.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Vlastimil Babka , kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev Content-Type: text/plain; charset="us-ascii" On Thu, Aug 27, 2026, Xiaoyao Li wrote: > On 8/26/2026 9:53 PM, Sean Christopherson wrote: > > > > Is this just a case of user error? That with gmem_in_place_conversion, > > > userspace should not use userspace_addr from something other than the > > > gmem for the same memslot? > > > > Yes. It's not just invalidations that will go sideways, > > > KVM accesses to guest > > memory won't hit the same physical page as actual guest accesses. > > Side topic. > > If KVM enforces KVM_MEMSLOT_GMEM_ONLY when gmem_in_place_conversion is true > like below proposal, should we update KVM's guest memory accessors to access > gmem directly? No, because routing KVM accesses through uaccess means we just have to ensure PRIVATE memory isn't mapped into userspace, which we already have to do anyways. Now, that doesn't preclude future work from optimizing certain flows to avoid GUP, or to harden KVM in other ways. But as a first step, bypassing uaccess is actively dangerous. > If still use the existing code of accessing userspace_addr of the memslot, I > think KVM needs to document clearly that when > KVM_CAP_GUEST_MEMFD_MEMORY_ATTRIBUTES is enumerated, if configuring the > guest_memfd for a memslot, the userspace_addr passed in needs to be the > mmaped address of the guest_memfd. Ya. > > Huh. But that isn't strictly guaranteed, because userspace could bind to a > > memslot that isn't configured with GUEST_MEMFD_FLAG_MMAP, in which case SHARED > > faults will go through the VMA, not kvm_mmu_faultin_pfn_gmem(). It's a bit early > > in the morning, but off the top of my head, I can't think of any reason we need > > to support such a setup. If userspace really, really wants to use a separate > > mapping, they could DELETE+CREATE an equivalent memslot without the guest_memfd > > file descriptor. > > > > So I think we should do this? > > > > diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c > > index 9c2d52bdf25e..86f53e53a136 100644 > > --- a/virt/kvm/guest_memfd.c > > +++ b/virt/kvm/guest_memfd.c > > @@ -1013,7 +1013,7 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot, > > */ > > WRITE_ONCE(slot->gmem.file, file); > > slot->gmem.pgoff = start; > > - if (kvm_gmem_supports_mmap(inode)) > > + if (gmem_in_place_conversion || kvm_gmem_supports_mmap(inode)) > > slot->flags |= KVM_MEMSLOT_GMEM_ONLY; > > I like this idea. It makes gmem_in_place_conversion a step closer to what > its name implies, though in-place conversion is not truly 100% > guaranteed[*]. > > [*] https://lore.kernel.org/all/akUbz_kJvYulaboo@google.com