From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f199.google.com (mail-pf1-f199.google.com [209.85.210.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7722349DB84 for ; Tue, 1 Sep 2026 18:41:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.199 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788288108; cv=none; b=pco0d7DNQdu4oi7jEI3hwwYob9+ra4UYaQZh3bFp/Ub7EIim2g4QF8J0wY+BDkIZniWgC1VXk6rh93XNGfLIPLyxcFHgNZpgIL7IdJgir+kxhb5WED4fB/vzJplqxLLGilo+a2meMd52bJMHcLKw9dAVIdjjjoUEU+SOADT9dOo= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788288108; c=relaxed/simple; bh=uVDf/Azs+Y2fNINbjdp5g3ddn/BJ/MorxkRf22rpgFo=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=c7wMgnjd1nlLMdvplr1k1nkqmGQVxyqrXspotSOa9SWanxTQlt+yPobRPxLVGrboihf3cS43ECETsLp7IOpZxPnAss7uA2oWzfPr6P5403FB2G2O9czoezeyQ+sqIQt84Axm+IJEIypIE2w3HAwj89KO6aZqRXCkCyAYP9F2Yl8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=tGfaEMjP; arc=none smtp.client-ip=209.85.210.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="tGfaEMjP" Received: by mail-pf1-f199.google.com with SMTP id d2e1a72fcca58-84842381150so165165b3a.3 for ; Tue, 01 Sep 2026 11:41:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788288106; x=1788892906; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=T0nDJQnfFlnDLvDCZVuLzJX537jxkgQCnIHPdf7AiK4=; b=tGfaEMjPuDzdeYIuXqDVqbgEHDl/qAa6HamdMl6CjYdq0DFtFXA0q7E14PzIps2cqG wn9kSsxKcAnLVP3msnNDulTXefo5PrzBEnUPQhIP9j0TbSFVYa0gt6fkUXZDSvF2iF8p /58Tz2A45lnIdeJAhRo5H9Sk8J6Qq28CxzUkmth4Usr4U/MZQ2bmOvoS8Hlt/+lR99EY arWGyKTNDimSB17bj3eCuXy5h6hU+ZxwaMs+bbMZgIURj8ue7Sx23m/+Jz/jdivNeJml W3kPJ2UWu6frvWVaYy7CTzxyaPqXiF2rzSxWw4pd194YekDi5RGKogOAiN/G9u4Ueifj c36w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788288106; x=1788892906; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=T0nDJQnfFlnDLvDCZVuLzJX537jxkgQCnIHPdf7AiK4=; b=pnXcR9GUYweubdPNQ6rmGzJm3dXDd1ULWPB6BCDWYTyJDFvpI0CD6LyVq/EfxSTDXQ JL94hRjdUzrgF4lVL6IGOxX1xoDHnnKt6CC1NJYMaIsaeF2d6b0gZB1eCfAYeuwlHQJd 3i75bPTraTT2vNIIH+WAwWnhlITlc2Sfi+nNrneAEGOJd2syaV9ZihIXuQcSNywYDEjI YiDM+DEhhyiQfBYhPrAJtbpDCcobr7Rsdc0cze63krEMVFEad/T7GbV26AzrwqFFvnL7 7aHqfx/hF6jZiacqrao+ClRME0yXexbv7yJ39ZKe9erHTEe4IIyLob8Z+R/U/ay+snga iAPw== X-Forwarded-Encrypted: i=1; AHgh+Rr8wfE2xuRfCXi69hcmqW39MOpndmSqHfWTdKEfyaRjGknuaPdNCxnOhS431kpXwvuVcqT5I7wtxqM=@vger.kernel.org X-Gm-Message-State: AFuF++m95Y8m2b5vqQf0UAbA+QmQoTwcYG0I1HdewGQUhSfYz+tWvIkw i1Qb99a5hOc1WJa5m/N4YWzFbOfyh8GxuqxMYpjbJDO80VozPCSRetyxi4KP99nvjYcP6ruaHGG OaSptpA== X-Received: from pfcy6.prod.google.com ([2002:a05:6a00:93c6:b0:848:49f6:6f5c]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:8016:b0:857:4dea:e20e with SMTP id d2e1a72fcca58-85b5693e833mr15207248b3a.0.1788288105188; Tue, 01 Sep 2026 11:41:45 -0700 (PDT) Date: Tue, 1 Sep 2026 11:41:44 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-doc@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260830-gmem-inplace-conversion-v12-0-85e5fd25252a@google.com> <20260830-gmem-inplace-conversion-v12-15-85e5fd25252a@google.com> Message-ID: Subject: Re: [PATCH v12 15/45] KVM: guest_memfd: Ensure pages are not in use before conversion From: Sean Christopherson To: Fuad Tabba Cc: ackerleytng@google.com, aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Randy Dunlap , Lorenzo Stoakes , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev Content-Type: text/plain; charset="us-ascii" On Tue, Sep 01, 2026, Sean Christopherson wrote: > On Tue, Sep 01, 2026, Fuad Tabba wrote: > > > diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c > > > index 803c7cdbbe0f6..fe02c47c85fb5 100644 > > > --- a/virt/kvm/guest_memfd.c > > > +++ b/virt/kvm/guest_memfd.c > > > @@ -538,8 +538,46 @@ static int kvm_gmem_mas_preallocate(struct ma_state *mas, u64 attributes, > > > return mas_preallocate(mas, xa_mk_value(attributes), GFP_KERNEL); > > > } > > > > > > +static bool kvm_gmem_has_outstanding_references(struct inode *inode, > > > + pgoff_t start, size_t nr_pages, > > > + pgoff_t *err_index) > > > +{ > > > + struct address_space *mapping = inode->i_mapping; > > > + pgoff_t last = start + nr_pages - 1; > > > + bool has_outstanding = false; > > > + struct folio_batch fbatch; > > > + pgoff_t next; > > > + int i; > > > + > > > + folio_batch_init(&fbatch); > > > + > > > + next = start; > > > + while (has_outstanding && filemap_get_folios(mapping, &next, last, &fbatch)) { > > > > has_outstanding starts as false, so the loop never runs and the function > > always returns false. The outstanding-reference check is dead at this > > patch, so a to-private conversion would not be rejected even when a page > > still has an outstanding reference. > > > > It's fixed later in "KVM: guest_memfd: Handle lru_add fbatch refcounts > > during conversion safety check", which changes the condition to > > !has_outstanding. I think that fix belongs in this patch, so the check > > works when it is introduced and the series bisects cleanly. Wait, why are there even separate patches for this? For all intents and purposes, "Ensure pages are not in use before conversion" introduces a bug and then the bug is fixed by "Handle lru_add fbatch refcounts during conversion safety check". Just don't introduce the bug. I also recommend splitting the export of lru_cache_drain_for_folio() to its own patch so that it can be more easily Acked by mm/ folks. If we want to squash it with the KVM change, then that's trivial to do when applying. > Why even bother with has_outstanding? Avoiding it requires copy+pasting > folio_batch_release(), but it's less code and IMO the end result is a lot easier > to follow: > > struct address_space *mapping = inode->i_mapping; > pgoff_t last = start + nr_pages - 1; > struct folio_batch fbatch; > pgoff_t next; > int i; > > folio_batch_init(&fbatch); > > next = start; > while (filemap_get_folios(mapping, &next, last, &fbatch)) { > for (i = 0; i < folio_batch_count(&fbatch); ++i) { > struct folio *folio = fbatch.folios[i]; > > /* > * Outstanding references are anything other than those > * from the page cache, plus 1 temporary reference held > * by filemap_get_folios() in the folio batch. > */ > if (folio_ref_count(folio) != folio_nr_pages(folio) + 1) { > *err_index = max(start, folio->index); > folio_batch_release(&fbatch); > return true; > } > } > > folio_batch_release(&fbatch); > cond_resched(); > } > > return false; > > and then we end up with: > > enum lru_cache_drained drained = LRU_CACHE_NOT_DRAINED; > struct address_space *mapping = inode->i_mapping; > pgoff_t last = start + nr_pages - 1; > struct folio_batch fbatch; > pgoff_t next; > int i; > > folio_batch_init(&fbatch); > > next = start; > while (filemap_get_folios(mapping, &next, last, &fbatch)) { > for (i = 0; i < folio_batch_count(&fbatch); ++i) { > struct folio *folio = fbatch.folios[i]; > > if (__folio_has_outstanding_references(folio, &drained)) { > *err_index = max(start, folio->index); > folio_batch_release(&fbatch); > return true; > } > } > > folio_batch_release(&fbatch); > cond_resched(); > } > > return false;