From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f197.google.com (mail-pf1-f197.google.com [209.85.210.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4EFA949EC4C for ; Tue, 1 Sep 2026 18:41:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.197 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788288107; cv=none; b=lb+eX+rKxN+2tcLYoEfiBsv9DEJwcEW1yiH2thOAqTuFRJfUvJ/XzY+I4ap/tl2M6mo9j1oRKNgu/9kBR71a4Zpz8OTb6HRvZfsuatxJogGgmglVjYRgJiOMHrnU4nyMT6tTTwKVPiEBsBCmluTjrme4c8sm48OFy1guFDv5gG8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788288107; c=relaxed/simple; bh=uVDf/Azs+Y2fNINbjdp5g3ddn/BJ/MorxkRf22rpgFo=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=SYMRo0tx1w742NuhpVcTCTV1rq4P1cQZrjqD+BLHgfXJ3y03RpmCgCQ9qm46twvB8cWcKJaoEZ2iL6wxyk2PPMGaJz5qlbvHYe2I9WaP4LvoAncX+JKCddGwIHJPEuS9vGgmSIuOLsRy2PvUmL4b6v++ChTpab/ZCBaQRmB67/Y= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=tGfaEMjP; arc=none smtp.client-ip=209.85.210.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="tGfaEMjP" Received: by mail-pf1-f197.google.com with SMTP id d2e1a72fcca58-853d0d87123so98994b3a.1 for ; Tue, 01 Sep 2026 11:41:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788288106; x=1788892906; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=T0nDJQnfFlnDLvDCZVuLzJX537jxkgQCnIHPdf7AiK4=; b=tGfaEMjPuDzdeYIuXqDVqbgEHDl/qAa6HamdMl6CjYdq0DFtFXA0q7E14PzIps2cqG wn9kSsxKcAnLVP3msnNDulTXefo5PrzBEnUPQhIP9j0TbSFVYa0gt6fkUXZDSvF2iF8p /58Tz2A45lnIdeJAhRo5H9Sk8J6Qq28CxzUkmth4Usr4U/MZQ2bmOvoS8Hlt/+lR99EY arWGyKTNDimSB17bj3eCuXy5h6hU+ZxwaMs+bbMZgIURj8ue7Sx23m/+Jz/jdivNeJml W3kPJ2UWu6frvWVaYy7CTzxyaPqXiF2rzSxWw4pd194YekDi5RGKogOAiN/G9u4Ueifj c36w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788288106; x=1788892906; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=T0nDJQnfFlnDLvDCZVuLzJX537jxkgQCnIHPdf7AiK4=; b=N/7/Q1gOth9CbqZzGmj5YPlwxrWw11QezSp8qF2eLYFqP3zGdCU69P9oiDUdmn6Z/e EXHemZmN1lt0CBc6PvCVjqBHKrBMjnNtTLBttJKMZXH+UOLEi3pSILmaojzlKs0WXMvG 9cFIU5so8BfV6DJb1UGUd3Q6ruCPnvQwqzTPIB6c6EZcEpzgU0hpDEpNCnG0nhCMcDWr +fI2/ZwueliZfkU/w6WGS2PYGSn060u1e6IfaCOpYwb6kgID1qFJrVcyEWK9LxgDCMl1 J0xWDPy3KmHm7GgQ5uEj1obCxtw9rDorBX05fq13TcHW+7s9VzIRmo4rFWujndNVTues Jd8A== X-Forwarded-Encrypted: i=1; AHgh+RriignQoOHn0vd1VUHN3LZi2Zu7tvX3jj3aClggAskZ96hpJGVQtRvYn/ybjAjxAbHXpjeZwYlYxSpKXmf5P0A=@vger.kernel.org X-Gm-Message-State: AFuF++mjG0jrTCdRotQj/Wq9dIG8WZM/APbAKHj/uCcnv6KnI+uw+KRY t57t7LKR0LMdVFTsuoMsZjDSD8pbRmbZj0DH25P9hNzaxyGpYbdQgDTi316xB655qyl5JKz8nbc krs9ocw== X-Received: from pfcy6.prod.google.com ([2002:a05:6a00:93c6:b0:848:49f6:6f5c]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:8016:b0:857:4dea:e20e with SMTP id d2e1a72fcca58-85b5693e833mr15207248b3a.0.1788288105188; Tue, 01 Sep 2026 11:41:45 -0700 (PDT) Date: Tue, 1 Sep 2026 11:41:44 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kselftest@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260830-gmem-inplace-conversion-v12-0-85e5fd25252a@google.com> <20260830-gmem-inplace-conversion-v12-15-85e5fd25252a@google.com> Message-ID: Subject: Re: [PATCH v12 15/45] KVM: guest_memfd: Ensure pages are not in use before conversion From: Sean Christopherson To: Fuad Tabba Cc: ackerleytng@google.com, aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Randy Dunlap , Lorenzo Stoakes , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev Content-Type: text/plain; charset="us-ascii" On Tue, Sep 01, 2026, Sean Christopherson wrote: > On Tue, Sep 01, 2026, Fuad Tabba wrote: > > > diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c > > > index 803c7cdbbe0f6..fe02c47c85fb5 100644 > > > --- a/virt/kvm/guest_memfd.c > > > +++ b/virt/kvm/guest_memfd.c > > > @@ -538,8 +538,46 @@ static int kvm_gmem_mas_preallocate(struct ma_state *mas, u64 attributes, > > > return mas_preallocate(mas, xa_mk_value(attributes), GFP_KERNEL); > > > } > > > > > > +static bool kvm_gmem_has_outstanding_references(struct inode *inode, > > > + pgoff_t start, size_t nr_pages, > > > + pgoff_t *err_index) > > > +{ > > > + struct address_space *mapping = inode->i_mapping; > > > + pgoff_t last = start + nr_pages - 1; > > > + bool has_outstanding = false; > > > + struct folio_batch fbatch; > > > + pgoff_t next; > > > + int i; > > > + > > > + folio_batch_init(&fbatch); > > > + > > > + next = start; > > > + while (has_outstanding && filemap_get_folios(mapping, &next, last, &fbatch)) { > > > > has_outstanding starts as false, so the loop never runs and the function > > always returns false. The outstanding-reference check is dead at this > > patch, so a to-private conversion would not be rejected even when a page > > still has an outstanding reference. > > > > It's fixed later in "KVM: guest_memfd: Handle lru_add fbatch refcounts > > during conversion safety check", which changes the condition to > > !has_outstanding. I think that fix belongs in this patch, so the check > > works when it is introduced and the series bisects cleanly. Wait, why are there even separate patches for this? For all intents and purposes, "Ensure pages are not in use before conversion" introduces a bug and then the bug is fixed by "Handle lru_add fbatch refcounts during conversion safety check". Just don't introduce the bug. I also recommend splitting the export of lru_cache_drain_for_folio() to its own patch so that it can be more easily Acked by mm/ folks. If we want to squash it with the KVM change, then that's trivial to do when applying. > Why even bother with has_outstanding? Avoiding it requires copy+pasting > folio_batch_release(), but it's less code and IMO the end result is a lot easier > to follow: > > struct address_space *mapping = inode->i_mapping; > pgoff_t last = start + nr_pages - 1; > struct folio_batch fbatch; > pgoff_t next; > int i; > > folio_batch_init(&fbatch); > > next = start; > while (filemap_get_folios(mapping, &next, last, &fbatch)) { > for (i = 0; i < folio_batch_count(&fbatch); ++i) { > struct folio *folio = fbatch.folios[i]; > > /* > * Outstanding references are anything other than those > * from the page cache, plus 1 temporary reference held > * by filemap_get_folios() in the folio batch. > */ > if (folio_ref_count(folio) != folio_nr_pages(folio) + 1) { > *err_index = max(start, folio->index); > folio_batch_release(&fbatch); > return true; > } > } > > folio_batch_release(&fbatch); > cond_resched(); > } > > return false; > > and then we end up with: > > enum lru_cache_drained drained = LRU_CACHE_NOT_DRAINED; > struct address_space *mapping = inode->i_mapping; > pgoff_t last = start + nr_pages - 1; > struct folio_batch fbatch; > pgoff_t next; > int i; > > folio_batch_init(&fbatch); > > next = start; > while (filemap_get_folios(mapping, &next, last, &fbatch)) { > for (i = 0; i < folio_batch_count(&fbatch); ++i) { > struct folio *folio = fbatch.folios[i]; > > if (__folio_has_outstanding_references(folio, &drained)) { > *err_index = max(start, folio->index); > folio_batch_release(&fbatch); > return true; > } > } > > folio_batch_release(&fbatch); > cond_resched(); > } > > return false;