From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6E27642A15C for ; Thu, 6 Aug 2026 23:00:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786057223; cv=none; b=WF0Cn3qCW8JCcS2ckGnHdW+2OujHVdUQt8LsKJtwLOpIf89+5jEe5EkWjnyUNaMadVdek11P30Ffaee/itbF70mjMkDSLxT3kZ5bNY9YkXu1zGbXrxQgEioh/qlq9BqWahMlWDa/k/E3OO+AatPK0aYZXWcbIA+IwhSNFTPqJUI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786057223; c=relaxed/simple; bh=I0wqDwZyoQSIsMb6537lKMySu1ARTn+0eN39hDtnsKo=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=WH3dn5ZiGxfnvc+LPJXlyGbg+NZqCHBjOb03GXmMzhecpzHMwSpxEeQkITm5UeG0RHzLONKjFEAeBjOqNOedrwPl52LkJcyptpNxT/fr2rC1juI8gDO0KUXn2c4B2J/EO35IWB1BeFoctp9NVq4X0uWkpLoPLgO8U1KcWTQAUVg= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=ByvyTw1N; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=kIcrDmog; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="ByvyTw1N"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="kIcrDmog" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1786057220; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=V/XyIJ0P1Ncw2ZC5RUbuOjY9InJ1Bze3sfj+kDZLJRI=; b=ByvyTw1N4MfMpMUK+t96yxbBqXKWefbzfGijHjLv7YxQQlixgy4L2VUEwNABGI9agKSNBa tZB2+7gw2o4wQG+/KMNWs7XSCSsGTxDkm8C/S6PbMTCbpqFLJm149P03wKNWRQFSFF4S5V 2raOUtMTC1cmmaMv3URHB1naKWV/HO0= Received: from mail-wr1-f71.google.com (mail-wr1-f71.google.com [209.85.221.71]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-604-XnwbHcpaOj6G91sFqUk-wQ-1; Thu, 06 Aug 2026 19:00:09 -0400 X-MC-Unique: XnwbHcpaOj6G91sFqUk-wQ-1 X-Mimecast-MFC-AGG-ID: XnwbHcpaOj6G91sFqUk-wQ_1786057208 Received: by mail-wr1-f71.google.com with SMTP id ffacd0b85a97d-473ac08a6a4so1838802f8f.0 for ; Thu, 06 Aug 2026 16:00:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1786057208; x=1786662008; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=V/XyIJ0P1Ncw2ZC5RUbuOjY9InJ1Bze3sfj+kDZLJRI=; b=kIcrDmogvSe+ZPn48xABRzkPuDHwWUGTaco7YvVvKSoUjcbTFtbviWBDZm8ZBogZsY sy4pG28HzFuw1gE+h7Ccaf8g/chQOCKEpFjzoAKjOumGrCG/4qouhRqcolR5LyKYRPzP Ov2z88E+oJyepXWiQ84+b8JvxMnMShQfMb7vnNIN7fhbJ/GTagZCljXFVvA+lzrfPpNl cj4mEytBrnKFdyKBX0RzAFsHUcHYVflwBay/9oAx7Ps/iaY2/0BK4ti7uIginI73WdiH 1zUIZ58t4+7fLI/sQN1E8ztMBUEyzQ7Vp/47IGtwQgbQRpOyPmBLMO9GY+sQyYdIDSqB J1dg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786057208; x=1786662008; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=V/XyIJ0P1Ncw2ZC5RUbuOjY9InJ1Bze3sfj+kDZLJRI=; b=fozbAJH7jEr1dvaWv2fY4YwUk3ZlZdOd+t6SHM7aRdJSuGYimO8SJWMka/G/TSbfmn WDTe34m7fwUyGEE3XW2jfyWQEKdhdmMACSdtE90k6ky48KLXYNZMQOgbbFc7fm28eQrs K7n2HQ3TiOhSqA+lJ0Tpgl/LH/hb1NNY+e6LlKurLtqTPZYjd+d+1opXWNj1i87G7yVe Uv7mAZPfcxHY+G/eDx14spr8EpPqJ8/IkFK0NWFbs8GRnHlHRBN3YP+4Xhd7YIvmBbVi zQXPCZrD9iky9fMfI+bxR50UNixRY0216+dVB5lgeVaD5U3cJwYHKObPZFmHl6z8mLDx YreQ== X-Forwarded-Encrypted: i=1; AHgh+RpMEu6pdCn05d1B8Zkh+gZ7wtQbbJMrBdqfcNyoctJhc7b9RagfaoGd+Vo9oGju9Jo0DUKi8TRjFQvbyLw=@vger.kernel.org X-Gm-Message-State: AOJu0YzLZnyR50IvKXWGRw00VzlCs3EEBR9LM+M/u8QwRGncBVZfefhR J1wzsnhWbeGRWfedR5ohhI/mnfb2sHpVldStMwOnI8/fcnLTIa8Sla7VVvKlEDrAhlO5He4uWc9 B2g77NlQ3cH6AZ2haEtKwl8AZLEaaRRkmEzI8jYu8qbl34TBitYn6E49GtmTBMyLWUg== X-Gm-Gg: AR+sD11ER2Q6xBbwThvEUJotLcqTABnoxgRlxu95lujShwiwcjsHozrzFA84gSlNaFk pcmnK3jhxke3opLGR6smL5dRtRavpU0HkyCMv5Et5bW1BqgjMyvavObr95CUR93mxB67nMRIIrd kakek17iCebDyMAP6hy6iK4+9h5c249km1iP8wxCWdG/tPDa48OJ0Tjm+qkM/omNjBvkMu49sCV BavQM7/ajmEdn3dm05LSWTg4SkK273C736eC4jCPB6IuZRKSPqshvtXYVq/uQTgwRRvQgSsTkgh QsxhQFKjzl9v0Lu9TVH9Bh/6kSB0Pmm0f5djiY0Jd5oRzsK8lyq/26TFE2bVgspTr5sGvp/px2a nlXDZMzhHoCHzGNvY9m4bVA== X-Received: by 2002:a05:600c:3b04:b0:495:5e86:4e11 with SMTP id 5b1f17b1804b1-4994e7baa85mr237163555e9.12.1786057207650; Thu, 06 Aug 2026 16:00:07 -0700 (PDT) X-Received: by 2002:a05:600c:3b04:b0:495:5e86:4e11 with SMTP id 5b1f17b1804b1-4994e7baa85mr237162725e9.12.1786057207114; Thu, 06 Aug 2026 16:00:07 -0700 (PDT) Received: from redhat.com (IGLD-80-230-28-14.inter.net.il. [80.230.28.14]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4995d85254asm1707125e9.1.2026.08.06.16.00.03 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 16:00:06 -0700 (PDT) Date: Thu, 6 Aug 2026 19:00:01 -0400 From: "Michael S. Tsirkin" To: Link Lin Cc: Jason Wang , Xuan Zhuo , Andrew Morton , David Hildenbrand , Vlastimil Babka , virtualization@lists.linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, prasin@google.com, rientjes@google.com, duenwen@google.com, jiaqiyan@google.com, ahwilkins@google.com, Greg Thelen , Alexander Duyck , jthoughton@google.com, stable@vger.kernel.org, Cory Maccarrone , Taylor Scanlon Subject: Re: [RFC PATCH] virtio_balloon: add VIRTIO_BALLOON_F_REPORTING_PM_SAFE feature bit Message-ID: <20260806185414-mutt-send-email-mst@kernel.org> References: <20260724014605.3377283-1-linkl@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260724014605.3377283-1-linkl@google.com> On Fri, Jul 24, 2026 at 01:46:05AM +0000, Link Lin wrote: > Following up on the fix for the PM suspend Use-After-Free race condition in > mm/page_reporting (merged in mm-hotfixes-unstable: > https://lore.kernel.org/all/20260723003650.CAAF01F000E9@smtp.kernel.org/), > we face a hypervisor-side deployment dilemma. > > Cloud hypervisors want to safely enable the Free Page Reporting (FPR) > virtqueue across their fleets, but enabling it indiscriminately on guests > without the recent suspend fix exposes them to UAF crashes. Relying on > out-of-band metadata (e.g., OS image tags) to selectively enable the > feature is fragile for custom user images or live-patched kernels. > > To address this at the protocol level, we propose adding a new feature bit > to the Virtio Specification: > VIRTIO_BALLOON_F_REPORTING_PM_SAFE (Bit 6) > > This establishes a formal device lifecycle contract for power management: > If negotiated, the driver MUST guarantee that all page reporting operations > are halted and pending requests are flushed before the device/system > transitions into a suspended state (e.g., ACPI S3/S4). > Deployment semantics: > - Hypervisors operating in a strict "safe mode" can offer Bit 6 exclusively > (suppressing Bit 5 / VIRTIO_BALLOON_F_REPORTING). > - Older, unpatched Linux guests will see Bit 5 is absent, ignore Bit 6, and > safely skip FPR initialization, preventing the suspend crash. Standard > ballooning remains 100% functional. > - Patched Linux guests will recognize Bit 6 and safely initialize FPR. > - Note for fleet deployments: Non-Linux guests (e.g., Windows, FreeBSD) > that rely on Bit 5 will temporarily lose FPR if the hypervisor exclusively > offers Bit 6. This is considered an acceptable trade-off to globally > protect unpatched guests without relying on OS image tags, until those > respective virtio drivers adopt Bit 6. This makes no sense to me. So there's a bug in the guest and it crashes. Patch the guest. There's just no chance we'll add flags every time some guest drivers on some OSes have a UAF. What makes this specific bug special? Trust me when I say UAF issues e.g. around hotplug are a dime a dozen. Now what, let's add another one around hotplug? And so on. > > Implementation Note on Upstream/Downstream Dependencies: > -------------------------------------------------------- > Because it is critical that downstream Linux distros do not accidentally > backport Bit 6 without the core MM UAF fix, the final upstream > implementation of this patch will enforce a strict compile-time dependency. > We plan to export a macro (e.g., PAGE_REPORTING_HAS_FREEZABLE_WQ) from the > core MM fix, and wrap Bit 6 behind an #ifdef of that macro in > virtio_balloon.c. This guarantees that compiler backports must consume the > entire dependency chain to advertise the feature. > > Below is the proposed Linux proof-of-concept based on upstream master. We > introduce a helper virtio_balloon_has_reporting() to ensure virtqueues are > properly allocated, torn down, and validated if either bit is negotiated. > > If this architectural approach is acceptable for cloud deployments, we will > formally submit this patch and open a corresponding issue for the OASIS > Virtio specification. > > Depends-on: > Signed-off-by: Link Lin > --- > drivers/virtio/virtio_balloon.c | 15 +++++++++++---- > include/uapi/linux/virtio_balloon.h | 1 + > 2 files changed, 12 insertions(+), 4 deletions(-) > > diff --git a/drivers/virtio/virtio_balloon.c b/drivers/virtio/virtio_balloon.c > index 581ac799d9..00e8273dc3 100644 > --- a/drivers/virtio/virtio_balloon.c > +++ b/drivers/virtio/virtio_balloon.c > @@ -39,6 +39,12 @@ > (1 << (VIRTIO_BALLOON_HINT_BLOCK_ORDER + PAGE_SHIFT)) > #define VIRTIO_BALLOON_HINT_BLOCK_PAGES (1 << VIRTIO_BALLOON_HINT_BLOCK_ORDER) > > +static inline bool virtio_balloon_has_reporting(struct virtio_device *vdev) > +{ > + return virtio_has_feature(vdev, VIRTIO_BALLOON_F_REPORTING) || > + virtio_has_feature(vdev, VIRTIO_BALLOON_F_REPORTING_PM_SAFE); > +} > + > enum virtio_balloon_vq { > VIRTIO_BALLOON_VQ_INFLATE, > VIRTIO_BALLOON_VQ_DEFLATE, > @@ -598,7 +604,7 @@ static int init_vqs(struct virtio_balloon *vb) > if (virtio_has_feature(vb->vdev, VIRTIO_BALLOON_F_FREE_PAGE_HINT)) > vqs_info[VIRTIO_BALLOON_VQ_FREE_PAGE].name = "free_page_vq"; > > - if (virtio_has_feature(vb->vdev, VIRTIO_BALLOON_F_REPORTING)) { > + if (virtio_balloon_has_reporting(vb->vdev)) { > vqs_info[VIRTIO_BALLOON_VQ_REPORTING].name = "reporting_vq"; > vqs_info[VIRTIO_BALLOON_VQ_REPORTING].callback = balloon_ack; > } > @@ -635,7 +641,7 @@ static int init_vqs(struct virtio_balloon *vb) > if (virtio_has_feature(vb->vdev, VIRTIO_BALLOON_F_FREE_PAGE_HINT)) > vb->free_page_vq = vqs[VIRTIO_BALLOON_VQ_FREE_PAGE]; > > - if (virtio_has_feature(vb->vdev, VIRTIO_BALLOON_F_REPORTING)) > + if (virtio_balloon_has_reporting(vb->vdev)) > vb->reporting_vq = vqs[VIRTIO_BALLOON_VQ_REPORTING]; > > return 0; > @@ -1013,7 +1019,7 @@ static int virtballoon_probe(struct virtio_device *vdev) > } > > vb->pr_dev_info.report = virtballoon_free_page_report; > - if (virtio_has_feature(vb->vdev, VIRTIO_BALLOON_F_REPORTING)) { > + if (virtio_balloon_has_reporting(vb->vdev)) { > unsigned int capacity; > > capacity = virtqueue_get_vring_size(vb->reporting_vq); > @@ -1099,7 +1105,7 @@ static void virtballoon_remove(struct virtio_device *vdev) > { > struct virtio_balloon *vb = vdev->priv; > > - if (virtio_has_feature(vb->vdev, VIRTIO_BALLOON_F_REPORTING)) > + if (virtio_balloon_has_reporting(vb->vdev)) > page_reporting_unregister(&vb->pr_dev_info); > if (virtio_has_feature(vb->vdev, VIRTIO_BALLOON_F_DEFLATE_ON_OOM)) > unregister_oom_notifier(&vb->oom_nb); > @@ -1162,8 +1168,10 @@ static int virtballoon_validate(struct virtio_device *vdev) > */ > if (!want_init_on_free() && !page_poisoning_enabled_static()) > __virtio_clear_bit(vdev, VIRTIO_BALLOON_F_PAGE_POISON); > - else if (!virtio_has_feature(vdev, VIRTIO_BALLOON_F_PAGE_POISON)) > + else if (!virtio_has_feature(vdev, VIRTIO_BALLOON_F_PAGE_POISON)) { > __virtio_clear_bit(vdev, VIRTIO_BALLOON_F_REPORTING); > + __virtio_clear_bit(vdev, VIRTIO_BALLOON_F_REPORTING_PM_SAFE); > + } > > __virtio_clear_bit(vdev, VIRTIO_F_ACCESS_PLATFORM); > return 0; > @@ -1176,6 +1184,7 @@ static unsigned int features[] = { > VIRTIO_BALLOON_F_FREE_PAGE_HINT, > VIRTIO_BALLOON_F_PAGE_POISON, > VIRTIO_BALLOON_F_REPORTING, > + VIRTIO_BALLOON_F_REPORTING_PM_SAFE, > }; > > static struct virtio_driver virtio_balloon_driver = { > diff --git a/include/uapi/linux/virtio_balloon.h b/include/uapi/linux/virtio_balloon.h > index ee35a37280..d206f156d6 100644 > --- a/include/uapi/linux/virtio_balloon.h > +++ b/include/uapi/linux/virtio_balloon.h > @@ -37,6 +37,7 @@ > #define VIRTIO_BALLOON_F_FREE_PAGE_HINT 3 /* VQ to report free pages */ > #define VIRTIO_BALLOON_F_PAGE_POISON 4 /* Guest is using page poisoning */ > #define VIRTIO_BALLOON_F_REPORTING 5 /* Page reporting virtqueue */ > +#define VIRTIO_BALLOON_F_REPORTING_PM_SAFE 6 /* PM-safe page reporting */ > > /* Size of a PFN in the balloon interface. */ > #define VIRTIO_BALLOON_PFN_SHIFT 12 > --