From: Shameer Kolothum <skolothumtho@nvidia.com>
To: <kvm@vger.kernel.org>, <linux-pci@vger.kernel.org>,
<linux-kernel@vger.kernel.org>
Cc: <alex@shazbot.org>, <jgg@ziepe.ca>, <kevin.tian@intel.com>,
<kbusch@meta.com>, <michal.winiarski@intel.com>,
<satyanarayana.k.v.p@intel.com>, <sonangp@nvidia.com>,
<nathanc@nvidia.com>, <mochs@nvidia.com>
Subject: [RFC PATCH 09/19] vfio/pci: Serialize interrupt operations with recovery
Date: Tue, 1 Sep 2026 10:32:07 +0100 [thread overview]
Message-ID: <20260901093217.8539-10-skolothumtho@nvidia.com> (raw)
In-Reply-To: <20260901093217.8539-1-skolothumtho@nvidia.com>
Hold recovery_lock for reading around INTx, MSI and MSI-X capability
queries and configuration changes. ERR and REQ are software-only indexes
and stay available while recovery blocks device access. INTx is covered
by the same test even though its count comes from the virtual config
space, so that one rule applies to every index which can reach hardware.
The test is on the index alone, so a blocked device also refuses the few
requests on those indexes which would not have touched it: signalling an
eventfd for test purposes, and adding or removing the virqfd behind INTx
masking. Both return -EIO until access is unblocked, which for a
non-fatal error is the time the host takes to log it. Reading the flags
or the count of a request is not enough to tell whether it reaches the
device, and refusing a few extra requests for the length of an error
event is cheaper than getting that classification wrong.
Copy the IRQ payload from userspace before taking recovery_lock. The copy
can fault, and with userfaultfd the fault is serviced by userspace, so
holding the lock across it would let a user stall error_detected() for as
long as it likes. The count read and the interrupt operation each take
the lock for themselves.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
drivers/vfio/pci/vfio_pci_core.c | 55 ++++++++++++++++++++++++++++++++
1 file changed, 55 insertions(+)
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 0b1b2398dc88..876ff51d6987 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -1313,11 +1313,29 @@ int vfio_pci_ioctl_get_region_info(struct vfio_device *core_vdev,
}
EXPORT_SYMBOL_GPL(vfio_pci_ioctl_get_region_info);
+/*
+ * Which IRQ indexes can reach the device. ERR and REQ are software only.
+ * An index added later gets no access guard until it is listed here.
+ */
+static bool vfio_pci_irq_index_is_device(u32 index)
+{
+ switch (index) {
+ case VFIO_PCI_INTX_IRQ_INDEX:
+ case VFIO_PCI_MSI_IRQ_INDEX:
+ case VFIO_PCI_MSIX_IRQ_INDEX:
+ return true;
+ default:
+ return false;
+ }
+}
+
static int vfio_pci_ioctl_get_irq_info(struct vfio_pci_core_device *vdev,
struct vfio_irq_info __user *arg)
{
unsigned long minsz = offsetofend(struct vfio_irq_info, count);
struct vfio_irq_info info;
+ bool device_irq;
+ int ret;
if (copy_from_user(&info, arg, minsz))
return -EFAULT;
@@ -1336,7 +1354,15 @@ static int vfio_pci_ioctl_get_irq_info(struct vfio_pci_core_device *vdev,
info.flags = VFIO_IRQ_INFO_EVENTFD;
+ device_irq = vfio_pci_irq_index_is_device(info.index);
+ if (device_irq) {
+ ret = vfio_pci_core_access_begin(vdev);
+ if (ret)
+ return ret;
+ }
info.count = vfio_pci_get_irq_count(vdev, info.index);
+ if (device_irq)
+ vfio_pci_core_access_end(vdev);
if (info.index == VFIO_PCI_INTX_IRQ_INDEX)
info.flags |=
@@ -1353,13 +1379,23 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,
unsigned long minsz = offsetofend(struct vfio_irq_set, count);
struct vfio_irq_set hdr;
u8 *data = NULL;
+ bool device_irq;
int max, ret = 0;
size_t data_size = 0;
if (copy_from_user(&hdr, arg, minsz))
return -EFAULT;
+ device_irq = vfio_pci_irq_index_is_device(hdr.index);
+ if (device_irq) {
+ ret = vfio_pci_core_access_begin(vdev);
+ if (ret)
+ return ret;
+ }
max = vfio_pci_get_irq_count(vdev, hdr.index);
+ /* Dropped for the user copy below, which can fault under userfaultfd. */
+ if (device_irq)
+ vfio_pci_core_access_end(vdev);
ret = vfio_set_irqs_validate_and_prepare(&hdr, max, VFIO_PCI_NUM_IRQS,
&data_size);
@@ -1372,12 +1408,31 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,
return PTR_ERR(data);
}
+ /*
+ * Interrupt teardown reaches vfio_virqfd_disable(), which flushes the
+ * global virqfd cleanup workqueue, so recovery_lock is held here for
+ * as long as work queued by any vfio device takes. Shutdown work waits
+ * for its inject worker, and an ioeventfd inject takes that device's
+ * memory_lock, so the wait can last as long as a reset there. That is
+ * only a wait. Nothing on that workqueue takes recovery_lock, which is
+ * why the ioeventfd write path reads the recovery state without it. A
+ * callback there which used the vfio_pci_core_iowrite*() accessors
+ * would break that and deadlock against a queued writer.
+ */
+ if (device_irq) {
+ ret = vfio_pci_core_access_begin(vdev);
+ if (ret)
+ goto out_free;
+ }
mutex_lock(&vdev->igate);
ret = vfio_pci_set_irqs_ioctl(vdev, hdr.flags, hdr.index, hdr.start,
hdr.count, data);
mutex_unlock(&vdev->igate);
+ if (device_irq)
+ vfio_pci_core_access_end(vdev);
+out_free:
kfree(data);
return ret;
--
2.43.0
next prev parent reply other threads:[~2026-09-01 9:34 UTC|newest]
Thread overview: 50+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-01 9:31 [RFC PATCH 00/19] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
2026-09-01 9:31 ` [RFC PATCH 01/19] vfio/pci: Add PCI error recovery support state Shameer Kolothum
2026-09-01 9:45 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 02/19] vfio/pci: Serialize generic device lifetime with recovery Shameer Kolothum
2026-09-01 9:47 ` sashiko-bot
2026-09-01 13:14 ` K V P, Satyanarayana
2026-09-01 13:37 ` Shameer Kolothum Thodi
2026-09-01 9:32 ` [RFC PATCH 03/19] vfio/pci: Add PCI recovery access guards Shameer Kolothum
2026-09-01 9:39 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 04/19] vfio/pci: Serialize function reset with recovery Shameer Kolothum
2026-09-01 9:45 ` sashiko-bot
2026-09-02 6:06 ` K V P, Satyanarayana
2026-09-03 11:20 ` Shameer Kolothum Thodi
2026-09-01 9:32 ` [RFC PATCH 05/19] vfio/pci: Serialize config access " Shameer Kolothum
2026-09-01 9:46 ` sashiko-bot
2026-09-02 6:27 ` K V P, Satyanarayana
2026-09-03 11:08 ` Shameer Kolothum Thodi
2026-09-01 9:32 ` [RFC PATCH 06/19] vfio/pci: Serialize ioeventfd writes " Shameer Kolothum
2026-09-01 9:43 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 07/19] vfio/pci: Retry BAR faults after temporary recovery Shameer Kolothum
2026-09-01 9:47 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 08/19] vfio/pci: Serialize BAR and ROM access with recovery Shameer Kolothum
2026-09-01 9:48 ` sashiko-bot
2026-09-01 9:32 ` Shameer Kolothum [this message]
2026-09-01 9:42 ` [RFC PATCH 09/19] vfio/pci: Serialize interrupt operations " sashiko-bot
2026-09-03 6:34 ` K V P, Satyanarayana
2026-09-03 10:39 ` Shameer Kolothum Thodi
2026-09-01 9:32 ` [RFC PATCH 10/19] vfio/pci: Serialize hot reset " Shameer Kolothum
2026-09-01 9:58 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 11/19] vfio/pci: Serialize runtime PM " Shameer Kolothum
2026-09-01 9:49 ` sashiko-bot
2026-09-03 6:43 ` K V P, Satyanarayana
2026-09-03 10:47 ` Shameer Kolothum Thodi
2026-09-01 9:32 ` [RFC PATCH 12/19] vfio/pci: Serialize physical device information queries " Shameer Kolothum
2026-09-01 9:48 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 13/19] vfio/pci: Serialize DMA-BUF export " Shameer Kolothum
2026-09-01 9:43 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 14/19] vfio/pci: Add generic PCI error slot reset handling Shameer Kolothum
2026-09-01 9:53 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 15/19] vfio/pci: Add INTx helpers for PCI recovery Shameer Kolothum
2026-09-01 9:59 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 16/19] vfio/pci: Quiesce INTx during " Shameer Kolothum
2026-09-01 9:53 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 17/19] vfio/pci: Add generic PCI error resume handling Shameer Kolothum
2026-09-01 9:55 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 18/19] vfio/pci: Coordinate generic device access with host recovery Shameer Kolothum
2026-09-01 9:56 ` sashiko-bot
2026-09-01 9:32 ` [RFC PATCH 19/19] vfio/pci: Expose and enable host PCI error recovery Shameer Kolothum
2026-09-01 9:56 ` sashiko-bot
2026-09-04 19:09 ` [RFC PATCH 00/19] vfio/pci: Handle PCI error recovery and report state to userspace Alex Williamson
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260901093217.8539-10-skolothumtho@nvidia.com \
--to=skolothumtho@nvidia.com \
--cc=alex@shazbot.org \
--cc=jgg@ziepe.ca \
--cc=kbusch@meta.com \
--cc=kevin.tian@intel.com \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=michal.winiarski@intel.com \
--cc=mochs@nvidia.com \
--cc=nathanc@nvidia.com \
--cc=satyanarayana.k.v.p@intel.com \
--cc=sonangp@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox