Linux PCI subsystem development
 help / color / mirror / Atom feed
From: Nicolin Chen <nicolinc@nvidia.com>
To: <will@kernel.org>, <robin.murphy@arm.com>, <jgg@nvidia.com>
Cc: <joro@8bytes.org>, <bhelgaas@google.com>, <praan@google.com>,
	<kevin.tian@intel.com>, <kees@kernel.org>, <smostafa@google.com>,
	<baolu.lu@linux.intel.com>,
	<linux-arm-kernel@lists.infradead.org>, <iommu@lists.linux.dev>,
	<linux-kernel@vger.kernel.org>, <linux-pci@vger.kernel.org>,
	<skaestle@nvidia.com>, <mmarrid@nvidia.com>,
	<skolothumtho@nvidia.com>, <bbiber@nvidia.com>,
	<harsha.v@oss.qualcomm.com>
Subject: [PATCH v3 03/13] iommu/arm-smmu-v3: Drain in-flight fault events on domain detach
Date: Mon, 31 Aug 2026 17:33:28 -0700	[thread overview]
Message-ID: <8717f3329313f40308cff20006066283ebbb3a56.1788222485.git.nicolinc@nvidia.com> (raw)
In-Reply-To: <cover.1788222485.git.nicolinc@nvidia.com>

When a device is switching away from a domain, either through a detach or a
replace operation, in-flight stall events for the old domain might still be
on the SMMU's hardware event queue or on the IOMMU core's IOPF queue. Thus,
if the IOMMU core swaps the device's attach_handle and frees the old domain
before those handlers complete, the IOPF work might hit use-after-free.

Two queues need to be drained: the SMMU hardware event queue and the IOMMU
core IOPF software workqueue. Start with the former: add a counting-based
arm_smmu_drain_queue() helper, and poll the evtq on a domain detach, so a
pending IRQ won't let the threaded handler run after the drain and queue a
fault referencing the domain being freed. Its until_empty mode serves the
suspend and runtime PM routines that would drain the CMDQ. Any timed-out
drain fires a WARN_ON as well, since reaching the timeout would take some
stuck consumer in any realistic case.

The existing queue_poll() API is not reusable for such a drain: it is the
atomic busy-wait for the command issuing paths, and it assumes a hardware
consumer making progress. A drain caller is sleepable, in contrast, while
the EVTQ/PRIQ consumer is a threaded IRQ handler that needs the CPU: such
a busy wait would starve the handler throughout an entire timeout, whenever
the waiter and the handler shared one CPU on a non-preemptible kernel. So,
this new sleeping helper is marked with a might_sleep() as well, given that
an atomic-context misuse would otherwise hide behind an empty queue.

Note that a drained event is dequeued, but not necessarily handled, since
queue_remove_raw() moves the MMIO CONS before the threaded IRQ handler gets
to push the event onto the IOPF workqueue. A subsequent change will invoke
synchronize_irq() and iopf_queue_flush_dev() to close that gap, and it will
act on the errno of a timed-out drain too.

The drain runs before the IOMMU core swaps the device's attach handle, so a
fault event generated on the new STE during this window resolves to the old
handle, completing with IOMMU_PAGE_RESP_INVALID that resumes the stall with
abort: the impact is bounded to that one failed transaction.

Also run the drain for every stall-capable master, even when the departing
attachment did not enable IOPF: such a stall event has to be aborted while
it still resolves to the old attach handle, otherwise the threaded handler
could pick it up right after the handle swap, mistakenly resuming it as if
it were a valid page fault against a new domain.

Fixes: cfea71aea921 ("iommu/arm-smmu-v3: Put iopf enablement in the domain attach path")
Cc: stable@vger.kernel.org # v6.16
Assisted-by: Claude:claude-fable-5
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
 drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 92 +++++++++++++++++++++
 1 file changed, 92 insertions(+)

diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index e00b6c88214f5..d255ff2519f9d 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -948,6 +948,86 @@ static int arm_smmu_cmdq_batch_submit(struct arm_smmu_device *smmu,
 					   cmds->num, true);
 }
 
+/**
+ * arm_smmu_drain_queue - Drain an SMMU queue
+ * @smmu: the SMMU device
+ * @q: the queue to drain
+ * @until_empty: target selection
+ *
+ * With @until_empty == true (for CMDQ), exit once the queue is observed empty:
+ *
+ *   cons0                cons                                prod
+ *     |                   |                                   |
+ *  ---+###################+=====================+=============+--->
+ *                         |<--------- undrained==0? --------->|
+ *
+ * With @until_empty == false (for EVTQ/PRIQ), exit once "drained" reaches its
+ * target: "pending" (i.e. prod0 - cons0, frozen at the entry time):
+ *
+ *   cons0                cons                 prod0         (prod)
+ *     |<---- drained ---->|                     |             |
+ *  ---+###################+=====================+=============+--->
+ *     |<--------------- pending --------------->|
+ *
+ * Note that a drained entry is dequeued, but not necessarily handled: the
+ * EVTQ/PRIQ callers must follow up with a synchronize_irq() to wait for the
+ * threaded IRQ handler to finish handling the dequeued entries.
+ *
+ * Context: Process context; may sleep.
+ * Return: 0 on success or a negative errno on timeout.
+ */
+static int arm_smmu_drain_queue(struct arm_smmu_device *smmu,
+				struct arm_smmu_queue *q, bool until_empty)
+{
+	ktime_t timeout = ktime_add_us(ktime_get(), ARM_SMMU_POLL_TIMEOUT_US);
+	u32 cons, prod, prev, undrained;
+	u32 drained = 0, pending;
+
+	might_sleep();
+
+	cons = readl_relaxed(q->cons_reg);
+	prod = readl_relaxed(q->prod_reg);
+	/* The exit target: the number of entries in the queue at entry */
+	pending = Q_POS(&q->llq, prod - cons);
+
+	while (true) {
+		/* Accumulate the entries consumed since the last poll */
+		prev = cons;
+		cons = readl_relaxed(q->cons_reg);
+		drained += Q_POS(&q->llq, cons - prev);
+
+		prod = readl_relaxed(q->prod_reg);
+		undrained = Q_POS(&q->llq, prod - cons);
+
+		/* Exit on an empty queue, regardless of until_empty */
+		if (!undrained)
+			return 0;
+
+		/* Snapshot mode: exit once the pending entries are drained */
+		if (!until_empty && drained >= pending)
+			return 0;
+
+		/*
+		 * A timeout means the consumer might be stuck. In theory, if it
+		 * moves 2 * qsize entries or more within a single poll interval
+		 * Q_POS() would wrap and undercount drained: that could trigger
+		 * a spurious warning too, if the queue was never once observed
+		 * empty. Yet, that much consumption in such a short interval is
+		 * unrealistic. WARN it only, as a stuck consumer is a real bug.
+		 */
+		if (WARN_ON(ktime_compare(ktime_get(), timeout) > 0))
+			break;
+
+		/* The consumer might be a threaded IRQ handler. Yield to it */
+		usleep_range(100, 200);
+	}
+
+	dev_warn_ratelimited(smmu->dev,
+			     "queue drain timed out at prod=0x%x cons=0x%x\n",
+			     prod, cons);
+	return -ETIMEDOUT;
+}
+
 static void arm_smmu_page_response(struct device *dev, struct iopf_fault *unused,
 				   struct iommu_page_response *resp)
 {
@@ -3318,11 +3398,23 @@ void arm_smmu_attach_release(struct arm_smmu_attach_state *state)
 {
 	struct arm_smmu_master_domain *master_domain = state->old_master_domain;
 	struct arm_smmu_master *master = state->master;
+	struct arm_smmu_device *smmu = master->smmu;
 
+	lockdep_assert_not_held(&arm_smmu_asid_lock);
 	iommu_group_mutex_assert(master->dev);
 
 	if (!master_domain)
 		return;
+
+	/*
+	 * Drain the hardware eventq, while stale events still resolve to the
+	 * old attach handle. Otherwise, the threaded handler could pick one
+	 * up once the IOMMU core swaps the handle, mistakenly resuming it
+	 * against the next domain.
+	 */
+	if (master->stall_enabled)
+		arm_smmu_drain_queue(smmu, &smmu->evtq.q, false);
+
 	arm_smmu_disable_iopf(master, master_domain);
 	kfree(master_domain);
 	state->old_master_domain = NULL;
-- 
2.43.0


  parent reply	other threads:[~2026-09-01  0:34 UTC|newest]

Thread overview: 40+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-01  0:33 [PATCH v3 00/13] iommu/arm-smmu-v3: Add PRI support Nicolin Chen
2026-09-01  0:33 ` [PATCH v3 01/13] iommu/arm-smmu-v3: Add arm_smmu_attach_release() Nicolin Chen
2026-09-01  0:46   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 02/13] iommu/arm-smmu-v3: Add Q_POS() macro Nicolin Chen
2026-09-01  0:38   ` sashiko-bot
2026-09-01  0:33 ` Nicolin Chen [this message]
2026-09-01  0:48   ` [PATCH v3 03/13] iommu/arm-smmu-v3: Drain in-flight fault events on domain detach sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 04/13] iommu/arm-smmu-v3: Flush in-flight fault work " Nicolin Chen
2026-09-01  0:55   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 05/13] iommu/arm-smmu-v3: Allocate IOPF queue without FEAT_SVA Nicolin Chen
2026-09-01  0:46   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 06/13] iommu/arm-smmu-v3: Submit CMDQ_OP_PRI_RESP for IOPF event Nicolin Chen
2026-09-01  0:53   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 07/13] iommu/arm-smmu-v3: Disable the queue IRQs before disabling the SMMU Nicolin Chen
2026-09-01  0:55   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 08/13] iommu/arm-smmu-v3: Disable PRI when no IRQ handler is registered Nicolin Chen
2026-09-01  0:47   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 09/13] iommu/arm-smmu-v3: Support PRI Page Request in arm_smmu_handle_ppr() Nicolin Chen
2026-09-01  0:50   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 10/13] iommu/arm-smmu-v3: Allocate IOPF queue for ARM_SMMU_FEAT_PRI Nicolin Chen
2026-09-01  0:43   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 11/13] PCI/ATS: Add PRI stubs Nicolin Chen
2026-09-01  0:42   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 12/13] PCI/ATS: Export pci_enable_pri() and pci_reset_pri() Nicolin Chen
2026-09-01  0:44   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-01  0:33 ` [PATCH v3 13/13] iommu/arm-smmu-v3: Enable PRI for PCI device in arm_smmu_probe_device() Nicolin Chen
2026-09-01  0:51   ` sashiko-bot
2026-09-03 19:18   ` Jonathan Cameron
2026-09-03 19:18 ` [PATCH v3 00/13] iommu/arm-smmu-v3: Add PRI support Jonathan Cameron

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=8717f3329313f40308cff20006066283ebbb3a56.1788222485.git.nicolinc@nvidia.com \
    --to=nicolinc@nvidia.com \
    --cc=baolu.lu@linux.intel.com \
    --cc=bbiber@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=harsha.v@oss.qualcomm.com \
    --cc=iommu@lists.linux.dev \
    --cc=jgg@nvidia.com \
    --cc=joro@8bytes.org \
    --cc=kees@kernel.org \
    --cc=kevin.tian@intel.com \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=mmarrid@nvidia.com \
    --cc=praan@google.com \
    --cc=robin.murphy@arm.com \
    --cc=skaestle@nvidia.com \
    --cc=skolothumtho@nvidia.com \
    --cc=smostafa@google.com \
    --cc=will@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox