The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* [PATCH v2] iommu/dma: Restore locking around msi_page_list
@ 2026-07-30 13:23 Andrew Jones
  2026-07-30 13:52 ` Robin Murphy
  2026-07-31  4:25 ` Nutty.Liu
  0 siblings, 2 replies; 3+ messages in thread
From: Andrew Jones @ 2026-07-30 13:23 UTC (permalink / raw)
  To: iommu, linux-kernel; +Cc: robin.murphy, joro, will, nicolinc, jgg

Unlike a group's default domain, which is always freshly allocated
and privately owned (iommu_group_alloc_default_domain()), VFIO type1's
legacy container merges any newly attached group into an existing
domain whenever their iommu_ops and cache-coherency enforcement match.

iommu_dma_get_msi_page() only asserts the caller's own group mutex is
held (iommu_group_mutex_assert()). On an IOMMU that publishes
IOMMU_RESV_SW_MSI, e.g. ARM SMMU, a VM with two such devices assigned
through the legacy container can have their guest drivers probe and
allocate MSIs in parallel; each host-side VFIO_DEVICE_SET_IRQS lands
on a different device fd and group mutex, but both devices' domains
are the same merged domain, so both can enter
iommu_dma_get_msi_page() concurrently and corrupt msi_page_list.

commit 288683c92b1a ("iommu: Make iommu_dma_prepare_msi() into a
generic operation") dropped the prior msi_prepare_lock on the
reasoning that "each iommu_domain is unique to a group," which holds
for default domains but not this VFIO type1 case. Restore the static
lock, since it's only guarding a corner case and will likely never
be contended.

iommufd avoids the equivalent problem by having its own callers
(iommufd_sw_map_msi()) take a ctx-wide sw_msi_lock before ever
reaching the shared list. VFIO type1 can't mirror that since it
dispatches to iommu_dma_sw_msi() which is outside VFIO's jurisdiction.

Fixes: 288683c92b1a ("iommu: Make iommu_dma_prepare_msi() into a generic operation")
Signed-off-by: Andrew Jones <andrew.jones@oss.qualcomm.com>
---
Sashiko reported this issue while reviewing a riscv iommu series[1].
I've only compile-tested this fix.

[1] https://sashiko.dev/#/patchset/20260724151218.965929-1-andrew.jones@oss.qualcomm.com

v2:
 - switched back to statick lock as 288683c92b1a had [Robin]


 drivers/iommu/dma-iommu.c | 13 +++++++++++++
 1 file changed, 13 insertions(+)

diff --git a/drivers/iommu/dma-iommu.c b/drivers/iommu/dma-iommu.c
index 9abaec0703ef..9a07eb39336e 100644
--- a/drivers/iommu/dma-iommu.c
+++ b/drivers/iommu/dma-iommu.c
@@ -2204,6 +2204,19 @@ static struct iommu_dma_msi_page *iommu_dma_get_msi_page(struct device *dev,
 	dma_addr_t iova;
 	int prot = IOMMU_WRITE | IOMMU_NOEXEC | IOMMU_MMIO;
 	size_t size = cookie_msi_granule(domain);
+	static DEFINE_MUTEX(msi_prepare_lock);
+
+	/*
+	 * Normally a device's default domain is only ever attached to that
+	 * device's own group, and the group mutex held by
+	 * iommu_group_mutex_assert()'s callers is enough on its own. A VFIO
+	 * type1 container is the one case that breaks that assumption: it
+	 * can merge devices from different groups onto one domain, so two
+	 * devices' group mutexes don't serialize each other here. A static
+	 * lock is sufficient due to the expectation that this is a corner
+	 * case that will never be contended in practice.
+	 */
+	guard(mutex)(&msi_prepare_lock);
 
 	msi_addr &= ~(phys_addr_t)(size - 1);
 	list_for_each_entry(msi_page, msi_page_list, list)
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-07-31  4:25 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-30 13:23 [PATCH v2] iommu/dma: Restore locking around msi_page_list Andrew Jones
2026-07-30 13:52 ` Robin Murphy
2026-07-31  4:25 ` Nutty.Liu

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox