All of lore.kernel.org
 help / color / mirror / Atom feed
From: Lu Baolu <baolu.lu@linux.intel.com>
To: Joerg Roedel <joro@8bytes.org>
Cc: ZhaoJinming <zhaojinming@uniontech.com>,
	Kevin Tian <kevin.tian@intel.com>,
	Dmitry Antipov <dmantipov@yandex.ru>,
	Guanghui Feng <guanghuifeng@linux.alibaba.com>,
	Li RongQing <lirongqing@baidu.com>,
	Desnes Nunes <desnesn@redhat.com>,
	iommu@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: [PATCH v2 01/19] iommu/vt-d: Fix UCTP context table slot when copying root entries
Date: Wed,  5 Aug 2026 07:42:55 +0800	[thread overview]
Message-ID: <20260804234314.3087110-2-baolu.lu@linux.intel.com> (raw)
In-Reply-To: <20260804234314.3087110-1-baolu.lu@linux.intel.com>

From: Desnes Nunes <desnesn@redhat.com>

When translation is already enabled at boot (e.g. kdump), the vt-d driver
copies context tables from the previous kernel's root table. In scalable
mode, buses that only populate the upper root half (UCTP, devfn >= 0x80)
should be written to ctxt_tbls[tbl_idx + 1] through copy_context_table().
However, the current copy path always uses tbl[tbl_idx + 0] in this situa-
tion. Since idx wraps to 0 at devfn 0x80 due to a zeroed LCTP, new_ce for
LCTP will be NULL and keep pos equals to 0. Thus, UCTP entries will be co-
pied into tbl[tbl_idx + 0] instead of tbl[tbl_idx + 1], and written after-
wards to root_entry[bus].lo instead of .hi in copy_translation_tables().

In short, devices on bus 0x80 with devfn >= 0x80 fail DMA with fault 0x39,
which will break drivers running in kernels with translation pre-enabled.
This fixes NO_PASID DMAR faults for UCTP-only buses such as:

DMAR: [DMA Read NO_PASID] Request device [80:14.0] fault addr 0xe81759000
      [fault reason 0x39] SM: Present bit in Root Entry is clear

For instance, this fault yielded to locking issues between systemd and
xHCI, blocking a system's reboot after a vmcore was captured with kdump:

 systemd-udevd[246]: usb3: Worker [255] processing SEQNUM=2193 is taking a long time
 dracut-initqueue[277]: Timed out while waiting for udev queue to empty.
 systemd-udevd[246]: usb3: Worker [255] processing SEQNUM=2193 killed
 systemd-udevd[246]: usb3: Worker [255] terminated by signal 9 (KILL).
 ...
 kdump[569]: saving vmcore complete
 ...
 systemd-shutdown[1]: Rebooting.
 INFO: task kworker/0:1:11 blocked for more than 122 seconds.
       Not tainted 7.0.0-clean #1
 "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
 task:kworker/0:1 state:D stack:0 pid:11 tgid:11 ppid:2 task_flags:0x4208160 flags:0x00080000
 Workqueue: usb_hub_wq hub_event
 Call Trace:
  <TASK>
  __schedule+0x299/0x5c0
  schedule+0x27/0x80
  schedule_timeout+0xbd/0x100
  __wait_for_common+0x97/0x1b0
  ? __pfx_schedule_timeout+0x10/0x10
  xhci_alloc_dev+0x9e/0x2b0
  usb_alloc_dev+0x7a/0x3b0
  hub_port_connect+0x285/0x960
  hub_port_connect_change+0x94/0x290
  port_event+0x4bb/0x840
  hub_event+0x141/0x460
  process_one_work+0x196/0x390
  worker_thread+0x1af/0x320
  ? __pfx_worker_thread+0x10/0x10
  kthread+0xe3/0x120
  ? __pfx_kthread+0x10/0x10
  ret_from_fork+0x199/0x260
  ? __pfx_kthread+0x10/0x10
  ret_from_fork_asm+0x1a/0x30
  </TASK>
 INFO: task systemd-shutdow:1 blocked for more than 122 seconds.
       Not tainted 7.0.0-clean #1
 "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
 task:systemd-shutdow state:D stack:0 pid:1 tgid:1 ppid:0 task_flags:0x400100 flags:0x00080000
 Call Trace:
  <TASK>
  __schedule+0x299/0x5c0
  schedule+0x27/0x80
  schedule_preempt_disabled+0x15/0x30
  __mutex_lock.constprop.0+0x547/0xac0
  device_shutdown+0xac/0x1b0
  kernel_restart+0x3a/0x70
  __do_sys_reboot+0x147/0x240
  do_syscall_64+0x11b/0x6a0
  ? handle_mm_fault+0x110/0x350
  ? do_user_addr_fault+0x206/0x680
  ? irqentry_exit+0x7a/0x4d0
  entry_SYSCALL_64_after_hwframe+0x76/0x7e
 RIP: 0033:0x7fe2958da917
 RSP: 002b:00007ffc5c458618 EFLAGS: 00000206 ORIG_RAX: 00000000000000a9
 RAX: ffffffffffffffda RBX: 0000000000000000 RCX: 00007fe2958da917
 RDX: 0000000001234567 RSI: 0000000028121969 RDI: 00000000fee1dead
 RBP: 00007ffc5c458790 R08: 0000000000000069 R09: 00000000ffffffff
 R10: 0000000000000000 R11: 0000000000000206 R12: 0000000000000000
 R13: 0000000000000000 R14: 00007ffc5c4588b8 R15: 0000000000000000
  </TASK>
 INFO: task systemd-shutdow:1 is blocked on a mutex likely owned by task kworker/0:1:11.

Fixes: 091d42e43d21 ("iommu/vt-d: Copy translation tables from old kernel")
Signed-off-by: Desnes Nunes <desnesn@redhat.com>
Tested-by: Tao Liu <ltao@redhat.com>
Signed-off-by: Lu Baolu <baolu.lu@linux.intel.com>
---
 drivers/iommu/intel/iommu.c | 10 ++++++----
 1 file changed, 6 insertions(+), 4 deletions(-)

diff --git a/drivers/iommu/intel/iommu.c b/drivers/iommu/intel/iommu.c
index 849d06dfe1ae..cf5f92619943 100644
--- a/drivers/iommu/intel/iommu.c
+++ b/drivers/iommu/intel/iommu.c
@@ -1446,7 +1446,7 @@ static int copy_context_table(struct intel_iommu *iommu,
 			      struct context_entry **tbl,
 			      int bus, bool ext)
 {
-	int tbl_idx, pos = 0, idx, devfn, ret = 0, did;
+	int tbl_idx, tbl_slot = 0, idx, devfn, ret = 0, did;
 	struct context_entry *new_ce = NULL, ce;
 	struct context_entry *old_ce = NULL;
 	struct root_entry re;
@@ -1462,10 +1462,9 @@ static int copy_context_table(struct intel_iommu *iommu,
 		if (idx == 0) {
 			/* First save what we may have and clean up */
 			if (new_ce) {
-				tbl[tbl_idx] = new_ce;
+				tbl[tbl_idx + tbl_slot] = new_ce;
 				__iommu_flush_cache(iommu, new_ce,
 						    VTD_PAGE_SIZE);
-				pos = 1;
 			}
 
 			if (old_ce)
@@ -1487,6 +1486,9 @@ static int copy_context_table(struct intel_iommu *iommu,
 				}
 			}
 
+			/* Track if saving UCTP or LCTP entries in scalable mode */
+			tbl_slot = ext && devfn >= 0x80 ? 1 : 0;
+
 			ret = -ENOMEM;
 			old_ce = memremap(old_ce_phys, PAGE_SIZE,
 					MEMREMAP_WB);
@@ -1515,7 +1517,7 @@ static int copy_context_table(struct intel_iommu *iommu,
 		new_ce[idx] = ce;
 	}
 
-	tbl[tbl_idx + pos] = new_ce;
+	tbl[tbl_idx + tbl_slot] = new_ce;
 
 	__iommu_flush_cache(iommu, new_ce, VTD_PAGE_SIZE);
 
-- 
2.43.0


  reply	other threads:[~2026-08-04 23:54 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-04 23:42 [PATCH v2 00/19][PULL REQUEST] Intel IOMMU updates for v7.3 Lu Baolu
2026-08-04 23:42 ` Lu Baolu [this message]
2026-08-04 23:42 ` [PATCH v2 02/19] iommu/vt-d: Use logical OR operator for privilege mode check Lu Baolu
2026-08-04 23:42 ` [PATCH v2 03/19] iommu/vt-d: Fix CACHE_TAG_NESTING_DEVTLB polluting shared variables in flush loop Lu Baolu
2026-08-04 23:42 ` [PATCH v2 04/19] iommu/vt-d: Use kstrtoint_from_user() in dmar_perf_latency_write() Lu Baolu
2026-08-04 23:42 ` [PATCH v2 05/19] iommu/vt-d: Fix no_iommu to disable platform opt-in Lu Baolu
2026-08-04 23:43 ` [PATCH v2 06/19] iommu/vt-d: Force requesting ACS when tboot is enabled Lu Baolu
2026-08-04 23:43 ` [PATCH v2 07/19] iommu/vt-d: Remove dead code when CONFIG_INTEL_IOMMU is not set Lu Baolu
2026-08-04 23:43 ` [PATCH v2 08/19] iommu/vt-d: Consolidate dmar policy management and force_on logic Lu Baolu
2026-08-04 23:43 ` [PATCH v2 09/19] iommu/vt-d: Use dmar_can_force_on() for platform opt-in Lu Baolu
2026-08-04 23:43 ` [PATCH v2 10/19] iommu/vt-d: Call dmar_can_force_on() for tboot opt-in Lu Baolu
2026-08-04 23:43 ` [PATCH v2 11/19] iommu/vt-d: Remove the 'force_on' variable Lu Baolu
2026-08-04 23:43 ` [PATCH v2 12/19] iommu/vt-d: Remove dmar_disabled Lu Baolu
2026-08-04 23:43 ` [PATCH v2 13/19] iommu/vt-d: Support the new DMA_REMAP_OPT_OUT flag bit Lu Baolu
2026-08-04 23:43 ` [PATCH v2 14/19] iommu/vt-d: Cache max domain ID to avoid redundant calculation Lu Baolu
2026-08-04 23:43 ` [PATCH v2 15/19] iommu/vt-d: Fix copied_tables bitmap leak on error in copy_translation_tables Lu Baolu
2026-08-04 23:43 ` [PATCH v2 16/19] iommu/vt-d: Clear Present bit before tearing down copied context entry Lu Baolu
2026-08-04 23:43 ` [PATCH v2 17/19] iommu/vt-d: Fix iopf_refcount leak on RID domain replacement Lu Baolu
2026-08-04 23:43 ` [PATCH v2 18/19] iommu/vt-d: Tear down scalable-mode context on probe failure Lu Baolu
2026-08-04 23:43 ` [PATCH v2 19/19] iommu/vt-d: Flush context cache with correct SID when tearing down aliases Lu Baolu
2026-08-10  8:04 ` [PATCH v2 00/19][PULL REQUEST] Intel IOMMU updates for v7.3 Joerg Roedel

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260804234314.3087110-2-baolu.lu@linux.intel.com \
    --to=baolu.lu@linux.intel.com \
    --cc=desnesn@redhat.com \
    --cc=dmantipov@yandex.ru \
    --cc=guanghuifeng@linux.alibaba.com \
    --cc=iommu@lists.linux.dev \
    --cc=joro@8bytes.org \
    --cc=kevin.tian@intel.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lirongqing@baidu.com \
    --cc=zhaojinming@uniontech.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.