dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4
@ 2026-09-30  3:32 David Zhang
  2026-09-30  3:32 ` [PATCH V1 01/20] accel/amdxdna: Rename NPU3 firmware files David Zhang
                   ` (19 more replies)
  0 siblings, 20 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

This patch series extends the amdxdna accelerator driver to support
AMD AIE4 (NPU3) platforms with kernel-mode command submission (KMQ),
device telemetry queries, comprehensive system and runtime power
management, and NPU3 classic device support.

This series establishes kernel-mode submission as the default execution
path. Additionally, it implements suspend/resume and runtime power
management across Physical Functions (PF), Virtual Functions (VF), and
Classic devices.

Changes in v1:
- Firmware version 6.0: return IOMMU_PASID_INVALID from aie4_msg_pasid()
- Power mode: consolidate cached power mode override restoration on
  hardware start directly into the power mode introduction patch, and
  respect user buffer size.
- Transport hooks: fix header include to use <drm/amdxdna_accel.h>.
- Command submission & fencing:
  - Fold fence timeline naming (using dev_name() backed by struct
    device to avoid use-after-free) and unique timeline context allocation
    into the command submission patch.

Series structure:

1. Firmware Interfaces, Versioning & Compatibility (Patches 1-4)
   - Patch 1: Rename NPU3 firmware filenames to match official release
     binaries.
   - Patch 2: Remove userspace doorbell mmap and default to kernel
     submission.
   - Patch 3: Add CERT (Column Embedded Run Time) firmware version
     queries and protocol validation.
   - Patch 4: Upgrade firmware interface version to 6.0.

2. Device Operations, Telemetry & Power Modes (Patches 5-8)
   - Patch 5: Add NPU3 classic device operations, and remove device IDs
     0x1B0B and 0x1B0C.
   - Patch 6: Expose AIE and NPU firmware versions via GET_INFO ioctl.
   - Patch 7: Add power mode support, and restore cached user power mode.
   - Patch 8: Add clock metadata, DPM frequency table initialization,
     and hardware resource info queries.

3. Hardware Initialization & Transport Hooks (Patches 9-11)
   - Patch 9: Add context switch hysteresis configuration with debugfs
     control (ctx_switch_hysteresis_us) applied on hardware start.
   - Patch 10: Refactor AIE4 hardware initialization sequence into
     distinct modular phases (query_fw, config_fw, setup_aie) across
     PF, VF, and classic devices.
   - Patch 11: Decouple PCI doorbell and MSI-X notification transport
     hooks from transport-neutral context management code.

4. Kernel-Mode Command Submission & Fencing (Patches 12-14)
   - Patch 12: Implement kernel queue lifecycle, job workqueue, and
     memory layout for kernel-mode submission.
   - Patch 13: Prepare command submission structures.
   - Patch 14: Implement full command packet building (direct/indirect
     packets).

5. System & Runtime Power Management (Patches 15-18)
   - Patch 15: Fix runtime PM deadlock on device removal by finalizing
     RPM before acquiring dev_lock, and symmetrically initialize RPM
     during probe across all device types.
   - Patch 16: Implement system suspend and resume for PF, VF, and
     classic devices, restoring contexts.
   - Patch 17: Link SR-IOV VFs via device_link_add() to ensure the PM
     core suspends VFs before the PF and resumes the PF before VFs.
   - Patch 18: Implement runtime suspend and resume support.

6. Userspace ABI Compatibility (Patch 19)
   - Patch 19: Add a stub hwctx_config callback returning 0 to support
     the DRM_AMDXDNA_CONFIG_HWCTX ioctl, enabling userspace validation
     runtimes (such as XRT GEMM tests) to run unmodified.

7. Firmware Logging (Patch 20)
   - Patch 20: Allocate a DRAM buffer for firmware logging on PF and
     classic devices and register it at hardware start, with the log
     level set to the least verbose setting.

Testing:
- Tested on AMD AIE4/NPU3 hardware in both classic and SR-IOV (PF/VF)
  modes.
- Verified kernel-mode command submission with direct and indirect
  execution packets under concurrent workloads.
- Verified system suspend/resume (S2idle/S3) and runtime autosuspend
  cycles during idle and active command submission.
- Verified SR-IOV VF binding, execution, and PM dependency sequencing.
- Confirmed no regression on existing AIE2 devices (NPU1/NPU4).

David Zhang (20):
  accel/amdxdna: Rename NPU3 firmware files
  accel/amdxdna: Remove mmap for doorbell
  accel/amdxdna: Add CERT firmware version support
  accel/amdxdna: Upgrade firmware version to 6.0
  accel/amdxdna: Add NPU3 classic device support
  accel/amdxdna: Add AIE version query to aie4_get_info
  accel/amdxdna: Add get and set power_mode for AIE4
  accel/amdxdna: Add clock, DPM frequency, and resource info queries for
    AIE4
  accel/amdxdna: Add context switch hysteresis with debugfs control
  accel/amdxdna: Refactor AIE4 hardware initialization sequence
  accel/amdxdna: Decouple AIE4 doorbell and MSI-X notification transport
    hooks
  accel/amdxdna: Implement AIE4 kernel queue lifecycle and memory layout
  accel/amdxdna: Prepare for AIE4 command submission
  accel/amdxdna: Implement AIE4 command packet building and submission
  accel/amdxdna: Finalize runtime PM before acquiring dev_lock on
    removal
  accel/amdxdna: Implement AIE4 suspend and resume
  accel/amdxdna: Link SR-IOV VFs for power management sequencing
  accel/amdxdna: Implement runtime suspend and resume support
  accel/amdxdna: Add stub hwctx_config for AIE4
  accel/amdxdna: Enable AIE4 firmware logging to DRAM

 drivers/accel/amdxdna/aie.c             |   45 +-
 drivers/accel/amdxdna/aie.h             |   45 +-
 drivers/accel/amdxdna/aie2_message.c    |    4 +-
 drivers/accel/amdxdna/aie2_pci.c        |   65 +-
 drivers/accel/amdxdna/aie2_pci.h        |   39 +-
 drivers/accel/amdxdna/aie2_pm.c         |   10 +-
 drivers/accel/amdxdna/aie4_ctx.c        | 1086 +++++++++++++++++++++--
 drivers/accel/amdxdna/aie4_host_queue.h |   79 +-
 drivers/accel/amdxdna/aie4_message.c    |  237 +++++
 drivers/accel/amdxdna/aie4_msg_priv.h   |  161 +++-
 drivers/accel/amdxdna/aie4_pci.c        |  867 +++++++++++++++++-
 drivers/accel/amdxdna/aie4_pci.h        |  143 ++-
 drivers/accel/amdxdna/aie4_sriov.c      |  104 ++-
 drivers/accel/amdxdna/amdxdna_ctx.c     |   20 +-
 drivers/accel/amdxdna/amdxdna_ctx.h     |   22 +
 drivers/accel/amdxdna/amdxdna_debugfs.c |    3 +
 drivers/accel/amdxdna/amdxdna_pci_drv.c |   53 +-
 drivers/accel/amdxdna/amdxdna_pci_drv.h |   18 +-
 drivers/accel/amdxdna/amdxdna_pm.c      |   35 +
 drivers/accel/amdxdna/amdxdna_pm.h      |    4 +-
 drivers/accel/amdxdna/amdxdna_sysfs.c   |    2 +-
 drivers/accel/amdxdna/npu1_regs.c       |   19 +-
 drivers/accel/amdxdna/npu3_regs.c       |   99 ++-
 drivers/accel/amdxdna/npu4_regs.c       |   30 +-
 24 files changed, 2907 insertions(+), 283 deletions(-)

-- 
2.34.1


^ permalink raw reply	[flat|nested] 33+ messages in thread

* [PATCH V1 01/20] accel/amdxdna: Rename NPU3 firmware files
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:32 ` [PATCH V1 02/20] accel/amdxdna: Remove mmap for doorbell David Zhang
                   ` (18 subsequent siblings)
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

The NPU3 firmware files will be released as npu.sbin and cert.sbin.
Rename the current firmware names to match the firmware release names,
and add MODULE_FIRMWARE() declarations for the new 17f2_10 npu.sbin and
cert.sbin names.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/amdxdna_pci_drv.c | 2 ++
 drivers/accel/amdxdna/npu3_regs.c       | 4 ++--
 2 files changed, 4 insertions(+), 2 deletions(-)

diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.c b/drivers/accel/amdxdna/amdxdna_pci_drv.c
index bb339e641416..d9e2e71d3e05 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.c
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.c
@@ -28,6 +28,8 @@ MODULE_FIRMWARE("amdnpu/17f0_20/npu.sbin");
 MODULE_FIRMWARE("amdnpu/1502_00/npu_7.sbin");
 MODULE_FIRMWARE("amdnpu/17f0_10/npu_7.sbin");
 MODULE_FIRMWARE("amdnpu/17f0_11/npu_7.sbin");
+MODULE_FIRMWARE("amdnpu/17f2_10/npu.sbin");
+MODULE_FIRMWARE("amdnpu/17f2_10/cert.sbin");
 
 /*
  * 0.0: Initial version
diff --git a/drivers/accel/amdxdna/npu3_regs.c b/drivers/accel/amdxdna/npu3_regs.c
index d76b2e99c308..8d287ef32fff 100644
--- a/drivers/accel/amdxdna/npu3_regs.c
+++ b/drivers/accel/amdxdna/npu3_regs.c
@@ -43,8 +43,8 @@ static const struct amdxdna_fw_feature_tbl npu3_fw_feature_table[] = {
 };
 
 static const struct amdxdna_dev_priv npu3_dev_priv = {
-	.npufw_path             = "npu.dev.sbin",
-	.certfw_path            = "cert.dev.sbin",
+	.npufw_path             = "npu.sbin",
+	.certfw_path            = "cert.sbin",
 	.mbox_bar		= NPU3_MBOX_BAR,
 	.mbox_rbuf_bar		= NPU3_MBOX_BUFFER_BAR,
 	.mbox_info_off		= NPU3_MBOX_INFO_OFF,
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 02/20] accel/amdxdna: Remove mmap for doorbell
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
  2026-09-30  3:32 ` [PATCH V1 01/20] accel/amdxdna: Rename NPU3 firmware files David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:32 ` [PATCH V1 03/20] accel/amdxdna: Add CERT firmware version support David Zhang
                   ` (17 subsequent siblings)
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

Make kernel submission the default, so mapping the doorbell back to user
space is not needed:
- Remove .mmap handler and use standard drm_gem_mmap.
- Set hwctx->doorbell_offset to AMDXDNA_INVALID_DOORBELL_OFFSET on
  context creation so userspace does not receive a valid-looking BAR
  offset.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_ctx.c        | 20 +--------------
 drivers/accel/amdxdna/aie4_pci.c        | 33 -------------------------
 drivers/accel/amdxdna/aie4_pci.h        |  1 -
 drivers/accel/amdxdna/amdxdna_pci_drv.c | 17 +------------
 drivers/accel/amdxdna/amdxdna_pci_drv.h |  1 -
 5 files changed, 2 insertions(+), 70 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
index 8408b0d2696f..8157f2a6fd10 100644
--- a/drivers/accel/amdxdna/aie4_ctx.c
+++ b/drivers/accel/amdxdna/aie4_ctx.c
@@ -158,7 +158,7 @@ static int aie4_hwctx_create(struct amdxdna_hwctx *hwctx)
 	}
 
 	priv->hw_ctx_id = resp.hw_context_id;
-	hwctx->doorbell_offset = resp.doorbell_offset;
+	hwctx->doorbell_offset = AMDXDNA_INVALID_DOORBELL_OFFSET;
 
 	return 0;
 }
@@ -313,21 +313,3 @@ int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout)
 
 	return ret <= 0 ? ret : 0;
 }
-
-int aie4_hwctx_valid_doorbell(struct amdxdna_client *client, u32 vm_pgoff)
-{
-	struct amdxdna_hwctx *hwctx;
-	unsigned long hwctx_id;
-	int idx;
-
-	idx = srcu_read_lock(&client->hwctx_srcu);
-	amdxdna_for_each_hwctx(client, hwctx_id, hwctx) {
-		if (vm_pgoff == (hwctx->doorbell_offset >> PAGE_SHIFT)) {
-			srcu_read_unlock(&client->hwctx_srcu, idx);
-			return 1;
-		}
-	}
-	srcu_read_unlock(&client->hwctx_srcu, idx);
-
-	return 0;
-}
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index a58a83af42a4..db02d25e3f4a 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -518,38 +518,6 @@ static int aie4m_pcidev_init(struct amdxdna_dev *xdna)
 	return 0;
 }
 
-static int aie4_doorbell_mmap(struct amdxdna_client *client, struct vm_area_struct *vma)
-{
-	struct amdxdna_dev *xdna = client->xdna;
-	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
-	const struct amdxdna_dev_priv *npriv = xdna->dev_info->dev_priv;
-	phys_addr_t res_start;
-	unsigned long pfn;
-	int ret;
-
-	if (!aie4_hwctx_valid_doorbell(client, vma->vm_pgoff)) {
-		XDNA_ERR(xdna, "Invalid doorbell page offset 0x%lx", vma->vm_pgoff);
-		return -EINVAL;
-	}
-
-	if (vma_pages(vma) != 1) {
-		XDNA_ERR(xdna, "can only map one page, got %ld", vma_pages(vma));
-		return -EINVAL;
-	}
-
-	res_start = pci_resource_start(pdev, xdna->dev_info->doorbell_bar) + npriv->doorbell_off;
-	pfn = PHYS_PFN(res_start) + vma->vm_pgoff;
-	vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);
-	vm_flags_set(vma, VM_IO | VM_DONTEXPAND | VM_DONTDUMP);
-	ret = io_remap_pfn_range(vma, vma->vm_start,
-				 pfn,
-				 PAGE_SIZE,
-				 vma->vm_page_prot);
-
-	XDNA_DBG(xdna, "doorbell ret %d", ret);
-	return ret;
-}
-
 static int aie4_get_info(struct amdxdna_client *client, struct amdxdna_drm_get_info *args)
 {
 	struct amdxdna_dev *xdna = client->xdna;
@@ -661,7 +629,6 @@ const struct amdxdna_dev_ops aie4_vf_ops = {
 	.fini			= aie4_vf_fini,
 	.hwctx_init		= aie4_hwctx_init,
 	.hwctx_fini		= aie4_hwctx_fini,
-	.mmap			= aie4_doorbell_mmap,
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
 };
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index 3fd5eace3ed7..c6219544dc0f 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -69,7 +69,6 @@ int aie4_attach_work_buffer(struct amdxdna_dev_hdl *ndev);
 int aie4_hwctx_init(struct amdxdna_hwctx *hwctx);
 void aie4_hwctx_fini(struct amdxdna_hwctx *hwctx);
 int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout);
-int aie4_hwctx_valid_doorbell(struct amdxdna_client *client, u32 vm_pgoff);
 
 /* aie4_sriov.c */
 #if IS_ENABLED(CONFIG_PCI_IOV)
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.c b/drivers/accel/amdxdna/amdxdna_pci_drv.c
index d9e2e71d3e05..3140af69e29c 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.c
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.c
@@ -250,21 +250,6 @@ static int amdxdna_drm_set_state_ioctl(struct drm_device *dev, void *data, struc
 	return ret;
 }
 
-static int amdxdna_drm_gem_mmap(struct file *filp, struct vm_area_struct *vma)
-{
-	struct drm_file *drm_filp = filp->private_data;
-	struct amdxdna_client *client = drm_filp->driver_priv;
-	struct amdxdna_dev *xdna = client->xdna;
-
-	if (likely(vma->vm_pgoff >= DRM_FILE_PAGE_OFFSET_START))
-		return drm_gem_mmap(filp, vma);
-
-	if (!xdna->dev_info->ops->mmap)
-		return -EOPNOTSUPP;
-
-	return xdna->dev_info->ops->mmap(client, vma);
-}
-
 static const struct drm_ioctl_desc amdxdna_drm_ioctls[] = {
 	/* Context */
 	DRM_IOCTL_DEF_DRV(AMDXDNA_CREATE_HWCTX, amdxdna_drm_create_hwctx_ioctl, 0),
@@ -323,7 +308,7 @@ static const struct file_operations amdxdna_fops = {
 	.poll		= drm_poll,
 	.read		= drm_read,
 	.llseek		= noop_llseek,
-	.mmap		= amdxdna_drm_gem_mmap,
+	.mmap		= drm_gem_mmap,
 	.show_fdinfo	= drm_show_fdinfo,
 	.fop_flags	= FOP_UNSIGNED_OFFSET,
 };
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.h b/drivers/accel/amdxdna/amdxdna_pci_drv.h
index a997d27a504d..84c8973e9197 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.h
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.h
@@ -57,7 +57,6 @@ struct amdxdna_dev_ops {
 	int (*resume)(struct amdxdna_dev *xdna);
 	int (*suspend)(struct amdxdna_dev *xdna);
 	int (*sriov_configure)(struct amdxdna_dev *xdna, int num_vfs);
-	int (*mmap)(struct amdxdna_client *client, struct vm_area_struct *vma);
 	int (*hwctx_init)(struct amdxdna_hwctx *hwctx);
 	void (*hwctx_fini)(struct amdxdna_hwctx *hwctx);
 	int (*hwctx_config)(struct amdxdna_hwctx *hwctx, u32 type, u64 value, void *buf, u32 size);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 03/20] accel/amdxdna: Add CERT firmware version support
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
  2026-09-30  3:32 ` [PATCH V1 01/20] accel/amdxdna: Rename NPU3 firmware files David Zhang
  2026-09-30  3:32 ` [PATCH V1 02/20] accel/amdxdna: Remove mmap for doorbell David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:53   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 04/20] accel/amdxdna: Upgrade firmware version to 6.0 David Zhang
                   ` (16 subsequent siblings)
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

AIE4 platforms run two separate firmware binaries: NPU firmware for
management and CERT (Column Embedded Run Time) firmware for
handling execution contexts and queues.

Add support to query and validate CERT firmware version:
- Add mailbox opcodes and structs to query NPU firmware version (identify)
  and CERT firmware version.
- Unify firmware version storage by using struct
  amdxdna_drm_query_firmware_version across the driver.
- Introduce aie_check_cert_protocol() and cert_feature_tbl to validate
  CERT firmware host queue protocol compatibility against driver
  capabilities.
- Add helper functions amdxdna_get_firmware_version() and
  amdxdna_get_aie_version() to share version query handling across
  generations.

Note on patch ordering:
Introducing CERT firmware protocol validation prior to the firmware 6.0
upgrade ensures host queue protocol compatibility (host_queue_major/minor)
is validated before the host queue layout restructure, preserving
bisectability.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie.c             | 45 ++++++++++++++++++++---
 drivers/accel/amdxdna/aie.h             |  9 ++++-
 drivers/accel/amdxdna/aie2_message.c    |  4 +--
 drivers/accel/amdxdna/aie2_pci.c        | 47 +++----------------------
 drivers/accel/amdxdna/aie2_pci.h        |  4 +--
 drivers/accel/amdxdna/aie4_message.c    | 45 +++++++++++++++++++++++
 drivers/accel/amdxdna/aie4_msg_priv.h   | 30 ++++++++++++++++
 drivers/accel/amdxdna/aie4_pci.c        | 17 ++++++++-
 drivers/accel/amdxdna/aie4_pci.h        | 11 ++++++
 drivers/accel/amdxdna/amdxdna_pci_drv.h | 10 ++----
 drivers/accel/amdxdna/amdxdna_sysfs.c   |  2 +-
 drivers/accel/amdxdna/npu3_regs.c       |  8 +++++
 12 files changed, 170 insertions(+), 62 deletions(-)

diff --git a/drivers/accel/amdxdna/aie.c b/drivers/accel/amdxdna/aie.c
index dd6f36f222c7..01a439c0ccf4 100644
--- a/drivers/accel/amdxdna/aie.c
+++ b/drivers/accel/amdxdna/aie.c
@@ -65,13 +65,12 @@ int aie_send_mgmt_msg_wait(struct aie_device *aie, struct xdna_mailbox_msg *msg)
 	return ret;
 }
 
-int aie_check_protocol(struct aie_device *aie, u32 fw_major, u32 fw_minor)
+static int aie_check_protocol_impl(struct aie_device *aie, u32 fw_major, u32 fw_minor,
+				   const struct amdxdna_fw_feature_tbl *feature)
 {
-	const struct amdxdna_fw_feature_tbl *feature;
 	bool found = false;
 
-	for (feature = aie->xdna->dev_info->fw_feature_tbl;
-	     feature->major; feature++) {
+	for (; feature && feature->major; feature++) {
 		if (feature->major != fw_major)
 			continue;
 		if (fw_minor < feature->min_minor)
@@ -88,6 +87,44 @@ int aie_check_protocol(struct aie_device *aie, u32 fw_major, u32 fw_minor)
 	return found ? 0 : -EOPNOTSUPP;
 }
 
+int aie_check_protocol(struct aie_device *aie, u32 fw_major, u32 fw_minor)
+{
+	return aie_check_protocol_impl(aie, fw_major, fw_minor,
+				       aie->xdna->dev_info->fw_feature_tbl);
+}
+
+int aie_check_cert_protocol(struct aie_device *aie, u32 cert_major, u32 cert_minor)
+{
+	return aie_check_protocol_impl(aie, cert_major, cert_minor,
+				       aie->xdna->dev_info->cert_feature_tbl);
+}
+
+int amdxdna_get_aie_version(struct amdxdna_client *client,
+			    struct amdxdna_drm_get_info *args,
+			    struct amdxdna_drm_query_aie_version *version)
+{
+	u32 buf_sz;
+
+	buf_sz = min_t(u32, args->buffer_size, sizeof(*version));
+	if (copy_to_user(u64_to_user_ptr(args->buffer), version, buf_sz))
+		return -EFAULT;
+
+	return 0;
+}
+
+int amdxdna_get_firmware_version(struct amdxdna_client *client,
+				 struct amdxdna_drm_get_info *args,
+				 struct amdxdna_drm_query_firmware_version *version)
+{
+	u32 buf_sz;
+
+	buf_sz = min_t(u32, args->buffer_size, sizeof(*version));
+	if (copy_to_user(u64_to_user_ptr(args->buffer), version, buf_sz))
+		return -EFAULT;
+
+	return 0;
+}
+
 static void amdxdna_update_vbnv(struct amdxdna_dev *xdna,
 				const struct amdxdna_rev_vbnv *tbl,
 				u32 rev)
diff --git a/drivers/accel/amdxdna/aie.h b/drivers/accel/amdxdna/aie.h
index 0483d582b7f8..899399756661 100644
--- a/drivers/accel/amdxdna/aie.h
+++ b/drivers/accel/amdxdna/aie.h
@@ -28,6 +28,7 @@ struct aie_device {
 	struct psp_device *psp_hdl;
 	struct smu_device *smu_hdl;
 
+	struct amdxdna_drm_query_aie_version version;
 	struct amdxdna_drm_query_aie_metadata metadata;
 };
 
@@ -96,6 +97,7 @@ void aie_dump_mgmt_chann_debug(struct aie_device *aie);
 void aie_destroy_chann(struct aie_device *aie, struct mailbox_channel **chann);
 int aie_send_mgmt_msg_wait(struct aie_device *aie, struct xdna_mailbox_msg *msg);
 int aie_check_protocol(struct aie_device *aie, u32 fw_major, u32 fw_minor);
+int aie_check_cert_protocol(struct aie_device *aie, u32 fw_major, u32 fw_minor);
 void amdxdna_vbnv_init(struct amdxdna_dev *xdna);
 int amdxdna_get_metadata(struct aie_device *aie, struct amdxdna_client *client,
 			 struct amdxdna_drm_get_info *args);
@@ -103,7 +105,12 @@ void *amdxdna_alloc_msg_buffer(struct amdxdna_dev *xdna, u32 *size,
 			       dma_addr_t *dma_addr);
 void amdxdna_free_msg_buffer(struct amdxdna_dev *xdna, size_t size,
 			     void *cpu_addr, dma_addr_t dma_addr);
-
+int amdxdna_get_aie_version(struct amdxdna_client *client,
+			    struct amdxdna_drm_get_info *args,
+			    struct amdxdna_drm_query_aie_version *version);
+int amdxdna_get_firmware_version(struct amdxdna_client *client,
+				 struct amdxdna_drm_get_info *args,
+				 struct amdxdna_drm_query_firmware_version *version);
 /* aie_psp.c */
 struct psp_device *aiem_psp_create(struct drm_device *ddev, struct psp_config *conf);
 int aie_psp_start(struct psp_device *psp);
diff --git a/drivers/accel/amdxdna/aie2_message.c b/drivers/accel/amdxdna/aie2_message.c
index f658760c3d48..bae0cc4c3580 100644
--- a/drivers/accel/amdxdna/aie2_message.c
+++ b/drivers/accel/amdxdna/aie2_message.c
@@ -149,7 +149,7 @@ int aie2_query_aie_metadata(struct amdxdna_dev_hdl *ndev,
 }
 
 int aie2_query_firmware_version(struct amdxdna_dev_hdl *ndev,
-				struct amdxdna_fw_ver *fw_ver)
+				struct amdxdna_drm_query_firmware_version *fw_ver)
 {
 	DECLARE_AIE_MSG(firmware_version, MSG_OP_GET_FIRMWARE_VERSION);
 	int ret;
@@ -160,7 +160,7 @@ int aie2_query_firmware_version(struct amdxdna_dev_hdl *ndev,
 
 	fw_ver->major = resp.major;
 	fw_ver->minor = resp.minor;
-	fw_ver->sub = resp.sub;
+	fw_ver->patch = resp.sub;
 	fw_ver->build = resp.build;
 
 	return 0;
diff --git a/drivers/accel/amdxdna/aie2_pci.c b/drivers/accel/amdxdna/aie2_pci.c
index 7a4314ca843b..5dc6e5b97afc 100644
--- a/drivers/accel/amdxdna/aie2_pci.c
+++ b/drivers/accel/amdxdna/aie2_pci.c
@@ -205,15 +205,16 @@ static int aie2_mgmt_fw_init(struct amdxdna_dev_hdl *ndev)
 
 static int aie2_mgmt_fw_query(struct amdxdna_dev_hdl *ndev)
 {
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
 	int ret;
 
-	ret = aie2_query_firmware_version(ndev, &ndev->aie.xdna->fw_ver);
+	ret = aie2_query_firmware_version(ndev, &xdna->fw_ver);
 	if (ret) {
 		XDNA_ERR(ndev->aie.xdna, "query firmware version failed");
 		return ret;
 	}
 
-	ret = aie2_query_aie_version(ndev, &ndev->version);
+	ret = aie2_query_aie_version(ndev, &ndev->aie.version);
 	if (ret) {
 		XDNA_ERR(ndev->aie.xdna, "Query AIE version failed");
 		return ret;
@@ -673,44 +674,6 @@ static int aie2_get_aie_status(struct amdxdna_client *client,
 	return 0;
 }
 
-static int aie2_get_aie_version(struct amdxdna_client *client,
-				struct amdxdna_drm_get_info *args)
-{
-	struct amdxdna_drm_query_aie_version version;
-	struct amdxdna_dev *xdna = client->xdna;
-	struct amdxdna_dev_hdl *ndev;
-	u32 buf_sz;
-
-	ndev = xdna->dev_handle;
-	version.major = ndev->version.major;
-	version.minor = ndev->version.minor;
-
-	buf_sz = min(args->buffer_size, sizeof(version));
-	if (copy_to_user(u64_to_user_ptr(args->buffer), &version, buf_sz))
-		return -EFAULT;
-
-	return 0;
-}
-
-static int aie2_get_firmware_version(struct amdxdna_client *client,
-				     struct amdxdna_drm_get_info *args)
-{
-	struct amdxdna_drm_query_firmware_version version;
-	struct amdxdna_dev *xdna = client->xdna;
-	u32 buf_sz;
-
-	version.major = xdna->fw_ver.major;
-	version.minor = xdna->fw_ver.minor;
-	version.patch = xdna->fw_ver.sub;
-	version.build = xdna->fw_ver.build;
-
-	buf_sz = min(args->buffer_size, sizeof(version));
-	if (copy_to_user(u64_to_user_ptr(args->buffer), &version, buf_sz))
-		return -EFAULT;
-
-	return 0;
-}
-
 static int aie2_get_power_mode(struct amdxdna_client *client,
 			       struct amdxdna_drm_get_info *args)
 {
@@ -1025,7 +988,7 @@ static int aie2_get_info(struct amdxdna_client *client, struct amdxdna_drm_get_i
 		ret = amdxdna_get_metadata(&ndev->aie, client, args);
 		break;
 	case DRM_AMDXDNA_QUERY_AIE_VERSION:
-		ret = aie2_get_aie_version(client, args);
+		ret = amdxdna_get_aie_version(client, args, &ndev->aie.version);
 		break;
 	case DRM_AMDXDNA_QUERY_CLOCK_METADATA:
 		ret = aie2_get_clock_metadata(client, args);
@@ -1037,7 +1000,7 @@ static int aie2_get_info(struct amdxdna_client *client, struct amdxdna_drm_get_i
 		ret = aie2_get_hwctx_status(client, args);
 		break;
 	case DRM_AMDXDNA_QUERY_FIRMWARE_VERSION:
-		ret = aie2_get_firmware_version(client, args);
+		ret = amdxdna_get_firmware_version(client, args, &xdna->fw_ver);
 		break;
 	case DRM_AMDXDNA_GET_POWER_MODE:
 		ret = aie2_get_power_mode(client, args);
diff --git a/drivers/accel/amdxdna/aie2_pci.h b/drivers/accel/amdxdna/aie2_pci.h
index 2c7019bd26b5..67971f0c4acf 100644
--- a/drivers/accel/amdxdna/aie2_pci.h
+++ b/drivers/accel/amdxdna/aie2_pci.h
@@ -74,7 +74,6 @@ enum aie2_sram_reg_idx {
 };
 
 struct amdxdna_client;
-struct amdxdna_fw_ver;
 struct amdxdna_hwctx;
 struct amdxdna_sched_job;
 
@@ -150,7 +149,6 @@ struct amdxdna_dev_hdl {
 	void			__iomem *mbox_base;
 
 	u32				total_col;
-	struct amdxdna_drm_query_aie_version version;
 	struct aie2_exec_msg_ops	*exec_msg_ops;
 	struct drm_gpu_scheduler	*hwctx_sched;
 	struct ida			hwctx_sched_ida;
@@ -263,7 +261,7 @@ int aie2_query_aie_version(struct amdxdna_dev_hdl *ndev,
 int aie2_query_aie_metadata(struct amdxdna_dev_hdl *ndev,
 			    struct amdxdna_drm_query_aie_metadata *metadata);
 int aie2_query_firmware_version(struct amdxdna_dev_hdl *ndev,
-				struct amdxdna_fw_ver *fw_ver);
+				struct amdxdna_drm_query_firmware_version *fw_ver);
 int aie2_query_app_health(struct amdxdna_dev_hdl *ndev, u32 context_id,
 			  struct app_health_report *report);
 int aie2_get_dev_revision(struct amdxdna_dev_hdl *ndev, enum aie2_dev_revision *rev);
diff --git a/drivers/accel/amdxdna/aie4_message.c b/drivers/accel/amdxdna/aie4_message.c
index 88037edbb02a..b137a2a40b34 100644
--- a/drivers/accel/amdxdna/aie4_message.c
+++ b/drivers/accel/amdxdna/aie4_message.c
@@ -64,6 +64,51 @@ int aie4_query_aie_metadata(struct amdxdna_dev_hdl *ndev,
 	return 0;
 }
 
+int aie4_query_npu_firmware_version(struct amdxdna_dev_hdl *ndev,
+				    struct amdxdna_drm_query_firmware_version *fw_version)
+{
+	DECLARE_AIE_MSG(aie4_msg_identify, AIE4_MSG_OP_IDENTIFY);
+	int ret;
+
+	ret = aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	if (ret)
+		return ret;
+
+	fw_version->major = resp.fw_major;
+	fw_version->minor = resp.fw_minor;
+	fw_version->patch = resp.fw_patch;
+	fw_version->build = resp.fw_build;
+
+	return 0;
+}
+
+int aie4_query_cert_firmware_version(struct amdxdna_dev_hdl *ndev,
+				     struct amdxdna_drm_query_firmware_version *cert_version)
+{
+	DECLARE_AIE_MSG(aie4_msg_query_cert_firmware_version,
+			AIE4_MSG_OP_QUERY_CERT_FIRMWARE_VERSION);
+	int ret;
+
+	ret = aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	if (ret)
+		return ret;
+
+	ret = aie_check_cert_protocol(&ndev->aie,
+				      resp.host_queue_major, resp.host_queue_minor);
+	if (ret) {
+		XDNA_ERR(ndev->aie.xdna, "host queue %d.%d is not supported",
+			 resp.host_queue_major, resp.host_queue_minor);
+		return ret;
+	}
+
+	cert_version->major = resp.major_version;
+	cert_version->minor = resp.minor_version;
+	cert_version->patch = resp.hotfix;
+	cert_version->build = resp.build;
+
+	return 0;
+}
+
 int aie4_attach_work_buffer(struct amdxdna_dev_hdl *ndev)
 {
 	DECLARE_AIE_MSG(aie4_msg_attach_work_buffer, AIE4_MSG_OP_ATTACH_WORK_BUFFER);
diff --git a/drivers/accel/amdxdna/aie4_msg_priv.h b/drivers/accel/amdxdna/aie4_msg_priv.h
index af0866045b91..5b97c8057de0 100644
--- a/drivers/accel/amdxdna/aie4_msg_priv.h
+++ b/drivers/accel/amdxdna/aie4_msg_priv.h
@@ -10,8 +10,10 @@
 #include <linux/types.h>
 
 enum aie4_msg_opcode {
+	AIE4_MSG_OP_IDENTIFY                         = 0x10002,
 	AIE4_MSG_OP_SUSPEND                          = 0x10003,
 	AIE4_MSG_OP_ATTACH_WORK_BUFFER               = 0x1000D,
+	AIE4_MSG_OP_QUERY_CERT_FIRMWARE_VERSION      = 0x1000F,
 
 	AIE4_MSG_OP_CREATE_VFS                       = 0x20001,
 	AIE4_MSG_OP_DESTROY_VFS                      = 0x20002,
@@ -30,6 +32,18 @@ enum aie4_msg_status {
 	MAX_AIE4_MSG_STATUS_CODE = 0x4,
 };
 
+struct aie4_msg_identify_req {
+	__u32 rsvd;
+} __packed;
+
+struct aie4_msg_identify_resp {
+	enum aie4_msg_status status;
+	__u32 fw_major;
+	__u32 fw_minor;
+	__u32 fw_patch;
+	__u32 fw_build;
+} __packed;
+
 struct aie4_msg_suspend_req {
 	__u32 rsvd;
 } __packed;
@@ -132,6 +146,22 @@ struct aie4_msg_aie4_tile_info_resp {
 	struct aie4_tile_info info;
 } __packed;
 
+struct aie4_msg_query_cert_firmware_version_req {
+	__u32 resvd;
+} __packed;
+
+struct aie4_msg_query_cert_firmware_version_resp {
+	enum aie4_msg_status status;
+	__u8 major_version;
+	__u8 minor_version;
+	__u8 git_hash[41];
+	__u8 date[11];
+	__u8 hotfix;
+	__u8 build;
+	__u16 host_queue_major;
+	__u16 host_queue_minor;
+} __packed;
+
 #define AIE4_WORK_BUFFER_MIN_SIZE      SZ_4M
 
 struct aie4_msg_attach_work_buffer_req {
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index db02d25e3f4a..3cb81bc1b627 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -271,7 +271,22 @@ static void aie4_partition_fini(struct amdxdna_dev_hdl *ndev)
 
 static int aie4_query(struct amdxdna_dev_hdl *ndev)
 {
-	return aie4_query_aie_metadata(ndev, &ndev->aie.metadata);
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	int ret;
+
+	ret = aie4_query_npu_firmware_version(ndev, &xdna->fw_ver);
+	if (ret)
+		return ret;
+
+	ret = aie4_query_cert_firmware_version(ndev, &ndev->cert_version);
+	if (ret)
+		return ret;
+
+	ret = aie4_query_aie_metadata(ndev, &ndev->aie.metadata);
+	if (ret)
+		return ret;
+
+	return 0;
 }
 
 static int aie4_pf_hw_start(struct amdxdna_dev_hdl *ndev)
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index c6219544dc0f..8c62ee6a9b23 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -57,6 +57,13 @@ struct amdxdna_dev_hdl {
 	void				*work_buf;
 	dma_addr_t			work_buf_addr;
 	u32				work_buf_size;
+
+	struct amdxdna_drm_query_firmware_version cert_version;
+};
+
+enum aie4_fw_feature {
+	AIE4_HSA_COMMAND = 5,
+	AIE4_FEATURE_MAX
 };
 
 /* aie4_message.c */
@@ -64,6 +71,10 @@ int aie4_query_aie_metadata(struct amdxdna_dev_hdl *ndev,
 			    struct amdxdna_drm_query_aie_metadata *metadata);
 int aie4_suspend_fw(struct amdxdna_dev_hdl *ndev);
 int aie4_attach_work_buffer(struct amdxdna_dev_hdl *ndev);
+int aie4_query_npu_firmware_version(struct amdxdna_dev_hdl *ndev,
+				    struct amdxdna_drm_query_firmware_version *fw_version);
+int aie4_query_cert_firmware_version(struct amdxdna_dev_hdl *ndev,
+				     struct amdxdna_drm_query_firmware_version *cert_version);
 
 /* aie4_ctx.c */
 int aie4_hwctx_init(struct amdxdna_hwctx *hwctx);
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.h b/drivers/accel/amdxdna/amdxdna_pci_drv.h
index 84c8973e9197..0002e6ef32ba 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.h
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.h
@@ -99,16 +99,10 @@ struct amdxdna_dev_info {
 	size_t				dev_heap_max_size;
 	const struct amdxdna_dev_priv	*dev_priv;
 	const struct amdxdna_fw_feature_tbl *fw_feature_tbl;
+	const struct amdxdna_fw_feature_tbl *cert_feature_tbl;
 	const struct amdxdna_dev_ops	*ops;
 };
 
-struct amdxdna_fw_ver {
-	u32 major;
-	u32 minor;
-	u32 sub;
-	u32 build;
-};
-
 struct amdxdna_carveout;
 
 struct amdxdna_dev {
@@ -120,7 +114,7 @@ struct amdxdna_dev {
 	struct mutex			dev_lock; /* per device lock */
 	struct list_head		client_list;
 	struct mutex			client_lock; /* client_list */
-	struct amdxdna_fw_ver		fw_ver;
+	struct amdxdna_drm_query_firmware_version fw_ver;
 	struct rw_semaphore		notifier_lock; /* for mmu notifier*/
 	struct workqueue_struct		*notifier_wq;
 
diff --git a/drivers/accel/amdxdna/amdxdna_sysfs.c b/drivers/accel/amdxdna/amdxdna_sysfs.c
index d9e359ee8182..e20b7fb1e5d1 100644
--- a/drivers/accel/amdxdna/amdxdna_sysfs.c
+++ b/drivers/accel/amdxdna/amdxdna_sysfs.c
@@ -37,7 +37,7 @@ static ssize_t fw_version_show(struct device *dev, struct device_attribute *attr
 	struct amdxdna_dev *xdna = dev_get_drvdata(dev);
 
 	return sprintf(buf, "%d.%d.%d.%d\n", xdna->fw_ver.major,
-		       xdna->fw_ver.minor, xdna->fw_ver.sub,
+		       xdna->fw_ver.minor, xdna->fw_ver.patch,
 		       xdna->fw_ver.build);
 }
 static DEVICE_ATTR_RO(fw_version);
diff --git a/drivers/accel/amdxdna/npu3_regs.c b/drivers/accel/amdxdna/npu3_regs.c
index 8d287ef32fff..31208c42ad5f 100644
--- a/drivers/accel/amdxdna/npu3_regs.c
+++ b/drivers/accel/amdxdna/npu3_regs.c
@@ -42,6 +42,12 @@ static const struct amdxdna_fw_feature_tbl npu3_fw_feature_table[] = {
 	{ 0 }
 };
 
+static const struct amdxdna_fw_feature_tbl npu3_cert_feature_table[] = {
+	{ .major = 1, .min_minor = 0 },
+	{ .features = BIT_U64(AIE4_HSA_COMMAND), .major = 1, .min_minor = 0 },
+	{ 0 }
+};
+
 static const struct amdxdna_dev_priv npu3_dev_priv = {
 	.npufw_path             = "npu.sbin",
 	.certfw_path            = "cert.sbin",
@@ -85,6 +91,7 @@ const struct amdxdna_dev_info dev_npu3_pf_info = {
 	.device_type		= AMDXDNA_DEV_TYPE_PF,
 	.dev_priv		= &npu3_dev_priv,
 	.fw_feature_tbl		= npu3_fw_feature_table,
+	.cert_feature_tbl	= npu3_cert_feature_table,
 	.ops			= &aie4_pf_ops,
 };
 
@@ -96,5 +103,6 @@ const struct amdxdna_dev_info dev_npu3_vf_info = {
 	.device_type		= AMDXDNA_DEV_TYPE_UMQ,
 	.dev_priv		= &npu3_dev_vf_priv,
 	.fw_feature_tbl		= npu3_fw_feature_table,
+	.cert_feature_tbl	= npu3_cert_feature_table,
 	.ops			= &aie4_vf_ops,
 };
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 04/20] accel/amdxdna: Upgrade firmware version to 6.0
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (2 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 03/20] accel/amdxdna: Add CERT firmware version support David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:32 ` [PATCH V1 05/20] accel/amdxdna: Add NPU3 classic device support David Zhang
                   ` (15 subsequent siblings)
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

Upgrade firmware interface version to 6.0. Update host queue layout,
opcode definitions, and context creation/destruction request structures.
Parse priority band and PASID for hardware context creation. Protocol
compatibility for this queue layout is validated against the CERT
firmware protocol version.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_ctx.c        | 23 ++++++++++++++++++++---
 drivers/accel/amdxdna/aie4_host_queue.h | 14 ++++++++++++--
 drivers/accel/amdxdna/aie4_message.c    | 11 +++++++++++
 drivers/accel/amdxdna/aie4_msg_priv.h   | 20 +++++++++++++++++---
 drivers/accel/amdxdna/aie4_pci.h        |  1 +
 drivers/accel/amdxdna/npu3_regs.c       |  2 +-
 6 files changed, 62 insertions(+), 9 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
index 8157f2a6fd10..5eb918e1d58c 100644
--- a/drivers/accel/amdxdna/aie4_ctx.c
+++ b/drivers/accel/amdxdna/aie4_ctx.c
@@ -9,6 +9,7 @@
 #include <drm/drm_gem_shmem_helper.h>
 #include <drm/drm_print.h>
 #include <drm/gpu_scheduler.h>
+#include <linux/iommu.h>
 #include <linux/types.h>
 
 #include "aie.h"
@@ -110,6 +111,22 @@ static int aie4_msg_destroy_context(struct amdxdna_dev_hdl *ndev, u32 hw_context
 	return aie_send_mgmt_msg_wait(&ndev->aie, &msg);
 }
 
+static u8 aie4_parse_priority_to_dev(u32 priority)
+{
+	switch (priority) {
+	case AMDXDNA_QOS_LOW_PRIORITY:
+		return AIE4_CONTEXT_PRIORITY_BAND_IDLE;
+	case AMDXDNA_QOS_NORMAL_PRIORITY:
+		return AIE4_CONTEXT_PRIORITY_BAND_NORMAL;
+	case AMDXDNA_QOS_HIGH_PRIORITY:
+		return AIE4_CONTEXT_PRIORITY_BAND_FOCUS;
+	case AMDXDNA_QOS_REALTIME_PRIORITY:
+		return AIE4_CONTEXT_PRIORITY_BAND_REAL_TIME;
+	default:
+		return AIE4_CONTEXT_PRIORITY_BAND_NORMAL;
+	}
+}
+
 static int aie4_hwctx_create(struct amdxdna_hwctx *hwctx)
 {
 	DECLARE_AIE_MSG(aie4_msg_create_hw_context, AIE4_MSG_OP_CREATE_HW_CONTEXT);
@@ -129,9 +146,9 @@ static int aie4_hwctx_create(struct amdxdna_hwctx *hwctx)
 
 	req.partition_id = ndev->partition_id;
 	req.request_num_tiles = hwctx->num_tiles;
-	req.pasid = FIELD_PREP(AIE4_MSG_PASID, client->pasid) |
-		FIELD_PREP(AIE4_MSG_PASID_VLD, 1);
-	req.priority_band = hwctx->qos.priority;
+	req.pasid = aie4_msg_pasid(client);
+	req.pasid = req.pasid == IOMMU_PASID_INVALID ? 0 : req.pasid;
+	req.priority_band = aie4_parse_priority_to_dev(hwctx->qos.priority);
 
 	req.hsa_addr_high = upper_32_bits(amdxdna_gem_dev_addr(priv->umq_bo));
 	req.hsa_addr_low = lower_32_bits(amdxdna_gem_dev_addr(priv->umq_bo));
diff --git a/drivers/accel/amdxdna/aie4_host_queue.h b/drivers/accel/amdxdna/aie4_host_queue.h
index 1b33eda3f727..97e535939b32 100644
--- a/drivers/accel/amdxdna/aie4_host_queue.h
+++ b/drivers/accel/amdxdna/aie4_host_queue.h
@@ -10,6 +10,14 @@
 
 #define CTX_MAX_CMDS                    32
 
+/*
+ * Host queue header layout.
+ *
+ * Note: Compatibility for this layout is checked against the CERT firmware
+ * protocol version (host_queue_major/minor) via aie_check_cert_protocol(),
+ * introduced in the preceding patch ("accel/amdxdna: Add CERT firmware
+ * version support").
+ */
 struct host_queue_header {
 	__u64 read_index;
 	struct {
@@ -17,8 +25,10 @@ struct host_queue_header {
 		__u16 minor;
 	} version;
 	__u32 capacity; /* Queue capacity, must be power of two. */
-	__u64 write_index;
+	__u64 padding0[6];
+	__u64 write_index; /* different cacheline from read_index to avoid false sharing */
+	__u64 padding1[6];
 	__u64 data_address; /* The xdna dev addr for payload. */
-};
+} __packed;
 
 #endif /* _AIE4_HOST_QUEUE_H_ */
diff --git a/drivers/accel/amdxdna/aie4_message.c b/drivers/accel/amdxdna/aie4_message.c
index b137a2a40b34..bdbd1d116b61 100644
--- a/drivers/accel/amdxdna/aie4_message.c
+++ b/drivers/accel/amdxdna/aie4_message.c
@@ -5,6 +5,8 @@
 
 #include <drm/amdxdna_accel.h>
 #include <drm/drm_print.h>
+#include <linux/bitfield.h>
+#include <linux/iommu.h>
 #include <linux/mutex.h>
 
 #include "aie.h"
@@ -14,6 +16,15 @@
 #include "amdxdna_mailbox_helper.h"
 #include "amdxdna_pci_drv.h"
 
+u32 aie4_msg_pasid(struct amdxdna_client *client)
+{
+	if (!amdxdna_pasid_on(client))
+		return IOMMU_PASID_INVALID;
+
+	return FIELD_PREP(AIE4_MSG_PASID, client->pasid) |
+	       FIELD_PREP(AIE4_MSG_PASID_VLD, 1);
+}
+
 int aie4_suspend_fw(struct amdxdna_dev_hdl *ndev)
 {
 	DECLARE_AIE_MSG(aie4_msg_suspend, AIE4_MSG_OP_SUSPEND);
diff --git a/drivers/accel/amdxdna/aie4_msg_priv.h b/drivers/accel/amdxdna/aie4_msg_priv.h
index 5b97c8057de0..b9f7c61f36e3 100644
--- a/drivers/accel/amdxdna/aie4_msg_priv.h
+++ b/drivers/accel/amdxdna/aie4_msg_priv.h
@@ -12,7 +12,6 @@
 enum aie4_msg_opcode {
 	AIE4_MSG_OP_IDENTIFY                         = 0x10002,
 	AIE4_MSG_OP_SUSPEND                          = 0x10003,
-	AIE4_MSG_OP_ATTACH_WORK_BUFFER               = 0x1000D,
 	AIE4_MSG_OP_QUERY_CERT_FIRMWARE_VERSION      = 0x1000F,
 
 	AIE4_MSG_OP_CREATE_VFS                       = 0x20001,
@@ -23,6 +22,8 @@ enum aie4_msg_opcode {
 	AIE4_MSG_OP_CREATE_HW_CONTEXT                = 0x30003,
 	AIE4_MSG_OP_DESTROY_HW_CONTEXT               = 0x30004,
 	AIE4_MSG_OP_AIE_TILE_INFO                    = 0x30006,
+
+	AIE4_MSG_OP_ATTACH_WORK_BUFFER               = 0x40001,
 };
 
 enum aie4_msg_status {
@@ -32,6 +33,14 @@ enum aie4_msg_status {
 	MAX_AIE4_MSG_STATUS_CODE = 0x4,
 };
 
+enum aie4_msg_context_priority_band {
+	AIE4_CONTEXT_PRIORITY_BAND_IDLE = 0,
+	AIE4_CONTEXT_PRIORITY_BAND_NORMAL,
+	AIE4_CONTEXT_PRIORITY_BAND_FOCUS,
+	AIE4_CONTEXT_PRIORITY_BAND_REAL_TIME,
+	AIE4_CONTEXT_PRIORITY_BAND_COUNT
+};
+
 struct aie4_msg_identify_req {
 	__u32 rsvd;
 } __packed;
@@ -94,7 +103,9 @@ struct aie4_msg_create_hw_context_req {
 #define AIE4_MSG_PASID GENMASK(19, 0)
 #define AIE4_MSG_PASID_VLD GENMASK(31, 31)
 	__u32 pasid;
-	__u32 priority_band;
+	__u8 priority_band;
+	__u8 priority_level;
+	__u16 restore_id;
 } __packed;
 
 struct aie4_msg_create_hw_context_resp {
@@ -106,11 +117,14 @@ struct aie4_msg_create_hw_context_resp {
 
 struct aie4_msg_destroy_hw_context_req {
 	__u32 hw_context_id;
-	__u32 resvd1;
+#define AIE4_MSG_GRACEFUL_FLAG GENMASK(0, 0)
+	__u32 graceful_flag;
 } __packed;
 
 struct aie4_msg_destroy_hw_context_resp {
 	enum aie4_msg_status status;
+	__u16 restore_id;
+	__u16 resvd;
 } __packed;
 
 struct aie4_tile_info {
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index 8c62ee6a9b23..bdbb2d7cf0e7 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -75,6 +75,7 @@ int aie4_query_npu_firmware_version(struct amdxdna_dev_hdl *ndev,
 				    struct amdxdna_drm_query_firmware_version *fw_version);
 int aie4_query_cert_firmware_version(struct amdxdna_dev_hdl *ndev,
 				     struct amdxdna_drm_query_firmware_version *cert_version);
+u32 aie4_msg_pasid(struct amdxdna_client *client);
 
 /* aie4_ctx.c */
 int aie4_hwctx_init(struct amdxdna_hwctx *hwctx);
diff --git a/drivers/accel/amdxdna/npu3_regs.c b/drivers/accel/amdxdna/npu3_regs.c
index 31208c42ad5f..891c5f243ae5 100644
--- a/drivers/accel/amdxdna/npu3_regs.c
+++ b/drivers/accel/amdxdna/npu3_regs.c
@@ -38,7 +38,7 @@
 #define MP1_C2PMSG_60_ALT_1     0x3B109F0
 
 static const struct amdxdna_fw_feature_tbl npu3_fw_feature_table[] = {
-	{ .major = 5, .min_minor = 10 },
+	{ .major = 6, .min_minor = 0 },
 	{ 0 }
 };
 
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 05/20] accel/amdxdna: Add NPU3 classic device support
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (3 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 04/20] accel/amdxdna: Upgrade firmware version to 6.0 David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:32 ` [PATCH V1 06/20] accel/amdxdna: Add AIE version query to aie4_get_info David Zhang
                   ` (14 subsequent siblings)
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

Add NPU3 classic device operations (aie4_classic_ops) and register
definitions, as well as the PCI device ID and firmware declarations for
NPU3 classic devices. Remove device IDs 0x1B0B and 0x1B0C because those
devices will not be released.

Note: System suspend/resume support requires draining and managing
in-flight commands by leveraging the kernel-mode queue (KMQ) infra.
Full suspend and resume callbacks for all AIE4 device types are
introduced in a subsequent patch with KMQ support.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_pci.c        | 83 +++++++++++++++++++++++++
 drivers/accel/amdxdna/aie4_pci.h        |  1 +
 drivers/accel/amdxdna/amdxdna_pci_drv.c | 15 +++--
 drivers/accel/amdxdna/amdxdna_pci_drv.h |  1 +
 drivers/accel/amdxdna/npu3_regs.c       | 14 +++++
 5 files changed, 110 insertions(+), 4 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index 3cb81bc1b627..880619e3a4a8 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -359,6 +359,51 @@ static void aie4_vf_hw_stop(struct amdxdna_dev_hdl *ndev)
 	aie4_mailbox_fini(ndev);
 }
 
+static int aie4_classic_hw_start(struct amdxdna_dev_hdl *ndev)
+{
+	int ret;
+
+	ret = aie4_fw_start(ndev);
+	if (ret)
+		return ret;
+
+	ret = aie4_mailbox_init(ndev);
+	if (ret)
+		goto stop_fw;
+
+	ret = aie4_query(ndev);
+	if (ret)
+		goto mailbox_fini;
+
+	ret = aie4_attach_work_buffer(ndev);
+	if (ret)
+		goto mailbox_fini;
+
+	ret = aie4_partition_init(ndev);
+	if (ret)
+		goto mailbox_fini;
+
+	return 0;
+
+mailbox_fini:
+	aie4_mailbox_fini(ndev);
+stop_fw:
+	aie4_fw_stop(ndev);
+	return ret;
+}
+
+static void aie4_classic_hw_stop(struct amdxdna_dev_hdl *ndev)
+{
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+
+	aie4_partition_fini(ndev);
+	aie4_suspend_fw(ndev);
+	aie4_mailbox_fini(ndev);
+	aie4_fw_stop(ndev);
+}
+
 static int aie4_request_firmware(struct amdxdna_dev_hdl *ndev,
 				 const struct firmware **npufw,
 				 const struct firmware **certfw)
@@ -621,6 +666,29 @@ static int aie4_vf_init(struct amdxdna_dev *xdna)
 	return aie4_vf_hw_start(xdna->dev_handle);
 }
 
+static int aie4_classic_init(struct amdxdna_dev *xdna)
+{
+	int ret;
+
+	ret = aie4m_pcidev_init(xdna);
+	if (ret)
+		return ret;
+
+	ret = aie4_alloc_work_buffer(xdna->dev_handle);
+	if (ret)
+		return ret;
+
+	ret = aie4_classic_hw_start(xdna->dev_handle);
+	if (ret)
+		goto free_work_buf;
+
+	return 0;
+
+free_work_buf:
+	aie4_free_work_buffer(xdna->dev_handle);
+	return ret;
+}
+
 static void aie4_pf_fini(struct amdxdna_dev *xdna)
 {
 	aie4_sriov_stop(xdna->dev_handle);
@@ -633,6 +701,12 @@ static void aie4_vf_fini(struct amdxdna_dev *xdna)
 	aie4_vf_hw_stop(xdna->dev_handle);
 }
 
+static void aie4_classic_fini(struct amdxdna_dev *xdna)
+{
+	aie4_classic_hw_stop(xdna->dev_handle);
+	aie4_free_work_buffer(xdna->dev_handle);
+}
+
 const struct amdxdna_dev_ops aie4_pf_ops = {
 	.init			= aie4_pf_init,
 	.fini			= aie4_pf_fini,
@@ -647,3 +721,12 @@ const struct amdxdna_dev_ops aie4_vf_ops = {
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
 };
+
+const struct amdxdna_dev_ops aie4_classic_ops = {
+	.init			= aie4_classic_init,
+	.fini			= aie4_classic_fini,
+	.hwctx_init		= aie4_hwctx_init,
+	.hwctx_fini		= aie4_hwctx_fini,
+	.cmd_wait		= aie4_cmd_wait,
+	.get_aie_info		= aie4_get_info,
+};
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index bdbb2d7cf0e7..940e67347d74 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -96,5 +96,6 @@ static inline int aie4_sriov_stop(struct amdxdna_dev_hdl *ndev)
 
 extern const struct amdxdna_dev_ops aie4_pf_ops;
 extern const struct amdxdna_dev_ops aie4_vf_ops;
+extern const struct amdxdna_dev_ops aie4_classic_ops;
 
 #endif /* _AIE4_PCI_H_ */
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.c b/drivers/accel/amdxdna/amdxdna_pci_drv.c
index 3140af69e29c..8b6e7283e057 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.c
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.c
@@ -28,8 +28,14 @@ MODULE_FIRMWARE("amdnpu/17f0_20/npu.sbin");
 MODULE_FIRMWARE("amdnpu/1502_00/npu_7.sbin");
 MODULE_FIRMWARE("amdnpu/17f0_10/npu_7.sbin");
 MODULE_FIRMWARE("amdnpu/17f0_11/npu_7.sbin");
+MODULE_FIRMWARE("amdnpu/17f1_10/npu.sbin");
+MODULE_FIRMWARE("amdnpu/17f1_10/cert.sbin");
 MODULE_FIRMWARE("amdnpu/17f2_10/npu.sbin");
 MODULE_FIRMWARE("amdnpu/17f2_10/cert.sbin");
+MODULE_FIRMWARE("amdnpu/17f1_13/npu.sbin");
+MODULE_FIRMWARE("amdnpu/17f1_13/cert.sbin");
+MODULE_FIRMWARE("amdnpu/17f2_13/npu.sbin");
+MODULE_FIRMWARE("amdnpu/17f2_13/cert.sbin");
 
 /*
  * 0.0: Initial version
@@ -55,10 +61,9 @@ MODULE_FIRMWARE("amdnpu/17f2_10/cert.sbin");
 static const struct pci_device_id pci_ids[] = {
 	{ PCI_DEVICE(PCI_VENDOR_ID_AMD, 0x1502) },
 	{ PCI_DEVICE(PCI_VENDOR_ID_AMD, 0x17f0) },
+	{ PCI_DEVICE(PCI_VENDOR_ID_AMD, 0x17f1) },
 	{ PCI_DEVICE(PCI_VENDOR_ID_AMD, 0x17f2) },
 	{ PCI_DEVICE(PCI_VENDOR_ID_AMD, 0x17f3) },
-	{ PCI_DEVICE(PCI_VENDOR_ID_AMD, 0x1B0B) },
-	{ PCI_DEVICE(PCI_VENDOR_ID_AMD, 0x1B0C) },
 	{0}
 };
 
@@ -69,10 +74,12 @@ static const struct amdxdna_device_id amdxdna_ids[] = {
 	{ 0x17f0, 0x10, &dev_npu4_info },
 	{ 0x17f0, 0x11, &dev_npu5_info },
 	{ 0x17f0, 0x20, &dev_npu6_info },
+	{ 0x17f1, 0x10, &dev_npu3_classic_info },
 	{ 0x17f2, 0x10, &dev_npu3_pf_info },
 	{ 0x17f3, 0x10, &dev_npu3_vf_info },
-	{ 0x1B0B, 0x10, &dev_npu3_pf_info },
-	{ 0x1B0C, 0x10, &dev_npu3_vf_info },
+	{ 0x17f1, 0x13, &dev_npu3_classic_info },
+	{ 0x17f2, 0x13, &dev_npu3_pf_info },
+	{ 0x17f3, 0x13, &dev_npu3_vf_info },
 	{0}
 };
 
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.h b/drivers/accel/amdxdna/amdxdna_pci_drv.h
index 0002e6ef32ba..953bf783b3f7 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.h
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.h
@@ -169,6 +169,7 @@ struct amdxdna_client {
 
 /* Add device info below */
 extern const struct amdxdna_dev_info dev_npu1_info;
+extern const struct amdxdna_dev_info dev_npu3_classic_info;
 extern const struct amdxdna_dev_info dev_npu3_pf_info;
 extern const struct amdxdna_dev_info dev_npu3_vf_info;
 extern const struct amdxdna_dev_info dev_npu4_info;
diff --git a/drivers/accel/amdxdna/npu3_regs.c b/drivers/accel/amdxdna/npu3_regs.c
index 891c5f243ae5..21e24901976c 100644
--- a/drivers/accel/amdxdna/npu3_regs.c
+++ b/drivers/accel/amdxdna/npu3_regs.c
@@ -106,3 +106,17 @@ const struct amdxdna_dev_info dev_npu3_vf_info = {
 	.cert_feature_tbl	= npu3_cert_feature_table,
 	.ops			= &aie4_vf_ops,
 };
+
+const struct amdxdna_dev_info dev_npu3_classic_info = {
+	.mbox_bar		= NPU3_MBOX_BAR,
+	.sram_bar		= NPU3_MBOX_BUFFER_BAR,
+	.psp_bar                = NPU3_PSP_BAR_INDEX,
+	.smu_bar		= NPU3_SMU_BAR_INDEX,
+	.doorbell_bar		= NPU3_DOORBELL_BAR,
+	.default_vbnv		= "RyzenAI-npu3",
+	.device_type		= AMDXDNA_DEV_TYPE_UMQ,
+	.dev_priv		= &npu3_dev_priv,
+	.fw_feature_tbl		= npu3_fw_feature_table,
+	.cert_feature_tbl	= npu3_cert_feature_table,
+	.ops			= &aie4_classic_ops,
+};
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 06/20] accel/amdxdna: Add AIE version query to aie4_get_info
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (4 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 05/20] accel/amdxdna: Add NPU3 classic device support David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:32 ` [PATCH V1 07/20] accel/amdxdna: Add get and set power_mode for AIE4 David Zhang
                   ` (13 subsequent siblings)
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello, Hayden Laccabue

Add AIE version query via DRM_AMDXDNA_GET_INFO. The NPU firmware
version query already existed internally; it is now also exposed
through this ioctl alongside the new AIE version query.

These versions are mandatory once released, the firmware management
interfaces must support those queries in their stable version.
Propagating any failure to abort device initialization is intentional.

Co-developed-by: Hayden Laccabue <hayden.laccabue@amd.com>
Signed-off-by: Hayden Laccabue <hayden.laccabue@amd.com>
Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_message.c  | 20 ++++++++++++++++++++
 drivers/accel/amdxdna/aie4_msg_priv.h | 11 +++++++++++
 drivers/accel/amdxdna/aie4_pci.c      | 10 ++++++++++
 drivers/accel/amdxdna/aie4_pci.h      |  2 ++
 4 files changed, 43 insertions(+)

diff --git a/drivers/accel/amdxdna/aie4_message.c b/drivers/accel/amdxdna/aie4_message.c
index bdbd1d116b61..4814cbb1f2c5 100644
--- a/drivers/accel/amdxdna/aie4_message.c
+++ b/drivers/accel/amdxdna/aie4_message.c
@@ -75,6 +75,26 @@ int aie4_query_aie_metadata(struct amdxdna_dev_hdl *ndev,
 	return 0;
 }
 
+int aie4_query_aie_version(struct amdxdna_dev_hdl *ndev,
+			   struct amdxdna_drm_query_aie_version *version)
+{
+	DECLARE_AIE_MSG(aie4_msg_aie4_version_info, AIE4_MSG_OP_AIE_VERSION_INFO);
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	int ret;
+
+	ret = aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	if (ret)
+		return ret;
+
+	XDNA_DBG(xdna, "Query AIE version - major: %u minor: %u",
+		 resp.major, resp.minor);
+
+	version->major = resp.major;
+	version->minor = resp.minor;
+
+	return 0;
+}
+
 int aie4_query_npu_firmware_version(struct amdxdna_dev_hdl *ndev,
 				    struct amdxdna_drm_query_firmware_version *fw_version)
 {
diff --git a/drivers/accel/amdxdna/aie4_msg_priv.h b/drivers/accel/amdxdna/aie4_msg_priv.h
index b9f7c61f36e3..81842fa1d6ce 100644
--- a/drivers/accel/amdxdna/aie4_msg_priv.h
+++ b/drivers/accel/amdxdna/aie4_msg_priv.h
@@ -22,6 +22,7 @@ enum aie4_msg_opcode {
 	AIE4_MSG_OP_CREATE_HW_CONTEXT                = 0x30003,
 	AIE4_MSG_OP_DESTROY_HW_CONTEXT               = 0x30004,
 	AIE4_MSG_OP_AIE_TILE_INFO                    = 0x30006,
+	AIE4_MSG_OP_AIE_VERSION_INFO                 = 0x30007,
 
 	AIE4_MSG_OP_ATTACH_WORK_BUFFER               = 0x40001,
 };
@@ -160,6 +161,16 @@ struct aie4_msg_aie4_tile_info_resp {
 	struct aie4_tile_info info;
 } __packed;
 
+struct aie4_msg_aie4_version_info_req {
+	__u32 resvd;
+} __packed;
+
+struct aie4_msg_aie4_version_info_resp {
+	enum aie4_msg_status status;
+	__u16 major;
+	__u16 minor;
+} __packed;
+
 struct aie4_msg_query_cert_firmware_version_req {
 	__u32 resvd;
 } __packed;
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index 880619e3a4a8..aea3edd51b4f 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -282,6 +282,10 @@ static int aie4_query(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		return ret;
 
+	ret = aie4_query_aie_version(ndev, &ndev->aie.version);
+	if (ret)
+		return ret;
+
 	ret = aie4_query_aie_metadata(ndev, &ndev->aie.metadata);
 	if (ret)
 		return ret;
@@ -588,6 +592,12 @@ static int aie4_get_info(struct amdxdna_client *client, struct amdxdna_drm_get_i
 	case DRM_AMDXDNA_QUERY_AIE_METADATA:
 		ret = amdxdna_get_metadata(&ndev->aie, client, args);
 		break;
+	case DRM_AMDXDNA_QUERY_AIE_VERSION:
+		ret = amdxdna_get_aie_version(client, args, &ndev->aie.version);
+		break;
+	case DRM_AMDXDNA_QUERY_FIRMWARE_VERSION:
+		ret = amdxdna_get_firmware_version(client, args, &xdna->fw_ver);
+		break;
 	default:
 		XDNA_ERR(xdna, "Not supported request parameter %u", args->param);
 		ret = -EOPNOTSUPP;
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index 940e67347d74..5ae5e8427a3b 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -69,6 +69,8 @@ enum aie4_fw_feature {
 /* aie4_message.c */
 int aie4_query_aie_metadata(struct amdxdna_dev_hdl *ndev,
 			    struct amdxdna_drm_query_aie_metadata *metadata);
+int aie4_query_aie_version(struct amdxdna_dev_hdl *ndev,
+			   struct amdxdna_drm_query_aie_version *version);
 int aie4_suspend_fw(struct amdxdna_dev_hdl *ndev);
 int aie4_attach_work_buffer(struct amdxdna_dev_hdl *ndev);
 int aie4_query_npu_firmware_version(struct amdxdna_dev_hdl *ndev,
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 07/20] accel/amdxdna: Add get and set power_mode for AIE4
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (5 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 06/20] accel/amdxdna: Add AIE version query to aie4_get_info David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:56   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 08/20] accel/amdxdna: Add clock, DPM frequency, and resource info queries " David Zhang
                   ` (12 subsequent siblings)
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello, Soham Donwalkar,
	Nishad Saraf

Add power mode support for AIE4 devices via DRM_AMDXDNA_GET_INFO and
DRM_AMDXDNA_SET_STATE:
- Add AIE4_MSG_OP_POWER_OVERRIDE mailbox support
- Add DRM_AMDXDNA_GET_POWER_MODE support
- Implement aie4_set_power_mode() and wire .set_aie_state

Firmware boots in POWER_MODE_DEFAULT after a reload, so re-send the
cached user power mode override whenever the hardware starts via
aie4_restore_power_mode(). On a fresh probe, pw_mode is
POWER_MODE_DEFAULT.

Co-developed-by: Soham Donwalkar <soham.donwalkar@amd.com>
Signed-off-by: Soham Donwalkar <soham.donwalkar@amd.com>
Co-developed-by: Nishad Saraf <nishad.saraf@amd.com>
Signed-off-by: Nishad Saraf <nishad.saraf@amd.com>
Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_message.c  |  17 ++++
 drivers/accel/amdxdna/aie4_msg_priv.h |   9 ++
 drivers/accel/amdxdna/aie4_pci.c      | 124 ++++++++++++++++++++++++++
 drivers/accel/amdxdna/aie4_pci.h      |   6 ++
 4 files changed, 156 insertions(+)

diff --git a/drivers/accel/amdxdna/aie4_message.c b/drivers/accel/amdxdna/aie4_message.c
index 4814cbb1f2c5..25a510bd7419 100644
--- a/drivers/accel/amdxdna/aie4_message.c
+++ b/drivers/accel/amdxdna/aie4_message.c
@@ -157,3 +157,20 @@ int aie4_attach_work_buffer(struct amdxdna_dev_hdl *ndev)
 
 	return ret;
 }
+
+int aie4_msg_set_power_mode(struct amdxdna_dev_hdl *ndev, u8 power_mode)
+{
+	DECLARE_AIE_MSG(aie4_msg_power_override, AIE4_MSG_OP_POWER_OVERRIDE);
+	int ret;
+
+	req.power_mode = power_mode;
+
+	ret = aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	if (ret)
+		XDNA_WARN(ndev->aie.xdna,
+			  "Failed to set power mode %u, ret %d", (u32)power_mode, ret);
+	else
+		XDNA_DBG(ndev->aie.xdna, "Power mode set to %u", (u32)power_mode);
+
+	return ret;
+}
diff --git a/drivers/accel/amdxdna/aie4_msg_priv.h b/drivers/accel/amdxdna/aie4_msg_priv.h
index 81842fa1d6ce..4c06792df1bd 100644
--- a/drivers/accel/amdxdna/aie4_msg_priv.h
+++ b/drivers/accel/amdxdna/aie4_msg_priv.h
@@ -23,6 +23,7 @@ enum aie4_msg_opcode {
 	AIE4_MSG_OP_DESTROY_HW_CONTEXT               = 0x30004,
 	AIE4_MSG_OP_AIE_TILE_INFO                    = 0x30006,
 	AIE4_MSG_OP_AIE_VERSION_INFO                 = 0x30007,
+	AIE4_MSG_OP_POWER_OVERRIDE                   = 0x3000B,
 
 	AIE4_MSG_OP_ATTACH_WORK_BUFFER               = 0x40001,
 };
@@ -187,6 +188,14 @@ struct aie4_msg_query_cert_firmware_version_resp {
 	__u16 host_queue_minor;
 } __packed;
 
+struct aie4_msg_power_override_req {
+	__u32 power_mode;
+} __packed;
+
+struct aie4_msg_power_override_resp {
+	enum aie4_msg_status status;
+} __packed;
+
 #define AIE4_WORK_BUFFER_MIN_SIZE      SZ_4M
 
 struct aie4_msg_attach_work_buffer_req {
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index aea3edd51b4f..d11bdbf16881 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -4,6 +4,7 @@
  */
 
 #include <drm/amdxdna_accel.h>
+#include <drm/drm_drv.h>
 #include <drm/drm_managed.h>
 #include <drm/drm_print.h>
 #include <linux/firmware.h>
@@ -15,6 +16,7 @@
 #include "amdxdna_mailbox.h"
 #include "amdxdna_mailbox_helper.h"
 #include "amdxdna_pci_drv.h"
+#include "amdxdna_pm.h"
 
 #define NO_IOHUB		0
 #define PSP_NOTIFY_INTR		0xD007BE11
@@ -293,6 +295,27 @@ static int aie4_query(struct amdxdna_dev_hdl *ndev)
 	return 0;
 }
 
+/*
+ * Firmware always boots in POWER_MODE_DEFAULT after a (re)load, so re-send the
+ * cached user override whenever the hardware starts. This keeps the driver
+ * cache (ndev->pw_mode) and the firmware power state consistent across
+ * suspend/resume and runtime PM cycles. On a fresh probe pw_mode is
+ * POWER_MODE_DEFAULT and this is a no-op.
+ *
+ * Power override is a per-VF property in firmware: each supervisor (VF) stores
+ * its own requested mode and the hypervisor arbitrates globally by taking the
+ * highest mode across all supervisors. A full firmware reload on suspend clears
+ * every supervisor override back to default, so each device type (PF, VF and
+ * classic) must re-send its own cached override on resume.
+ */
+int aie4_restore_power_mode(struct amdxdna_dev_hdl *ndev)
+{
+	if (ndev->pw_mode == POWER_MODE_DEFAULT)
+		return 0;
+
+	return aie4_msg_set_power_mode(ndev, ndev->pw_mode);
+}
+
 static int aie4_pf_hw_start(struct amdxdna_dev_hdl *ndev)
 {
 	int ret;
@@ -309,6 +332,10 @@ static int aie4_pf_hw_start(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		goto mbox_fini;
 
+	ret = aie4_restore_power_mode(ndev);
+	if (ret)
+		goto mbox_fini;
+
 	return 0;
 
 mbox_fini:
@@ -346,6 +373,10 @@ static int aie4_vf_hw_start(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		goto mailbox_fini;
 
+	ret = aie4_restore_power_mode(ndev);
+	if (ret)
+		goto mailbox_fini;
+
 	return 0;
 
 mailbox_fini:
@@ -387,8 +418,14 @@ static int aie4_classic_hw_start(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		goto mailbox_fini;
 
+	ret = aie4_restore_power_mode(ndev);
+	if (ret)
+		goto partition_fini;
+
 	return 0;
 
+partition_fini:
+	aie4_partition_fini(ndev);
 mailbox_fini:
 	aie4_mailbox_fini(ndev);
 stop_fw:
@@ -528,6 +565,7 @@ static int aie4m_pcidev_init(struct amdxdna_dev *xdna)
 
 	ndev->priv = xdna->dev_info->dev_priv;
 	ndev->aie.xdna = xdna;
+	ndev->pw_mode = POWER_MODE_DEFAULT;
 	xdna->dev_handle = ndev;
 
 	xa_init_flags(&ndev->cert_comp_xa, XA_FLAGS_ALLOC);
@@ -582,6 +620,24 @@ static int aie4m_pcidev_init(struct amdxdna_dev *xdna)
 	return 0;
 }
 
+static int aie4_get_power_mode(struct amdxdna_client *client,
+			       struct amdxdna_drm_get_info *args)
+{
+	struct amdxdna_drm_get_power_mode mode = {};
+	struct amdxdna_dev *xdna = client->xdna;
+	struct amdxdna_dev_hdl *ndev;
+	u32 buf_sz;
+
+	ndev = xdna->dev_handle;
+	mode.power_mode = ndev->pw_mode;
+
+	buf_sz = min_t(u32, args->buffer_size, sizeof(mode));
+	if (copy_to_user(u64_to_user_ptr(args->buffer), &mode, buf_sz))
+		return -EFAULT;
+
+	return 0;
+}
+
 static int aie4_get_info(struct amdxdna_client *client, struct amdxdna_drm_get_info *args)
 {
 	struct amdxdna_dev *xdna = client->xdna;
@@ -598,6 +654,9 @@ static int aie4_get_info(struct amdxdna_client *client, struct amdxdna_drm_get_i
 	case DRM_AMDXDNA_QUERY_FIRMWARE_VERSION:
 		ret = amdxdna_get_firmware_version(client, args, &xdna->fw_ver);
 		break;
+	case DRM_AMDXDNA_GET_POWER_MODE:
+		ret = aie4_get_power_mode(client, args);
+		break;
 	default:
 		XDNA_ERR(xdna, "Not supported request parameter %u", args->param);
 		ret = -EOPNOTSUPP;
@@ -608,6 +667,69 @@ static int aie4_get_info(struct amdxdna_client *client, struct amdxdna_drm_get_i
 	return ret;
 }
 
+static int aie4_set_power_mode(struct amdxdna_client *client,
+			       struct amdxdna_drm_set_state *args)
+{
+	struct amdxdna_drm_set_power_mode power_state = { 0 };
+	struct amdxdna_dev *xdna = client->xdna;
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	u32 buf_sz;
+	u8 power_mode;
+	int ret;
+
+	buf_sz = min_t(u32, args->buffer_size, sizeof(power_state));
+	if (copy_from_user(&power_state, u64_to_user_ptr(args->buffer), buf_sz)) {
+		XDNA_ERR(xdna, "Failed to copy power mode request into kernel");
+		return -EFAULT;
+	}
+
+	if (XDNA_MBZ_DBG(xdna, power_state.pad, sizeof(power_state.pad)))
+		return -EINVAL;
+
+	power_mode = power_state.power_mode;
+	if (power_mode > POWER_MODE_TURBO) {
+		XDNA_ERR(xdna, "Invalid power mode %d", power_mode);
+		return -EINVAL;
+	}
+
+	ret = aie4_msg_set_power_mode(xdna->dev_handle, power_mode);
+	if (ret)
+		return ret;
+
+	ndev->pw_mode = power_mode;
+	return 0;
+}
+
+static int aie4_set_state(struct amdxdna_client *client,
+			  struct amdxdna_drm_set_state *args)
+{
+	struct amdxdna_dev *xdna = client->xdna;
+	int ret, idx;
+
+	if (!drm_dev_enter(&xdna->ddev, &idx))
+		return -ENODEV;
+
+	ret = amdxdna_pm_resume_get_locked(xdna);
+	if (ret)
+		goto dev_exit;
+
+	switch (args->param) {
+	case DRM_AMDXDNA_SET_POWER_MODE:
+		ret = aie4_set_power_mode(client, args);
+		break;
+	default:
+		XDNA_ERR(xdna, "Not supported request parameter %u", args->param);
+		ret = -EOPNOTSUPP;
+		break;
+	}
+
+	amdxdna_pm_suspend_put(xdna);
+
+dev_exit:
+	drm_dev_exit(idx);
+	return ret;
+}
+
 static int aie4_alloc_work_buffer(struct amdxdna_dev_hdl *ndev)
 {
 	struct amdxdna_dev *xdna = ndev->aie.xdna;
@@ -730,6 +852,7 @@ const struct amdxdna_dev_ops aie4_vf_ops = {
 	.hwctx_fini		= aie4_hwctx_fini,
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
+	.set_aie_state		= aie4_set_state,
 };
 
 const struct amdxdna_dev_ops aie4_classic_ops = {
@@ -739,4 +862,5 @@ const struct amdxdna_dev_ops aie4_classic_ops = {
 	.hwctx_fini		= aie4_hwctx_fini,
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
+	.set_aie_state		= aie4_set_state,
 };
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index 5ae5e8427a3b..fd2c50dc8080 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -58,6 +58,8 @@ struct amdxdna_dev_hdl {
 	dma_addr_t			work_buf_addr;
 	u32				work_buf_size;
 
+	u8				pw_mode;
+
 	struct amdxdna_drm_query_firmware_version cert_version;
 };
 
@@ -77,6 +79,7 @@ int aie4_query_npu_firmware_version(struct amdxdna_dev_hdl *ndev,
 				    struct amdxdna_drm_query_firmware_version *fw_version);
 int aie4_query_cert_firmware_version(struct amdxdna_dev_hdl *ndev,
 				     struct amdxdna_drm_query_firmware_version *cert_version);
+int aie4_msg_set_power_mode(struct amdxdna_dev_hdl *ndev, u8 power_mode);
 u32 aie4_msg_pasid(struct amdxdna_client *client);
 
 /* aie4_ctx.c */
@@ -84,6 +87,9 @@ int aie4_hwctx_init(struct amdxdna_hwctx *hwctx);
 void aie4_hwctx_fini(struct amdxdna_hwctx *hwctx);
 int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout);
 
+/* aie4_pci.c */
+int aie4_restore_power_mode(struct amdxdna_dev_hdl *ndev);
+
 /* aie4_sriov.c */
 #if IS_ENABLED(CONFIG_PCI_IOV)
 int aie4_sriov_configure(struct amdxdna_dev *xdna, int num_vfs);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 08/20] accel/amdxdna: Add clock, DPM frequency, and resource info queries for AIE4
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (6 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 07/20] accel/amdxdna: Add get and set power_mode for AIE4 David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:32 ` [PATCH V1 09/20] accel/amdxdna: Add context switch hysteresis with debugfs control David Zhang
                   ` (11 subsequent siblings)
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello, Soham Donwalkar

Add support for querying clock metadata, DPM frequency table, and
hardware resource information for AIE4/NPU3 devices:
- Move struct dpm_clk_freq and common clock/TOPS counters into struct
  aie_device.
- Add NPU3 DPM clock table, DPM control, and counter updates by querying
  active DPM levels via AIE4_MSG_OP_GET_CURRENT_DPM_LEVEL.
- Query AIE4 DPM frequency table from firmware via
  AIE4_MSG_OP_GET_DPM_FREQ_TABLE.

Co-developed-by: Soham Donwalkar <soham.donwalkar@amd.com>
Signed-off-by: Soham Donwalkar <soham.donwalkar@amd.com>
Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie.h           | 36 ++++++++++++
 drivers/accel/amdxdna/aie2_pci.c      | 14 ++---
 drivers/accel/amdxdna/aie2_pci.h      | 35 +----------
 drivers/accel/amdxdna/aie2_pm.c       | 10 ++--
 drivers/accel/amdxdna/aie4_message.c  | 74 +++++++++++++++++++++++
 drivers/accel/amdxdna/aie4_msg_priv.h | 31 ++++++++++
 drivers/accel/amdxdna/aie4_pci.c      | 84 ++++++++++++++++++++++++++-
 drivers/accel/amdxdna/aie4_pci.h      | 12 ++++
 drivers/accel/amdxdna/npu1_regs.c     | 19 +++---
 drivers/accel/amdxdna/npu3_regs.c     | 71 ++++++++++++++++++++++
 drivers/accel/amdxdna/npu4_regs.c     | 30 +++++-----
 11 files changed, 348 insertions(+), 68 deletions(-)

diff --git a/drivers/accel/amdxdna/aie.h b/drivers/accel/amdxdna/aie.h
index 899399756661..6268b708d17b 100644
--- a/drivers/accel/amdxdna/aie.h
+++ b/drivers/accel/amdxdna/aie.h
@@ -30,8 +30,44 @@ struct aie_device {
 
 	struct amdxdna_drm_query_aie_version version;
 	struct amdxdna_drm_query_aie_metadata metadata;
+
+	u32 clk_gating;
+	u32 npuclk_freq;
+	u32 hclk_freq;
+	u32 max_tops;
+	u32 curr_tops;
+};
+
+struct aie_hw_ops {
+	int (*set_dpm)(struct aie_device *aie, u32 dpm_level);
+	int (*update_counters)(struct aie_device *aie);
 };
 
+#define aie_update_counters(ndev)					\
+({									\
+	typeof(ndev) _ndev = ndev;					\
+	if ((_ndev)->priv->hw_ops && (_ndev)->priv->hw_ops->update_counters) \
+		(_ndev)->priv->hw_ops->update_counters(&(_ndev)->aie);	\
+})
+
+struct dpm_clk_freq {
+	u32	npuclk;
+	u32	hclk;
+};
+
+#include <linux/amd-pmf-io.h>
+
+#if IS_ENABLED(CONFIG_AMD_PMF)
+#define AIE_GET_PMF_NPU_METRICS(metrics) amd_pmf_get_npu_data(metrics)
+#else
+#define AIE_GET_PMF_NPU_METRICS(metrics)				\
+({									\
+	typeof(metrics) _m = metrics;					\
+	memset(_m, 0xff, sizeof(*_m));					\
+	(-EOPNOTSUPP);							\
+})
+#endif
+
 #define DECLARE_AIE_MSG(name, op) \
 	DECLARE_XDNA_MSG_COMMON(name, op, -1)
 #define AIE_FEATURE_ON(aie, feature) test_bit(feature, &(aie)->feature_mask)
diff --git a/drivers/accel/amdxdna/aie2_pci.c b/drivers/accel/amdxdna/aie2_pci.c
index 5dc6e5b97afc..b70af1923643 100644
--- a/drivers/accel/amdxdna/aie2_pci.c
+++ b/drivers/accel/amdxdna/aie2_pci.c
@@ -294,7 +294,7 @@ static struct xrs_action_ops aie2_xrs_actions = {
 
 static void aie2_smu_fini(struct amdxdna_dev_hdl *ndev)
 {
-	ndev->priv->hw_ops->set_dpm(ndev, 0);
+	ndev->priv->hw_ops->set_dpm(&ndev->aie, 0);
 	aie_smu_fini(ndev->aie.smu_hdl);
 }
 
@@ -706,12 +706,12 @@ static int aie2_get_clock_metadata(struct amdxdna_client *client,
 	if (!clock)
 		return -ENOMEM;
 
-	aie2_update_counters(ndev);
+	aie_update_counters(ndev);
 	snprintf(clock->mp_npu_clock.name, sizeof(clock->mp_npu_clock.name),
 		 "MP-NPU Clock");
-	clock->mp_npu_clock.freq_mhz = ndev->npuclk_freq;
+	clock->mp_npu_clock.freq_mhz = ndev->aie.npuclk_freq;
 	snprintf(clock->h_clock.name, sizeof(clock->h_clock.name), "H Clock");
-	clock->h_clock.freq_mhz = ndev->hclk_freq;
+	clock->h_clock.freq_mhz = ndev->aie.hclk_freq;
 
 	buf_sz = min(args->buffer_size, sizeof(*clock));
 	if (copy_to_user(u64_to_user_ptr(args->buffer), clock, buf_sz))
@@ -867,11 +867,11 @@ static int aie2_query_resource_info(struct amdxdna_client *client,
 	ndev = xdna->dev_handle;
 	priv = ndev->priv;
 
-	aie2_update_counters(ndev);
+	aie_update_counters(ndev);
 	res_info.npu_clk_max = priv->dpm_clk_tbl[ndev->max_dpm_level].hclk;
-	res_info.npu_tops_max = ndev->max_tops;
+	res_info.npu_tops_max = ndev->aie.max_tops;
 	res_info.npu_task_max = priv->hwctx_limit;
-	res_info.npu_tops_curr = ndev->curr_tops;
+	res_info.npu_tops_curr = ndev->aie.curr_tops;
 	res_info.npu_task_curr = ndev->hwctx_num;
 
 	buf_sz = min(args->buffer_size, sizeof(res_info));
diff --git a/drivers/accel/amdxdna/aie2_pci.h b/drivers/accel/amdxdna/aie2_pci.h
index 67971f0c4acf..0c8dd6510292 100644
--- a/drivers/accel/amdxdna/aie2_pci.h
+++ b/drivers/accel/amdxdna/aie2_pci.h
@@ -40,8 +40,8 @@
 	pci_resource_len(NDEV2PDEV(_ndev), (_ndev)->aie.xdna->dev_info->mbox_bar); \
 })
 
+#define AIE2_GET_PMF_NPU_METRICS(metrics) AIE_GET_PMF_NPU_METRICS(metrics)
 #if IS_ENABLED(CONFIG_AMD_PMF)
-#define AIE2_GET_PMF_NPU_METRICS(metrics) amd_pmf_get_npu_data(metrics)
 #define AIE2_GET_PMF_NPU_DATA(field, val)				\
 ({									\
 	struct amd_pmf_npu_metrics _npu_metrics;			\
@@ -52,13 +52,6 @@
 	(_ret);								\
 })
 #else
-#define AIE2_GET_PMF_NPU_METRICS(metrics)				\
-({									\
-	typeof(metrics) _m = metrics;					\
-	memset(_m, 0xff, sizeof(*_m));					\
-	(-EOPNOTSUPP);							\
-})
-
 #define SENSOR_DEFAULT_npu_power	U32_MAX
 #define AIE2_GET_PMF_NPU_DATA(field, val)				\
 ({									\
@@ -91,11 +84,6 @@ struct rt_config {
 	unsigned long feature_mask;
 };
 
-struct dpm_clk_freq {
-	u32	npuclk;
-	u32	hclk;
-};
-
 /*
  * Define the maximum number of pending commands in a hardware context.
  * Must be power of 2!
@@ -158,11 +146,6 @@ struct amdxdna_dev_hdl {
 	u32				dpm_level;
 	u32				dft_dpm_level;
 	u32				max_dpm_level;
-	u32				clk_gating;
-	u32				npuclk_freq;
-	u32				hclk_freq;
-	u32				max_tops;
-	u32				curr_tops;
 	u32				force_preempt_enabled;
 	u32				frame_boundary_preempt;
 
@@ -177,18 +160,6 @@ struct amdxdna_dev_hdl {
 	unsigned long			last_signal_ts;
 };
 
-struct aie2_hw_ops {
-	int (*set_dpm)(struct amdxdna_dev_hdl *ndev, u32 dpm_level);
-	int (*update_counters)(struct amdxdna_dev_hdl *ndev);
-};
-
-#define aie2_update_counters(ndev)				\
-({								\
-	typeof(ndev) _ndev = ndev;				\
-	if (_ndev->priv->hw_ops->update_counters)		\
-		_ndev->priv->hw_ops->update_counters(_ndev);	\
-})
-
 enum aie2_fw_feature {
 	AIE2_NPU_COMMAND,
 	AIE2_PREEMPT,
@@ -219,7 +190,7 @@ struct amdxdna_dev_priv {
 	struct aie_bar_off_pair		sram_offs[SRAM_MAX_INDEX];
 	struct aie_bar_off_pair		psp_regs_off[PSP_MAX_REGS];
 	struct aie_bar_off_pair		smu_regs_off[SMU_MAX_REGS];
-	const struct aie2_hw_ops	*hw_ops;
+	const struct aie_hw_ops		*hw_ops;
 };
 
 extern const struct amdxdna_dev_ops aie2_ops;
@@ -234,7 +205,7 @@ extern const struct rt_config npu1_default_rt_cfg[];
 extern const struct rt_config npu4_default_rt_cfg[];
 extern const struct amdxdna_fw_feature_tbl npu4_fw_feature_table[];
 extern const struct amdxdna_rev_vbnv npu4_rev_vbnv_tbl[];
-extern const struct aie2_hw_ops npu4_hw_ops;
+extern const struct aie_hw_ops npu4_hw_ops;
 
 /* aie2_pm.c */
 int aie2_pm_init(struct amdxdna_dev_hdl *ndev);
diff --git a/drivers/accel/amdxdna/aie2_pm.c b/drivers/accel/amdxdna/aie2_pm.c
index 4fe6030d2c41..f4ced7b67c25 100644
--- a/drivers/accel/amdxdna/aie2_pm.c
+++ b/drivers/accel/amdxdna/aie2_pm.c
@@ -23,7 +23,7 @@ static int aie2_pm_set_clk_gating(struct amdxdna_dev_hdl *ndev, u32 val)
 	if (ret)
 		return ret;
 
-	ndev->clk_gating = val;
+	ndev->aie.clk_gating = val;
 	return 0;
 }
 
@@ -35,7 +35,7 @@ int aie2_pm_set_dpm(struct amdxdna_dev_hdl *ndev, u32 dpm_level)
 	if (ret)
 		return ret;
 
-	ret = ndev->priv->hw_ops->set_dpm(ndev, dpm_level);
+	ret = ndev->priv->hw_ops->set_dpm(&ndev->aie, dpm_level);
 	if (!ret)
 		ndev->dpm_level = dpm_level;
 	amdxdna_pm_suspend_put(ndev->aie.xdna);
@@ -49,11 +49,11 @@ int aie2_pm_init(struct amdxdna_dev_hdl *ndev)
 
 	if (ndev->dev_status != AIE2_DEV_UNINIT) {
 		/* Resume device */
-		ret = ndev->priv->hw_ops->set_dpm(ndev, ndev->dpm_level);
+		ret = ndev->priv->hw_ops->set_dpm(&ndev->aie, ndev->dpm_level);
 		if (ret)
 			return ret;
 
-		ret = aie2_pm_set_clk_gating(ndev, ndev->clk_gating);
+		ret = aie2_pm_set_clk_gating(ndev, ndev->aie.clk_gating);
 		if (ret)
 			return ret;
 
@@ -64,7 +64,7 @@ int aie2_pm_init(struct amdxdna_dev_hdl *ndev)
 		ndev->max_dpm_level++;
 	ndev->max_dpm_level--;
 
-	ret = ndev->priv->hw_ops->set_dpm(ndev, ndev->max_dpm_level);
+	ret = ndev->priv->hw_ops->set_dpm(&ndev->aie, ndev->max_dpm_level);
 	if (ret)
 		return ret;
 	ndev->dpm_level = ndev->max_dpm_level;
diff --git a/drivers/accel/amdxdna/aie4_message.c b/drivers/accel/amdxdna/aie4_message.c
index 25a510bd7419..d48ef855fee7 100644
--- a/drivers/accel/amdxdna/aie4_message.c
+++ b/drivers/accel/amdxdna/aie4_message.c
@@ -140,6 +140,80 @@ int aie4_query_cert_firmware_version(struct amdxdna_dev_hdl *ndev,
 	return 0;
 }
 
+int aie4_init_dpm_freq_table(struct amdxdna_dev_hdl *ndev)
+{
+	DECLARE_AIE_MSG(aie4_msg_get_dpm_freq_table, AIE4_MSG_OP_GET_DPM_FREQ_TABLE);
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	u32 aie_levels, npu_levels, i;
+	int ret;
+
+	for (i = 0; i < AIE4_MAX_DPM_LEVEL_COUNT && ndev->priv->dpm_clk_tbl &&
+	     ndev->priv->dpm_clk_tbl[i].hclk; i++)
+		ndev->dpm_clk_tbl[i] = ndev->priv->dpm_clk_tbl[i];
+	ndev->max_aieclk_level = i ? i - 1 : 0;
+	ndev->max_npuhclk_level = i ? i - 1 : 0;
+
+	ret = aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	if (ret) {
+		XDNA_WARN(xdna, "Get DPM freq table failed, ret %d status 0x%x",
+			  ret, resp.status);
+		return ret;
+	}
+
+	aie_levels = resp.aieclk_table.num_levels;
+	npu_levels = resp.npuhclk_table.num_levels;
+
+	if (!aie_levels || !npu_levels ||
+	    aie_levels > AIE4_MAX_DPM_LEVEL_COUNT ||
+	    npu_levels > AIE4_MAX_DPM_LEVEL_COUNT) {
+		XDNA_ERR(xdna, "invalid dpm levels, aieclk: %u, npuhclk: %u",
+			 aie_levels, npu_levels);
+		return -EINVAL;
+	}
+
+	memset(ndev->dpm_clk_tbl, 0, sizeof(ndev->dpm_clk_tbl));
+	for (i = 0; i < aie_levels; i++)
+		ndev->dpm_clk_tbl[i].npuclk = resp.aieclk_table.values[i];
+
+	for (i = 0; i < npu_levels; i++)
+		ndev->dpm_clk_tbl[i].hclk = resp.npuhclk_table.values[i];
+
+	ndev->max_aieclk_level = aie_levels - 1;
+	ndev->max_npuhclk_level = npu_levels - 1;
+
+	return 0;
+}
+
+int aie4_query_dpm_level(struct amdxdna_dev_hdl *ndev,
+			 u32 *aieclk_dpm_level, u32 *npuhclk_dpm_level)
+{
+	DECLARE_AIE_MSG(aie4_msg_get_dpm_level, AIE4_MSG_OP_GET_CURRENT_DPM_LEVEL);
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	int ret;
+
+	ret = aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	if (ret)
+		return ret;
+
+	/*
+	 * Validate against ndev->max_aieclk_level and ndev->max_npuhclk_level
+	 * to ensure reported levels index into populated entries in dpm_clk_tbl.
+	 */
+	if (resp.aieclk_dpm_level > ndev->max_aieclk_level ||
+	    resp.npuhclk_dpm_level > ndev->max_npuhclk_level) {
+		XDNA_ERR(xdna,
+			 "invalid dpm level, aie: %u/%u, npu: %u/%u",
+			 resp.aieclk_dpm_level, ndev->max_aieclk_level,
+			 resp.npuhclk_dpm_level, ndev->max_npuhclk_level);
+		return -EINVAL;
+	}
+
+	*aieclk_dpm_level = resp.aieclk_dpm_level;
+	*npuhclk_dpm_level = resp.npuhclk_dpm_level;
+
+	return 0;
+}
+
 int aie4_attach_work_buffer(struct amdxdna_dev_hdl *ndev)
 {
 	DECLARE_AIE_MSG(aie4_msg_attach_work_buffer, AIE4_MSG_OP_ATTACH_WORK_BUFFER);
diff --git a/drivers/accel/amdxdna/aie4_msg_priv.h b/drivers/accel/amdxdna/aie4_msg_priv.h
index 4c06792df1bd..fe78df9e23c8 100644
--- a/drivers/accel/amdxdna/aie4_msg_priv.h
+++ b/drivers/accel/amdxdna/aie4_msg_priv.h
@@ -24,6 +24,8 @@ enum aie4_msg_opcode {
 	AIE4_MSG_OP_AIE_TILE_INFO                    = 0x30006,
 	AIE4_MSG_OP_AIE_VERSION_INFO                 = 0x30007,
 	AIE4_MSG_OP_POWER_OVERRIDE                   = 0x3000B,
+	AIE4_MSG_OP_GET_DPM_FREQ_TABLE               = 0x30012,
+	AIE4_MSG_OP_GET_CURRENT_DPM_LEVEL            = 0x30013,
 
 	AIE4_MSG_OP_ATTACH_WORK_BUFFER               = 0x40001,
 };
@@ -196,6 +198,35 @@ struct aie4_msg_power_override_resp {
 	enum aie4_msg_status status;
 } __packed;
 
+#define AIE4_MAX_DPM_LEVEL_COUNT	10
+
+struct aie4_dpm_table {
+	__u32 num_levels;
+	__u32 values[AIE4_MAX_DPM_LEVEL_COUNT];
+} __packed;
+
+/* AIE4_MSG_OP_GET_DPM_FREQ_TABLE */
+struct aie4_msg_get_dpm_freq_table_req {
+	__u32 rsvd;
+} __packed;
+
+struct aie4_msg_get_dpm_freq_table_resp {
+	enum aie4_msg_status status;
+	struct aie4_dpm_table aieclk_table;
+	struct aie4_dpm_table npuhclk_table;
+} __packed;
+
+/* AIE4_MSG_OP_GET_CURRENT_DPM_LEVEL */
+struct aie4_msg_get_dpm_level_req {
+	__u32 rsvd;
+} __packed;
+
+struct aie4_msg_get_dpm_level_resp {
+	enum aie4_msg_status status;
+	__u32 aieclk_dpm_level;
+	__u32 npuhclk_dpm_level;
+} __packed;
+
 #define AIE4_WORK_BUFFER_MIN_SIZE      SZ_4M
 
 struct aie4_msg_attach_work_buffer_req {
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index d11bdbf16881..95e682a3a4b7 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -292,6 +292,15 @@ static int aie4_query(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		return ret;
 
+	ndev->total_col = min_t(u32, AIE4_TOTAL_COLUMN, ndev->aie.metadata.cols);
+
+	ret = aie4_init_dpm_freq_table(ndev);
+	if (ret) {
+		/* if query dpm from fw failed, using default value */
+		if (ndev->priv->hw_ops && ndev->priv->hw_ops->set_dpm)
+			(void)ndev->priv->hw_ops->set_dpm(&ndev->aie, 0);
+	}
+
 	return 0;
 }
 
@@ -638,11 +647,75 @@ static int aie4_get_power_mode(struct amdxdna_client *client,
 	return 0;
 }
 
+static int aie4_query_clock_metadata(struct amdxdna_client *client,
+				     struct amdxdna_drm_get_info *args)
+{
+	struct amdxdna_drm_query_clock_metadata *clock;
+	struct amdxdna_dev *xdna = client->xdna;
+	struct amdxdna_dev_hdl *ndev;
+	int ret = 0;
+	u32 buf_sz;
+
+	ndev = xdna->dev_handle;
+	clock = kzalloc_obj(*clock);
+	if (!clock)
+		return -ENOMEM;
+
+	aie_update_counters(ndev);
+	snprintf(clock->mp_npu_clock.name, sizeof(clock->mp_npu_clock.name),
+		 "MP-NPU Clock");
+	clock->mp_npu_clock.freq_mhz = ndev->aie.npuclk_freq;
+	snprintf(clock->h_clock.name, sizeof(clock->h_clock.name), "H Clock");
+	clock->h_clock.freq_mhz = ndev->aie.hclk_freq;
+
+	buf_sz = min_t(u32, args->buffer_size, sizeof(*clock));
+	if (copy_to_user(u64_to_user_ptr(args->buffer), clock, buf_sz))
+		ret = -EFAULT;
+
+	kfree(clock);
+	return ret;
+}
+
+static int aie4_query_resource_info(struct amdxdna_client *client,
+				    struct amdxdna_drm_get_info *args)
+{
+	struct amdxdna_drm_get_resource_info res_info = {};
+	struct amdxdna_dev_hdl *ndev;
+	struct amdxdna_dev *xdna;
+	u32 buf_sz;
+
+	xdna = client->xdna;
+	ndev = xdna->dev_handle;
+
+	aie_update_counters(ndev);
+	res_info.npu_clk_max = ndev->dpm_clk_tbl[ndev->max_npuhclk_level].hclk;
+	res_info.npu_tops_max = ndev->aie.max_tops;
+	res_info.npu_tops_curr = ndev->aie.curr_tops;
+	/*
+	 * res_info.npu_task_max/npu_task_curr are left zero-initialized;
+	 * hardware context accounting for AIE4 will populate them in a
+	 * future patch.
+	 */
+
+	buf_sz = min_t(u32, args->buffer_size, sizeof(res_info));
+	if (copy_to_user(u64_to_user_ptr(args->buffer), &res_info, buf_sz))
+		return -EFAULT;
+
+	return 0;
+}
+
 static int aie4_get_info(struct amdxdna_client *client, struct amdxdna_drm_get_info *args)
 {
 	struct amdxdna_dev *xdna = client->xdna;
 	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
-	int ret;
+	int ret, idx;
+
+	if (!drm_dev_enter(&xdna->ddev, &idx))
+		return -ENODEV;
+
+	ret = amdxdna_pm_resume_get_locked(xdna);
+	if (ret)
+		goto dev_exit;
 
 	switch (args->param) {
 	case DRM_AMDXDNA_QUERY_AIE_METADATA:
@@ -651,19 +724,28 @@ static int aie4_get_info(struct amdxdna_client *client, struct amdxdna_drm_get_i
 	case DRM_AMDXDNA_QUERY_AIE_VERSION:
 		ret = amdxdna_get_aie_version(client, args, &ndev->aie.version);
 		break;
+	case DRM_AMDXDNA_QUERY_CLOCK_METADATA:
+		ret = aie4_query_clock_metadata(client, args);
+		break;
 	case DRM_AMDXDNA_QUERY_FIRMWARE_VERSION:
 		ret = amdxdna_get_firmware_version(client, args, &xdna->fw_ver);
 		break;
 	case DRM_AMDXDNA_GET_POWER_MODE:
 		ret = aie4_get_power_mode(client, args);
 		break;
+	case DRM_AMDXDNA_QUERY_RESOURCE_INFO:
+		ret = aie4_query_resource_info(client, args);
+		break;
 	default:
 		XDNA_ERR(xdna, "Not supported request parameter %u", args->param);
 		ret = -EOPNOTSUPP;
 	}
 
+	amdxdna_pm_suspend_put(xdna);
 	XDNA_DBG(xdna, "Got param %d", args->param);
 
+dev_exit:
+	drm_dev_exit(idx);
 	return ret;
 }
 
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index fd2c50dc8080..6e9e7f874a44 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -11,6 +11,7 @@
 #include <linux/pci.h>
 
 #include "aie.h"
+#include "aie4_msg_priv.h"
 #include "amdxdna_mailbox.h"
 
 struct cert_comp {
@@ -40,6 +41,9 @@ struct amdxdna_dev_priv {
 
 	struct aie_bar_off_pair	psp_regs_off[PSP_MAX_REGS];
 	struct aie_bar_off_pair	smu_regs_off[SMU_MAX_REGS];
+
+	const struct dpm_clk_freq	*dpm_clk_tbl;
+	const struct aie_hw_ops		*hw_ops;
 };
 
 struct amdxdna_dev_hdl {
@@ -50,6 +54,11 @@ struct amdxdna_dev_hdl {
 
 	struct mailbox			*mbox;
 	u32				partition_id;
+	u32				total_col;
+	u32				max_aieclk_level;
+	u32				max_npuhclk_level;
+
+	struct dpm_clk_freq		dpm_clk_tbl[AIE4_MAX_DPM_LEVEL_COUNT];
 
 	struct xarray                   cert_comp_xa; /* device level indexed by msix id */
 	struct mutex                    cert_comp_lock; /* protects cert_comp operations*/
@@ -79,6 +88,9 @@ int aie4_query_npu_firmware_version(struct amdxdna_dev_hdl *ndev,
 				    struct amdxdna_drm_query_firmware_version *fw_version);
 int aie4_query_cert_firmware_version(struct amdxdna_dev_hdl *ndev,
 				     struct amdxdna_drm_query_firmware_version *cert_version);
+int aie4_init_dpm_freq_table(struct amdxdna_dev_hdl *ndev);
+int aie4_query_dpm_level(struct amdxdna_dev_hdl *ndev,
+			 u32 *aieclk_dpm_level, u32 *npuhclk_dpm_level);
 int aie4_msg_set_power_mode(struct amdxdna_dev_hdl *ndev, u8 power_mode);
 u32 aie4_msg_pasid(struct amdxdna_client *client);
 
diff --git a/drivers/accel/amdxdna/npu1_regs.c b/drivers/accel/amdxdna/npu1_regs.c
index ca779674017a..b4a0ede636f0 100644
--- a/drivers/accel/amdxdna/npu1_regs.c
+++ b/drivers/accel/amdxdna/npu1_regs.c
@@ -71,24 +71,25 @@ static const struct amdxdna_fw_feature_tbl npu1_fw_feature_table[] = {
 	{ 0 }
 };
 
-static int npu1_set_dpm(struct amdxdna_dev_hdl *ndev, u32 dpm_level)
+static int npu1_set_dpm(struct aie_device *aie, u32 dpm_level)
 {
+	struct amdxdna_dev_hdl *ndev = aie->xdna->dev_handle;
 	u32 npuclk, hclk;
 	int ret;
 
 	npuclk = ndev->priv->dpm_clk_tbl[dpm_level].npuclk;
 	hclk = ndev->priv->dpm_clk_tbl[dpm_level].hclk;
-	ret = aie_smu_set_clocks(ndev->aie.smu_hdl, &npuclk, &hclk);
+	ret = aie_smu_set_clocks(aie->smu_hdl, &npuclk, &hclk);
 	if (ret)
 		return ret;
 
-	ndev->npuclk_freq = npuclk;
-	ndev->hclk_freq = hclk;
-	ndev->max_tops = 2 * ndev->total_col;
-	ndev->curr_tops = ndev->max_tops * hclk / 1028;
+	aie->npuclk_freq = npuclk;
+	aie->hclk_freq = hclk;
+	aie->max_tops = 2 * ndev->total_col;
+	aie->curr_tops = aie->max_tops * hclk / 1028;
 
-	XDNA_DBG(ndev->aie.xdna, "MP-NPU clock %d, H clock %d\n",
-		 ndev->npuclk_freq, ndev->hclk_freq);
+	XDNA_DBG(aie->xdna, "MP-NPU clock %d, H clock %d\n",
+		 aie->npuclk_freq, aie->hclk_freq);
 	return 0;
 }
 
@@ -123,7 +124,7 @@ static const struct amdxdna_dev_priv npu1_dev_priv = {
 		DEFINE_BAR_OFFSET(SMU_RESP_REG, NPU1_SMU, MPNPU_PUB_SCRATCH6),
 		DEFINE_BAR_OFFSET(SMU_OUT_REG,  NPU1_SMU, MPNPU_PUB_SCRATCH7),
 	},
-	.hw_ops		= &(const struct aie2_hw_ops) {
+	.hw_ops		= &(const struct aie_hw_ops) {
 		.set_dpm = npu1_set_dpm,
 	},
 };
diff --git a/drivers/accel/amdxdna/npu3_regs.c b/drivers/accel/amdxdna/npu3_regs.c
index 21e24901976c..c531fcca62bb 100644
--- a/drivers/accel/amdxdna/npu3_regs.c
+++ b/drivers/accel/amdxdna/npu3_regs.c
@@ -37,6 +37,8 @@
 #define MP1_C2PMSG_61_ALT_1     0x3B109F4
 #define MP1_C2PMSG_60_ALT_1     0x3B109F0
 
+#define NPU3_DPM_TOPS(ndev, hclk) (4096 * (ndev)->total_col * (hclk) / 1000000)
+
 static const struct amdxdna_fw_feature_tbl npu3_fw_feature_table[] = {
 	{ .major = 6, .min_minor = 0 },
 	{ 0 }
@@ -48,9 +50,75 @@ static const struct amdxdna_fw_feature_tbl npu3_cert_feature_table[] = {
 	{ 0 }
 };
 
+static const struct dpm_clk_freq npu3_dpm_clk_table[] = {
+	{  400,  400 },
+	{  960,  576 },
+	{ 1108,  576 },
+	{ 1200,  847 },
+	{ 1200, 1200 },
+	{ 1200, 1200 },
+	{ 1200, 1200 },
+	{ 1200, 1200 },
+	{ 0 }
+};
+
+static int npu3_set_dpm(struct aie_device *aie, u32 dpm_level)
+{
+	struct amdxdna_dev_hdl *ndev = aie->xdna->dev_handle;
+	u32 aie_lvl, npu_lvl;
+
+	if (dpm_level > max(ndev->max_aieclk_level, ndev->max_npuhclk_level)) {
+		XDNA_ERR(aie->xdna, "Invalid dpm level %u (max aie %u, npu %u)",
+			 dpm_level, ndev->max_aieclk_level, ndev->max_npuhclk_level);
+		return -EINVAL;
+	}
+
+	aie_lvl = min(dpm_level, ndev->max_aieclk_level);
+	npu_lvl = min(dpm_level, ndev->max_npuhclk_level);
+
+	aie->npuclk_freq = ndev->dpm_clk_tbl[aie_lvl].npuclk;
+	aie->hclk_freq = ndev->dpm_clk_tbl[npu_lvl].hclk;
+	aie->max_tops = NPU3_DPM_TOPS(ndev, ndev->dpm_clk_tbl[ndev->max_npuhclk_level].hclk);
+	aie->curr_tops = NPU3_DPM_TOPS(ndev, aie->hclk_freq);
+
+	XDNA_DBG(aie->xdna, "MP-NPU clock %d, H clock %d\n",
+		 aie->npuclk_freq, aie->hclk_freq);
+
+	return 0;
+}
+
+static int npu3_update_counters(struct aie_device *aie)
+{
+	struct amdxdna_dev_hdl *ndev = aie->xdna->dev_handle;
+	u32 aieclk_level, npuhclk_level;
+	int ret;
+
+	ret = aie4_query_dpm_level(ndev, &aieclk_level, &npuhclk_level);
+	if (!ret) {
+		aie->npuclk_freq = ndev->dpm_clk_tbl[aieclk_level].npuclk;
+		aie->hclk_freq = ndev->dpm_clk_tbl[npuhclk_level].hclk;
+		aie->max_tops = NPU3_DPM_TOPS(ndev,
+					      ndev->dpm_clk_tbl[ndev->max_npuhclk_level].hclk);
+		if (!aie->hclk_freq)
+			XDNA_WARN(aie->xdna, "dpm freq table not populated, clk is 0");
+	} else {
+		XDNA_WARN(aie->xdna, "cannot get dpm level from fw, using default");
+	}
+
+	aie->curr_tops = NPU3_DPM_TOPS(ndev, aie->hclk_freq);
+
+	return 0;
+}
+
+static const struct aie_hw_ops npu3_hw_ops = {
+	.set_dpm = npu3_set_dpm,
+	.update_counters = npu3_update_counters,
+};
+
 static const struct amdxdna_dev_priv npu3_dev_priv = {
 	.npufw_path             = "npu.sbin",
 	.certfw_path            = "cert.sbin",
+	.dpm_clk_tbl		= npu3_dpm_clk_table,
 	.mbox_bar		= NPU3_MBOX_BAR,
 	.mbox_rbuf_bar		= NPU3_MBOX_BUFFER_BAR,
 	.mbox_info_off		= NPU3_MBOX_INFO_OFF,
@@ -72,14 +140,17 @@ static const struct amdxdna_dev_priv npu3_dev_priv = {
 		DEFINE_BAR_OFFSET(SMU_RESP_REG, NPU3_SMU, MP1_C2PMSG_60_ALT_1),
 		DEFINE_BAR_OFFSET(SMU_OUT_REG,  NPU3_SMU, MP1_C2PMSG_61_ALT_1),
 	},
+	.hw_ops			= &npu3_hw_ops,
 };
 
 static const struct amdxdna_dev_priv npu3_dev_vf_priv = {
 	/* vf device does not load firmware */
+	.dpm_clk_tbl		= npu3_dpm_clk_table,
 	.mbox_bar		= NPU3_MBOX_BAR,
 	.mbox_rbuf_bar		= NPU3_MBOX_BUFFER_BAR,
 	.mbox_info_off		= NPU3_MBOX_INFO_OFF,
 	/* vf device does not have smu and psp */
+	.hw_ops			= &npu3_hw_ops,
 };
 
 const struct amdxdna_dev_info dev_npu3_pf_info = {
diff --git a/drivers/accel/amdxdna/npu4_regs.c b/drivers/accel/amdxdna/npu4_regs.c
index 15a161384625..c26380050c79 100644
--- a/drivers/accel/amdxdna/npu4_regs.c
+++ b/drivers/accel/amdxdna/npu4_regs.c
@@ -104,42 +104,44 @@ const struct amdxdna_fw_feature_tbl npu4_fw_feature_table[] = {
 	{ 0 }
 };
 
-static int npu4_set_dpm(struct amdxdna_dev_hdl *ndev, u32 dpm_level)
+static int npu4_set_dpm(struct aie_device *aie, u32 dpm_level)
 {
+	struct amdxdna_dev_hdl *ndev = aie->xdna->dev_handle;
 	int ret;
 
-	ret = aie_smu_set_dpm(ndev->aie.smu_hdl, dpm_level);
+	ret = aie_smu_set_dpm(aie->smu_hdl, dpm_level);
 	if (ret)
 		return ret;
 
-	ndev->npuclk_freq = ndev->priv->dpm_clk_tbl[dpm_level].npuclk;
-	ndev->hclk_freq = ndev->priv->dpm_clk_tbl[dpm_level].hclk;
-	ndev->max_tops = NPU4_DPM_TOPS(ndev, ndev->priv->dpm_clk_tbl[ndev->max_dpm_level].hclk);
-	ndev->curr_tops = NPU4_DPM_TOPS(ndev, ndev->hclk_freq);
+	aie->npuclk_freq = ndev->priv->dpm_clk_tbl[dpm_level].npuclk;
+	aie->hclk_freq = ndev->priv->dpm_clk_tbl[dpm_level].hclk;
+	aie->max_tops = NPU4_DPM_TOPS(ndev, ndev->priv->dpm_clk_tbl[ndev->max_dpm_level].hclk);
+	aie->curr_tops = NPU4_DPM_TOPS(ndev, aie->hclk_freq);
 
-	XDNA_DBG(ndev->aie.xdna, "MP-NPU clock %d, H clock %d\n",
-		 ndev->npuclk_freq, ndev->hclk_freq);
+	XDNA_DBG(aie->xdna, "MP-NPU clock %d, H clock %d\n",
+		 aie->npuclk_freq, aie->hclk_freq);
 
 	return 0;
 }
 
-static int npu4_update_counters(struct amdxdna_dev_hdl *ndev)
+static int npu4_update_counters(struct aie_device *aie)
 {
+	struct amdxdna_dev_hdl *ndev = aie->xdna->dev_handle;
 	struct amd_pmf_npu_metrics npu_metrics;
 	int ret;
 
-	ret = AIE2_GET_PMF_NPU_METRICS(&npu_metrics);
+	ret = AIE_GET_PMF_NPU_METRICS(&npu_metrics);
 	if (ret)
 		return ret;
 
-	ndev->npuclk_freq = npu_metrics.mpnpuclk_freq;
-	ndev->hclk_freq = npu_metrics.npuclk_freq;
-	ndev->curr_tops = NPU4_DPM_TOPS(ndev, ndev->hclk_freq);
+	aie->npuclk_freq = npu_metrics.mpnpuclk_freq;
+	aie->hclk_freq = npu_metrics.npuclk_freq;
+	aie->curr_tops = NPU4_DPM_TOPS(ndev, aie->hclk_freq);
 
 	return 0;
 }
 
-const struct aie2_hw_ops npu4_hw_ops = {
+const struct aie_hw_ops npu4_hw_ops = {
 	.set_dpm = npu4_set_dpm,
 	.update_counters = npu4_update_counters,
 };
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 09/20] accel/amdxdna: Add context switch hysteresis with debugfs control
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (7 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 08/20] accel/amdxdna: Add clock, DPM frequency, and resource info queries " David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:52   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 10/20] accel/amdxdna: Refactor AIE4 hardware initialization sequence David Zhang
                   ` (10 subsequent siblings)
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello, Nishad Saraf

Add aie4_set_ctx_hysteresis() to configure the AIE4 context switch
hysteresis timeout via the SET_RUNTIME_CONFIG message, applied at hw
start (PF and classic paths) with a default of 1000 us.

Expose a debugfs node 'ctx_switch_hysteresis_us' to change the timeout
at runtime (0 disables hysteresis). The stored value is re-applied on
every hw start so it survives runtime suspend/resume.

Co-developed-by: Nishad Saraf <nishads@amd.com>
Signed-off-by: Nishad Saraf <nishads@amd.com>
Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_message.c    | 50 +++++++++++++++
 drivers/accel/amdxdna/aie4_msg_priv.h   | 36 +++++++++++
 drivers/accel/amdxdna/aie4_pci.c        | 85 ++++++++++++++++++++++++-
 drivers/accel/amdxdna/aie4_pci.h        | 12 ++++
 drivers/accel/amdxdna/amdxdna_debugfs.c |  3 +
 drivers/accel/amdxdna/amdxdna_pci_drv.h |  1 +
 6 files changed, 185 insertions(+), 2 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_message.c b/drivers/accel/amdxdna/aie4_message.c
index d48ef855fee7..1bddcb183db6 100644
--- a/drivers/accel/amdxdna/aie4_message.c
+++ b/drivers/accel/amdxdna/aie4_message.c
@@ -248,3 +248,53 @@ int aie4_msg_set_power_mode(struct amdxdna_dev_hdl *ndev, u8 power_mode)
 
 	return ret;
 }
+
+int aie4_set_runtime_cfg(struct amdxdna_dev_hdl *ndev, u32 type,
+			 const void *data, size_t size)
+{
+	DECLARE_AIE_MSG(aie4_msg_set_runtime_cfg, AIE4_MSG_OP_SET_RUNTIME_CONFIG);
+	u8 buf[sizeof(req.type) + AIE4_RUNTIME_CFG_MAX_DATA_SIZE] = { 0 };
+	int ret;
+
+	if (size > AIE4_RUNTIME_CFG_MAX_DATA_SIZE)
+		return -EINVAL;
+
+	/*
+	 * Firmware expects a 4-byte @type immediately followed by the
+	 * per-type payload (size validated against the struct npu_msg_-
+	 * runtime_config_* picked by @type). The shared request struct
+	 * carries an inline @data[4] slot, so stage only the 4-byte @type
+	 * header plus the variable payload contiguously and send exactly
+	 * that many bytes on the wire.
+	 */
+	req.type = type;
+	memcpy(buf, &req.type, sizeof(req.type));
+	memcpy(buf + sizeof(req.type), data, size);
+
+	msg.send_data = buf;
+	msg.send_size = sizeof(req.type) + size;
+
+	ret = aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	if (ret)
+		XDNA_ERR(ndev->aie.xdna, "Failed to set runtime cfg %u: %d", type, ret);
+	return ret;
+}
+
+int aie4_set_ctx_hysteresis(struct amdxdna_dev_hdl *ndev, u32 timeout_us)
+{
+	struct aie4_msg_runtime_config_ctx_switch_hysteresis cfg = {
+		.timeout_us = timeout_us,
+	};
+	int ret;
+
+	ret = aie4_set_runtime_cfg(ndev, AIE4_RUNTIME_CONFIG_CTX_SWITCH_HYSTERESIS,
+				   &cfg, sizeof(cfg));
+	if (ret)
+		XDNA_WARN(ndev->aie.xdna,
+			  "Failed to set ctx switch hysteresis to %u us (%d), using fw default",
+			  timeout_us, ret);
+	else
+		XDNA_DBG(ndev->aie.xdna, "Context switch hysteresis set to %u us", timeout_us);
+
+	return ret;
+}
diff --git a/drivers/accel/amdxdna/aie4_msg_priv.h b/drivers/accel/amdxdna/aie4_msg_priv.h
index fe78df9e23c8..77984683a7b6 100644
--- a/drivers/accel/amdxdna/aie4_msg_priv.h
+++ b/drivers/accel/amdxdna/aie4_msg_priv.h
@@ -12,6 +12,7 @@
 enum aie4_msg_opcode {
 	AIE4_MSG_OP_IDENTIFY                         = 0x10002,
 	AIE4_MSG_OP_SUSPEND                          = 0x10003,
+	AIE4_MSG_OP_SET_RUNTIME_CONFIG               = 0x10007,
 	AIE4_MSG_OP_QUERY_CERT_FIRMWARE_VERSION      = 0x1000F,
 
 	AIE4_MSG_OP_CREATE_VFS                       = 0x20001,
@@ -65,6 +66,41 @@ struct aie4_msg_suspend_resp {
 	enum aie4_msg_status status;
 } __packed;
 
+/*
+ * Type selector for AIE4_MSG_OP_SET_RUNTIME_CONFIG. Values match the firmware
+ * ABI enum npu_msg_runtime_config_type; only the configs the driver programs
+ * are enumerated here.
+ */
+enum aie4_msg_runtime_config_type {
+	AIE4_RUNTIME_CONFIG_CTX_SWITCH_HYSTERESIS	= 0xD,
+	AIE4_MAX_RUNTIME_CONFIG
+};
+
+struct aie4_msg_set_runtime_cfg_req {
+	__u32 type;
+	__u8 data[4];
+} __packed;
+
+struct aie4_msg_set_runtime_cfg_resp {
+	enum aie4_msg_status status;
+} __packed;
+
+/* Maximum trailing per-type payload (struct npu_msg_runtime_config_*) in the
+ * firmware ABI; today the largest is npu_msg_runtime_config_event_trace_status
+ * at 12 bytes. Rounded up to leave headroom for future configs.
+ */
+#define AIE4_RUNTIME_CFG_MAX_DATA_SIZE 16
+
+/*
+ * Context switch hysteresis configuration.
+ *
+ * @timeout_us: Hysteresis time in microseconds for keeping a context loaded
+ *              in the AIE after it becomes idle, or 0 to disable hysteresis.
+ */
+struct aie4_msg_runtime_config_ctx_switch_hysteresis {
+	__u32 timeout_us;
+} __packed;
+
 struct aie4_msg_create_vfs_req {
 	__u32 vf_cnt;
 } __packed;
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index 95e682a3a4b7..60348ec5bc53 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -5,8 +5,10 @@
 
 #include <drm/amdxdna_accel.h>
 #include <drm/drm_drv.h>
+#include <drm/drm_file.h>
 #include <drm/drm_managed.h>
 #include <drm/drm_print.h>
+#include <linux/debugfs.h>
 #include <linux/firmware.h>
 #include <linux/sizes.h>
 
@@ -325,6 +327,20 @@ int aie4_restore_power_mode(struct amdxdna_dev_hdl *ndev)
 	return aie4_msg_set_power_mode(ndev, ndev->pw_mode);
 }
 
+static int aie4_config_fw(struct amdxdna_dev_hdl *ndev)
+{
+	int ret;
+
+	ret = aie4_attach_work_buffer(ndev);
+	if (ret)
+		return ret;
+
+	/* Best-effort tuning knob; failure is warned inside and does not fail hw start */
+	aie4_set_ctx_hysteresis(ndev, ndev->ctx_switch_hysteresis_us);
+
+	return 0;
+}
+
 static int aie4_pf_hw_start(struct amdxdna_dev_hdl *ndev)
 {
 	int ret;
@@ -337,7 +353,7 @@ static int aie4_pf_hw_start(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		goto stop_fw;
 
-	ret = aie4_attach_work_buffer(ndev);
+	ret = aie4_config_fw(ndev);
 	if (ret)
 		goto mbox_fini;
 
@@ -419,7 +435,7 @@ static int aie4_classic_hw_start(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		goto mailbox_fini;
 
-	ret = aie4_attach_work_buffer(ndev);
+	ret = aie4_config_fw(ndev);
 	if (ret)
 		goto mailbox_fini;
 
@@ -574,6 +590,7 @@ static int aie4m_pcidev_init(struct amdxdna_dev *xdna)
 
 	ndev->priv = xdna->dev_info->dev_priv;
 	ndev->aie.xdna = xdna;
+	ndev->ctx_switch_hysteresis_us = AIE4_CTX_HYSTERESIS_US;
 	ndev->pw_mode = POWER_MODE_DEFAULT;
 	xdna->dev_handle = ndev;
 
@@ -921,15 +938,78 @@ static void aie4_classic_fini(struct amdxdna_dev *xdna)
 	aie4_free_work_buffer(xdna->dev_handle);
 }
 
+static int aie4_ctx_hysteresis_get(void *data, u64 *val)
+{
+	struct amdxdna_dev_hdl *ndev = data;
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+
+	guard(mutex)(&xdna->dev_lock);
+	*val = ndev->ctx_switch_hysteresis_us;
+
+	return 0;
+}
+
+static int aie4_ctx_hysteresis_set(void *data, u64 val)
+{
+	struct amdxdna_dev_hdl *ndev = data;
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	int ret, idx;
+
+	if (val > U32_MAX)
+		return -EINVAL;
+
+	if (!drm_dev_enter(&xdna->ddev, &idx))
+		return -ENODEV;
+
+	mutex_lock(&xdna->dev_lock);
+
+	ret = amdxdna_pm_resume_get_locked(xdna);
+	if (ret)
+		goto unlock;
+
+	ret = aie4_set_ctx_hysteresis(ndev, (u32)val);
+	if (!ret)
+		ndev->ctx_switch_hysteresis_us = (u32)val;
+
+	amdxdna_pm_suspend_put(xdna);
+
+unlock:
+	mutex_unlock(&xdna->dev_lock);
+	drm_dev_exit(idx);
+
+	return ret;
+}
+
+/* Context switch hysteresis timeout in microseconds; 0 disables hysteresis. */
+DEFINE_DEBUGFS_ATTRIBUTE(aie4_ctx_hysteresis_fops, aie4_ctx_hysteresis_get,
+			 aie4_ctx_hysteresis_set, "%llu\n");
+
+static void aie4_debugfs_init(struct amdxdna_dev *xdna)
+{
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+
+	/*
+	 * Context switch hysteresis is a system-control runtime config that is
+	 * only programmed on the PF/classic hw start paths, never on a VF.
+	 * Only expose the knob where the driver actually applies it.
+	 */
+	if (!to_pci_dev(xdna->ddev.dev)->is_virtfn)
+		debugfs_create_file_unsafe("ctx_switch_hysteresis_us", 0600,
+					   xdna->ddev.accel->debugfs_root, ndev,
+					   &aie4_ctx_hysteresis_fops);
+}
+
 const struct amdxdna_dev_ops aie4_pf_ops = {
 	.init			= aie4_pf_init,
 	.fini			= aie4_pf_fini,
+	.debugfs_init		= aie4_debugfs_init,
 	.sriov_configure        = aie4_sriov_configure,
 };
 
 const struct amdxdna_dev_ops aie4_vf_ops = {
 	.init			= aie4_vf_init,
 	.fini			= aie4_vf_fini,
+	.debugfs_init		= aie4_debugfs_init,
 	.hwctx_init		= aie4_hwctx_init,
 	.hwctx_fini		= aie4_hwctx_fini,
 	.cmd_wait		= aie4_cmd_wait,
@@ -940,6 +1020,7 @@ const struct amdxdna_dev_ops aie4_vf_ops = {
 const struct amdxdna_dev_ops aie4_classic_ops = {
 	.init			= aie4_classic_init,
 	.fini			= aie4_classic_fini,
+	.debugfs_init		= aie4_debugfs_init,
 	.hwctx_init		= aie4_hwctx_init,
 	.hwctx_fini		= aie4_hwctx_fini,
 	.cmd_wait		= aie4_cmd_wait,
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index 6e9e7f874a44..063cedfe3c9d 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -14,6 +14,9 @@
 #include "aie4_msg_priv.h"
 #include "amdxdna_mailbox.h"
 
+/* Default context switch hysteresis timeout in microseconds. */
+#define AIE4_CTX_HYSTERESIS_US	1000
+
 struct cert_comp {
 	struct amdxdna_dev_hdl          *ndev;
 	u32                             msix_idx;
@@ -69,6 +72,12 @@ struct amdxdna_dev_hdl {
 
 	u8				pw_mode;
 
+	/*
+	 * Context switch hysteresis timeout in microseconds; pushed to the
+	 * firmware at hw start and tunable at runtime via debugfs.
+	 */
+	u32				ctx_switch_hysteresis_us;
+
 	struct amdxdna_drm_query_firmware_version cert_version;
 };
 
@@ -92,6 +101,9 @@ int aie4_init_dpm_freq_table(struct amdxdna_dev_hdl *ndev);
 int aie4_query_dpm_level(struct amdxdna_dev_hdl *ndev,
 			 u32 *aieclk_dpm_level, u32 *npuhclk_dpm_level);
 int aie4_msg_set_power_mode(struct amdxdna_dev_hdl *ndev, u8 power_mode);
+int aie4_set_runtime_cfg(struct amdxdna_dev_hdl *ndev, u32 type,
+			 const void *data, size_t size);
+int aie4_set_ctx_hysteresis(struct amdxdna_dev_hdl *ndev, u32 timeout_us);
 u32 aie4_msg_pasid(struct amdxdna_client *client);
 
 /* aie4_ctx.c */
diff --git a/drivers/accel/amdxdna/amdxdna_debugfs.c b/drivers/accel/amdxdna/amdxdna_debugfs.c
index a6ec17c63629..1f63cc91b168 100644
--- a/drivers/accel/amdxdna/amdxdna_debugfs.c
+++ b/drivers/accel/amdxdna/amdxdna_debugfs.c
@@ -126,4 +126,7 @@ void amdxdna_debugfs_init(struct amdxdna_dev *xdna)
 				    xdna,
 				    amdxdna_dbgfs_files[i].fops);
 	}
+
+	if (xdna->dev_info->ops->debugfs_init)
+		xdna->dev_info->ops->debugfs_init(xdna);
 }
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.h b/drivers/accel/amdxdna/amdxdna_pci_drv.h
index 953bf783b3f7..11f46ec738d7 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.h
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.h
@@ -54,6 +54,7 @@ struct amdxdna_sched_job;
 struct amdxdna_dev_ops {
 	int (*init)(struct amdxdna_dev *xdna);
 	void (*fini)(struct amdxdna_dev *xdna);
+	void (*debugfs_init)(struct amdxdna_dev *xdna);
 	int (*resume)(struct amdxdna_dev *xdna);
 	int (*suspend)(struct amdxdna_dev *xdna);
 	int (*sriov_configure)(struct amdxdna_dev *xdna, int num_vfs);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 10/20] accel/amdxdna: Refactor AIE4 hardware initialization sequence
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (8 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 09/20] accel/amdxdna: Add context switch hysteresis with debugfs control David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:32 ` [PATCH V1 11/20] accel/amdxdna: Decouple AIE4 doorbell and MSI-X notification transport hooks David Zhang
                   ` (9 subsequent siblings)
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

Reorganize AIE4 hardware initialization into distinct phases:
- aie4_query_fw(): Query NPU and CERT firmware versions.
- aie4_config_fw(): Attach work buffer and configure context switch
  hysteresis.
- aie4_setup_aie(): Query AIE version, metadata, initialize DPM frequency
  table, and initialize partitions.

Update aie4_pf_hw_start(), aie4_vf_hw_start(), and aie4_classic_hw_start()
to use these phases and unify error unwinding labels. As part of this,
aie4_pf_hw_start() now also calls aie4_query_fw(), which it previously
did not do.

Additionally:
- Zero-initialize struct smu_config smu_conf in aie4_prepare_firmware().
- Clean up iomem pointer type in aie4_fw_is_alive() to void __iomem *.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_pci.c | 82 +++++++++++++++++++-------------
 1 file changed, 50 insertions(+), 32 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index 60348ec5bc53..b1742116bedd 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -49,7 +49,7 @@ static int aie4_fw_is_alive(struct amdxdna_dev *xdna)
 {
 	const struct amdxdna_dev_priv *npriv = xdna->dev_info->dev_priv;
 	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
-	u32 __iomem *src;
+	void __iomem *src;
 	u32 fw_is_valid;
 	int ret;
 
@@ -273,7 +273,13 @@ static void aie4_partition_fini(struct amdxdna_dev_hdl *ndev)
 		XDNA_ERR(xdna, "partition fini failed: %d", ret);
 }
 
-static int aie4_query(struct amdxdna_dev_hdl *ndev)
+/*
+ * Called by all three hw_start paths (PF, VF, classic) right after mailbox
+ * init. aie4_query_cert_firmware_version() runs a CERT protocol
+ * compatibility check, so firmware/driver compatibility is intentionally
+ * verified before any other firmware operation is attempted.
+ */
+static int aie4_query_fw(struct amdxdna_dev_hdl *ndev)
 {
 	struct amdxdna_dev *xdna = ndev->aie.xdna;
 	int ret;
@@ -286,23 +292,6 @@ static int aie4_query(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		return ret;
 
-	ret = aie4_query_aie_version(ndev, &ndev->aie.version);
-	if (ret)
-		return ret;
-
-	ret = aie4_query_aie_metadata(ndev, &ndev->aie.metadata);
-	if (ret)
-		return ret;
-
-	ndev->total_col = min_t(u32, AIE4_TOTAL_COLUMN, ndev->aie.metadata.cols);
-
-	ret = aie4_init_dpm_freq_table(ndev);
-	if (ret) {
-		/* if query dpm from fw failed, using default value */
-		if (ndev->priv->hw_ops && ndev->priv->hw_ops->set_dpm)
-			(void)ndev->priv->hw_ops->set_dpm(&ndev->aie, 0);
-	}
-
 	return 0;
 }
 
@@ -341,6 +330,30 @@ static int aie4_config_fw(struct amdxdna_dev_hdl *ndev)
 	return 0;
 }
 
+static int aie4_setup_aie(struct amdxdna_dev_hdl *ndev)
+{
+	int ret;
+
+	ret = aie4_query_aie_version(ndev, &ndev->aie.version);
+	if (ret)
+		return ret;
+
+	ret = aie4_query_aie_metadata(ndev, &ndev->aie.metadata);
+	if (ret)
+		return ret;
+
+	ndev->total_col = min_t(u32, AIE4_TOTAL_COLUMN, ndev->aie.metadata.cols);
+
+	ret = aie4_init_dpm_freq_table(ndev);
+	if (ret) {
+		/* if query dpm from fw failed, using default value */
+		if (ndev->priv->hw_ops && ndev->priv->hw_ops->set_dpm)
+			(void)ndev->priv->hw_ops->set_dpm(&ndev->aie, 0);
+	}
+
+	return aie4_partition_init(ndev);
+}
+
 static int aie4_pf_hw_start(struct amdxdna_dev_hdl *ndev)
 {
 	int ret;
@@ -353,6 +366,10 @@ static int aie4_pf_hw_start(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		goto stop_fw;
 
+	ret = aie4_query_fw(ndev);
+	if (ret)
+		goto mbox_fini;
+
 	ret = aie4_config_fw(ndev);
 	if (ret)
 		goto mbox_fini;
@@ -390,21 +407,21 @@ static int aie4_vf_hw_start(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		return ret;
 
-	ret = aie4_query(ndev);
+	ret = aie4_query_fw(ndev);
 	if (ret)
-		goto mailbox_fini;
+		goto mbox_fini;
 
-	ret = aie4_partition_init(ndev);
+	ret = aie4_setup_aie(ndev);
 	if (ret)
-		goto mailbox_fini;
+		goto mbox_fini;
 
 	ret = aie4_restore_power_mode(ndev);
 	if (ret)
-		goto mailbox_fini;
+		goto mbox_fini;
 
 	return 0;
 
-mailbox_fini:
+mbox_fini:
 	aie4_mailbox_fini(ndev);
 	return ret;
 }
@@ -431,17 +448,17 @@ static int aie4_classic_hw_start(struct amdxdna_dev_hdl *ndev)
 	if (ret)
 		goto stop_fw;
 
-	ret = aie4_query(ndev);
+	ret = aie4_query_fw(ndev);
 	if (ret)
-		goto mailbox_fini;
+		goto mbox_fini;
 
 	ret = aie4_config_fw(ndev);
 	if (ret)
-		goto mailbox_fini;
+		goto mbox_fini;
 
-	ret = aie4_partition_init(ndev);
+	ret = aie4_setup_aie(ndev);
 	if (ret)
-		goto mailbox_fini;
+		goto mbox_fini;
 
 	ret = aie4_restore_power_mode(ndev);
 	if (ret)
@@ -451,10 +468,11 @@ static int aie4_classic_hw_start(struct amdxdna_dev_hdl *ndev)
 
 partition_fini:
 	aie4_partition_fini(ndev);
-mailbox_fini:
+mbox_fini:
 	aie4_mailbox_fini(ndev);
 stop_fw:
 	aie4_fw_stop(ndev);
+
 	return ret;
 }
 
@@ -528,8 +546,8 @@ static int aie4_prepare_firmware(struct amdxdna_dev_hdl *ndev,
 				 void __iomem *tbl[PCI_NUM_RESOURCES])
 {
 	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	struct smu_config smu_conf = {};
 	struct psp_config psp_conf;
-	struct smu_config smu_conf;
 	int i;
 
 	psp_conf.fw_size = npufw->size;
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 11/20] accel/amdxdna: Decouple AIE4 doorbell and MSI-X notification transport hooks
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (9 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 10/20] accel/amdxdna: Refactor AIE4 hardware initialization sequence David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:32 ` [PATCH V1 12/20] accel/amdxdna: Implement AIE4 kernel queue lifecycle and memory layout David Zhang
                   ` (8 subsequent siblings)
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello, Wendy Liang

Separate PCI-specific doorbell and interrupt notification handling from
the transport-neutral context code:
- Move MSI-X ISR and registration out of aie4_ctx.c into transport hooks
  aie4_request_notification() and aie4_free_notification() in aie4_pci.c.
- Add transport hooks aie4_doorbell_setup() and aie4_doorbell_ring() to
  validate the doorbell offset against the mapped doorbell BAR and ring
  the hardware doorbell.
- Map the doorbell BAR (BAR 2) via pcim_iomap() in aie4m_pcidev_init()
  and record ndev->doorbell_base. The doorbells are used exclusively by
  kernel submit driver on VF and classic devices. PF devices only perform
  management functions, and never host hardware contexts, thus never use
  the doorbells.

Co-developed-by: Wendy Liang <wendy.liang@amd.com>
Signed-off-by: Wendy Liang <wendy.liang@amd.com>
Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_ctx.c    | 29 +++--------
 drivers/accel/amdxdna/aie4_pci.c    | 79 +++++++++++++++++++++++++++++
 drivers/accel/amdxdna/aie4_pci.h    | 18 +++++++
 drivers/accel/amdxdna/amdxdna_ctx.h |  2 +
 4 files changed, 107 insertions(+), 21 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
index 5eb918e1d58c..5a2fc19bad20 100644
--- a/drivers/accel/amdxdna/aie4_ctx.c
+++ b/drivers/accel/amdxdna/aie4_ctx.c
@@ -22,18 +22,9 @@
 #include "amdxdna_mailbox_helper.h"
 #include "amdxdna_pci_drv.h"
 
-static irqreturn_t cert_comp_isr(int irq, void *p)
-{
-	struct cert_comp *cert_comp = p;
-
-	wake_up_all(&cert_comp->waitq);
-	return IRQ_HANDLED;
-}
-
 static struct cert_comp *aie4_lookup_cert_comp(struct amdxdna_dev_hdl *ndev, u32 msix_idx)
 {
 	struct amdxdna_dev *xdna = ndev->aie.xdna;
-	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
 	struct cert_comp *cert_comp;
 	int ret;
 
@@ -51,32 +42,27 @@ static struct cert_comp *aie4_lookup_cert_comp(struct amdxdna_dev_hdl *ndev, u32
 
 	cert_comp->ndev = ndev;
 	cert_comp->msix_idx = msix_idx;
+	cert_comp->irq = -ENOENT;
 	init_waitqueue_head(&cert_comp->waitq);
 	kref_init(&cert_comp->kref);
 
-	ret = pci_irq_vector(pdev, cert_comp->msix_idx);
-	if (ret < 0) {
-		XDNA_ERR(xdna, "MSI-X idx %u is invalid, ret:%d", msix_idx, ret);
-		goto free_cert_comp;
-	}
-	cert_comp->irq = ret;
-
-	ret = request_irq(cert_comp->irq, cert_comp_isr, 0, "xdna_hsa", cert_comp);
+	/* Transport-specific: PCI wires an MSI-X irq, platform an IPI callback. */
+	ret = aie4_request_notification(cert_comp);
 	if (ret) {
-		XDNA_ERR(xdna, "request irq %d failed %d", cert_comp->irq, ret);
+		XDNA_ERR(xdna, "request notification for msix idx %u failed %d", msix_idx, ret);
 		goto free_cert_comp;
 	}
 
 	ret = xa_err(xa_store(&ndev->cert_comp_xa, msix_idx, cert_comp, GFP_KERNEL));
 	if (ret) {
-		XDNA_ERR(xdna, "store cert_comp for msix idx %d failed %d", msix_idx, ret);
+		XDNA_ERR(xdna, "store cert_comp for msix idx %u failed %d", msix_idx, ret);
 		goto free_irq;
 	}
 
 	return cert_comp;
 
 free_irq:
-	free_irq(cert_comp->irq, cert_comp);
+	aie4_free_notification(cert_comp);
 free_cert_comp:
 	kfree(cert_comp);
 	return NULL;
@@ -90,7 +76,7 @@ static void cert_comp_release(struct kref *kref)
 	drm_WARN_ON(&ndev->aie.xdna->ddev, !mutex_is_locked(&ndev->cert_comp_lock));
 
 	xa_erase(&ndev->cert_comp_xa, cert_comp->msix_idx);
-	free_irq(cert_comp->irq, cert_comp);
+	aie4_free_notification(cert_comp);
 	kfree(cert_comp);
 }
 
@@ -100,6 +86,7 @@ static void aie4_put_cert_comp(struct cert_comp *cert_comp)
 
 	ndev = cert_comp->ndev;
 	guard(mutex)(&ndev->cert_comp_lock);
+
 	kref_put(&cert_comp->kref, cert_comp_release);
 }
 
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index b1742116bedd..f5fdc24689f8 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -14,6 +14,7 @@
 
 #include "aie.h"
 #include "aie4_msg_priv.h"
+#include "amdxdna_ctx.h"
 #include "aie4_pci.h"
 #include "amdxdna_mailbox.h"
 #include "amdxdna_mailbox_helper.h"
@@ -110,6 +111,82 @@ static void aie4_mailbox_fini(struct amdxdna_dev_hdl *ndev)
 	ndev->mbox = NULL;
 }
 
+static irqreturn_t cert_comp_isr(int irq, void *p)
+{
+	struct cert_comp *cert_comp = p;
+
+	wake_up_all(&cert_comp->waitq);
+	return IRQ_HANDLED;
+}
+
+/*
+ * Transport hook: wire the per-cert completion notification.  PCI maps the
+ * firmware-provided MSI-X index to a Linux irq and registers cert_comp_isr;
+ * the platform build registers an IPI mailbox callback instead.
+ */
+int aie4_request_notification(struct cert_comp *comp)
+{
+	struct pci_dev *pdev = to_pci_dev(comp->ndev->aie.xdna->ddev.dev);
+	int ret;
+
+	ret = pci_irq_vector(pdev, comp->msix_idx);
+	if (ret < 0)
+		return ret;
+	comp->irq = ret;
+
+	ret = request_irq(comp->irq, cert_comp_isr, 0, "xdna_hsa", comp);
+	if (ret) {
+		comp->irq = -ENOENT;
+		return ret;
+	}
+
+	return 0;
+}
+
+/* Transport hook: tear down the completion notification wired by the hook above. */
+void aie4_free_notification(struct cert_comp *comp)
+{
+	if (comp->irq >= 0)
+		free_irq(comp->irq, comp);
+}
+
+/*
+ * Transport hook: take what this transport needs from the create-context
+ * response.  PCI validates the firmware-provided doorbell offset against the
+ * mapped doorbell BAR and stores this context's kick target.
+ */
+int aie4_doorbell_setup(struct amdxdna_hwctx *hwctx,
+			const struct aie4_msg_create_hw_context_resp *resp)
+{
+	struct amdxdna_dev *xdna = hwctx->client->xdna;
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
+	u64 db_off = (u64)ndev->priv->doorbell_off + resp->doorbell_offset;
+
+	/*
+	 * doorbell_base is a pcim_iomap() of the whole doorbell BAR.  The offset
+	 * comes from firmware (or, on a VF, the PF/hypervisor); reject one that
+	 * would place the u32 doorbell write past the mapped BAR before
+	 * aie4_doorbell_ring() ever dereferences priv->doorbell_addr.
+	 */
+	if (db_off + sizeof(u32) >
+	    pci_resource_len(pdev, xdna->dev_info->doorbell_bar)) {
+		XDNA_ERR(xdna, "doorbell offset 0x%llx out of BAR", db_off);
+		return -EINVAL;
+	}
+
+	priv->doorbell_addr = ndev->doorbell_base + ndev->priv->doorbell_off +
+			      resp->doorbell_offset;
+	return 0;
+}
+
+/* Transport hook: ring this context's doorbell (kick CERT). */
+void aie4_doorbell_ring(struct amdxdna_hwctx *hwctx)
+{
+	writel(0, hwctx->priv->doorbell_addr);
+}
+
 static int aie4_irq_init(struct amdxdna_dev *xdna)
 {
 	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
@@ -634,6 +711,7 @@ static int aie4m_pcidev_init(struct amdxdna_dev *xdna)
 		set_bit(SMU_REG_BAR(ndev, i), &bars);
 	set_bit(xdna->dev_info->mbox_bar, &bars);
 	set_bit(xdna->dev_info->sram_bar, &bars);
+	set_bit(xdna->dev_info->doorbell_bar, &bars);
 
 	for (i = 0; i < PCI_NUM_RESOURCES; i++) {
 		if (!test_bit(i, &bars))
@@ -647,6 +725,7 @@ static int aie4m_pcidev_init(struct amdxdna_dev *xdna)
 
 	ndev->mbox_base = tbl[xdna->dev_info->mbox_bar];
 	ndev->rbuf_base = tbl[xdna->dev_info->sram_bar];
+	ndev->doorbell_base = tbl[xdna->dev_info->doorbell_bar];
 
 	pci_set_master(pdev);
 
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index 063cedfe3c9d..c6e7f6a80f69 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -32,6 +32,8 @@ struct amdxdna_hwctx_priv {
 
 	struct cert_comp                *cert_comp;
 	u32                             hw_ctx_id;
+
+	void                    __iomem *doorbell_addr;
 };
 
 struct amdxdna_dev_priv {
@@ -54,6 +56,7 @@ struct amdxdna_dev_hdl {
 	const struct amdxdna_dev_priv	*priv;
 	void			__iomem *mbox_base;
 	void			__iomem *rbuf_base;
+	void			__iomem *doorbell_base;
 
 	struct mailbox			*mbox;
 	u32				partition_id;
@@ -114,6 +117,21 @@ int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout);
 /* aie4_pci.c */
 int aie4_restore_power_mode(struct amdxdna_dev_hdl *ndev);
 
+/*
+ * Transport hooks: one definition per build (aie4_pci.c for PCI; a future
+ * OF/platform transport provides its own), selected at compile time.  aie4_ctx.c is
+ * transport-neutral and reaches the doorbell kick and the completion interrupt
+ * only through these.  The cert_comp object itself (allocation/xarray/kref/
+ * waitq) is firmware-driven and stays neutral in aie4_ctx.c; only the notification
+ * wiring (PCI MSI-X vs platform IPI callback) is transport-specific.
+ */
+struct aie4_msg_create_hw_context_resp;
+int aie4_doorbell_setup(struct amdxdna_hwctx *hwctx,
+			const struct aie4_msg_create_hw_context_resp *resp);
+void aie4_doorbell_ring(struct amdxdna_hwctx *hwctx);
+int aie4_request_notification(struct cert_comp *comp);
+void aie4_free_notification(struct cert_comp *comp);
+
 /* aie4_sriov.c */
 #if IS_ENABLED(CONFIG_PCI_IOV)
 int aie4_sriov_configure(struct amdxdna_dev *xdna, int num_vfs);
diff --git a/drivers/accel/amdxdna/amdxdna_ctx.h b/drivers/accel/amdxdna/amdxdna_ctx.h
index 6e78bab8a02c..9bbc3db4ebde 100644
--- a/drivers/accel/amdxdna/amdxdna_ctx.h
+++ b/drivers/accel/amdxdna/amdxdna_ctx.h
@@ -6,6 +6,8 @@
 #ifndef _AMDXDNA_CTX_H_
 #define _AMDXDNA_CTX_H_
 
+#include <drm/amdxdna_accel.h>
+#include <drm/gpu_scheduler.h>
 #include <linux/bitfield.h>
 
 #include "amdxdna_gem.h"
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 12/20] accel/amdxdna: Implement AIE4 kernel queue lifecycle and memory layout
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (10 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 11/20] accel/amdxdna: Decouple AIE4 doorbell and MSI-X notification transport hooks David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  4:00   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 13/20] accel/amdxdna: Prepare for AIE4 command submission David Zhang
                   ` (7 subsequent siblings)
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello, Wendy Liang

Initialize kernel-mode submission required buffers, workqueue, and
hardware contexts:
- Update queue definition that is being used to send requests.
- Add job workqueue for pending and running jobs.
- Add mutex protection for each io.
- Initialize queue with direct and indirect packet, format queue header.
- Add kernel-mode submission required steps in hwctx create/destroy.
- Add aie4_get_cert_comp() to safely acquire a reference to the
  per-hwctx completion tracker under io_lock in aie4_cmd_wait() before
  waiting on the queue, preventing race conditions with
  aie4_hwctx_destroy().

Co-developed-by: Max Zhen <max.zhen@amd.com>
Signed-off-by: Max Zhen <max.zhen@amd.com>
Co-developed-by: Wendy Liang <wendy.liang@amd.com>
Signed-off-by: Wendy Liang <wendy.liang@amd.com>
Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_ctx.c        | 225 ++++++++++++++++++++----
 drivers/accel/amdxdna/aie4_host_queue.h |  65 +++++++
 drivers/accel/amdxdna/aie4_pci.h        |  54 ++++++
 3 files changed, 311 insertions(+), 33 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
index 5a2fc19bad20..3927c9fef05f 100644
--- a/drivers/accel/amdxdna/aie4_ctx.c
+++ b/drivers/accel/amdxdna/aie4_ctx.c
@@ -22,6 +22,13 @@
 #include "amdxdna_mailbox_helper.h"
 #include "amdxdna_pci_drv.h"
 
+#define CTX_INVALID_ID			(~0U)
+#define CTX_INVALID_DOORBELL		AMDXDNA_INVALID_DOORBELL_OFFSET
+
+static void job_worker(struct work_struct *work)
+{
+}
+
 static struct cert_comp *aie4_lookup_cert_comp(struct amdxdna_dev_hdl *ndev, u32 msix_idx)
 {
 	struct amdxdna_dev *xdna = ndev->aie.xdna;
@@ -38,7 +45,7 @@ static struct cert_comp *aie4_lookup_cert_comp(struct amdxdna_dev_hdl *ndev, u32
 
 	cert_comp = kzalloc_obj(*cert_comp);
 	if (!cert_comp)
-		return NULL;
+		return ERR_PTR(-ENOMEM);
 
 	cert_comp->ndev = ndev;
 	cert_comp->msix_idx = msix_idx;
@@ -65,7 +72,7 @@ static struct cert_comp *aie4_lookup_cert_comp(struct amdxdna_dev_hdl *ndev, u32
 	aie4_free_notification(cert_comp);
 free_cert_comp:
 	kfree(cert_comp);
-	return NULL;
+	return ERR_PTR(ret);
 }
 
 static void cert_comp_release(struct kref *kref)
@@ -73,8 +80,6 @@ static void cert_comp_release(struct kref *kref)
 	struct cert_comp *cert_comp = container_of(kref, struct cert_comp, kref);
 	struct amdxdna_dev_hdl *ndev = cert_comp->ndev;
 
-	drm_WARN_ON(&ndev->aie.xdna->ddev, !mutex_is_locked(&ndev->cert_comp_lock));
-
 	xa_erase(&ndev->cert_comp_xa, cert_comp->msix_idx);
 	aie4_free_notification(cert_comp);
 	kfree(cert_comp);
@@ -82,20 +87,42 @@ static void cert_comp_release(struct kref *kref)
 
 static void aie4_put_cert_comp(struct cert_comp *cert_comp)
 {
-	struct amdxdna_dev_hdl *ndev;
+	struct amdxdna_dev_hdl *ndev = cert_comp->ndev;
 
-	ndev = cert_comp->ndev;
 	guard(mutex)(&ndev->cert_comp_lock);
 
 	kref_put(&cert_comp->kref, cert_comp_release);
 }
 
-static int aie4_msg_destroy_context(struct amdxdna_dev_hdl *ndev, u32 hw_context_id)
+static struct cert_comp *aie4_get_cert_comp(struct amdxdna_hwctx *hwctx)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+	struct cert_comp *cert_comp;
+
+	/*
+	 * priv->cert_comp is the per-hwctx field, guarded by io_lock. A non-NULL
+	 * value means this ctx still holds its link-ref, so the object is alive and
+	 * the kref_get here cannot race the free.
+	 */
+	guard(mutex)(&priv->io_lock);
+
+	cert_comp = READ_ONCE(priv->cert_comp);
+	if (cert_comp)
+		kref_get(&cert_comp->kref);
+
+	return cert_comp;
+}
+
+static void aie4_msg_destroy_context(struct amdxdna_dev_hdl *ndev, u32 hw_context_id)
 {
 	DECLARE_AIE_MSG(aie4_msg_destroy_hw_context, AIE4_MSG_OP_DESTROY_HW_CONTEXT);
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	int ret;
 
 	req.hw_context_id = hw_context_id;
-	return aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	ret = aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	if (ret)
+		XDNA_WARN(xdna, "destroy ctx id %d failed %d", hw_context_id, ret);
 }
 
 static u8 aie4_parse_priority_to_dev(u32 priority)
@@ -114,19 +141,20 @@ static u8 aie4_parse_priority_to_dev(u32 priority)
 	}
 }
 
-static int aie4_hwctx_create(struct amdxdna_hwctx *hwctx)
+int aie4_hwctx_create(struct amdxdna_hwctx *hwctx)
 {
 	DECLARE_AIE_MSG(aie4_msg_create_hw_context, AIE4_MSG_OP_CREATE_HW_CONTEXT);
 	struct amdxdna_client *client = hwctx->client;
 	struct amdxdna_hwctx_priv *priv = hwctx->priv;
-	struct amdxdna_dev *xdna = hwctx->client->xdna;
+	struct amdxdna_dev *xdna = client->xdna;
 	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct cert_comp *cert_comp;
 	int ret;
 
 	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
 
 	if (!ndev->partition_id || !hwctx->num_tiles) {
-		XDNA_ERR(xdna, "invalid request partition_id %d, num_tiles %d",
+		XDNA_ERR(xdna, "invalid request partition_id %u, num_tiles %d",
 			 ndev->partition_id, hwctx->num_tiles);
 		return -EINVAL;
 	}
@@ -136,7 +164,6 @@ static int aie4_hwctx_create(struct amdxdna_hwctx *hwctx)
 	req.pasid = aie4_msg_pasid(client);
 	req.pasid = req.pasid == IOMMU_PASID_INVALID ? 0 : req.pasid;
 	req.priority_band = aie4_parse_priority_to_dev(hwctx->qos.priority);
-
 	req.hsa_addr_high = upper_32_bits(amdxdna_gem_dev_addr(priv->umq_bo));
 	req.hsa_addr_low = lower_32_bits(amdxdna_gem_dev_addr(priv->umq_bo));
 
@@ -150,72 +177,166 @@ static int aie4_hwctx_create(struct amdxdna_hwctx *hwctx)
 	}
 
 	XDNA_DBG(xdna, "resp msix: %d, ctx id: %d, doorbell: %d",
-		 resp.job_complete_msix_idx,
-		 resp.hw_context_id,
+		 resp.job_complete_msix_idx, resp.hw_context_id,
 		 resp.doorbell_offset);
 
 	/* setup interrupt completion per msix index */
-	priv->cert_comp = aie4_lookup_cert_comp(ndev, resp.job_complete_msix_idx);
-	if (!priv->cert_comp) {
+	cert_comp = aie4_lookup_cert_comp(ndev, resp.job_complete_msix_idx);
+	if (IS_ERR(cert_comp)) {
 		aie4_msg_destroy_context(ndev, resp.hw_context_id);
-		return -EINVAL;
+		return PTR_ERR(cert_comp);
 	}
 
 	priv->hw_ctx_id = resp.hw_context_id;
-	hwctx->doorbell_offset = AMDXDNA_INVALID_DOORBELL_OFFSET;
+
+	hwctx->fw_ctx_id = resp.hw_context_id;
+	hwctx->start_col = 0;
+	hwctx->num_col = ndev->total_col;
+
+	/*
+	 * Kernel-mode submission: set up this context's doorbell kick target
+	 * (transport-specific, via aie4_doorbell_setup) so the driver can ring
+	 * it, and keep it out of user space (hand back an invalid offset so the
+	 * doorbell cannot be mmap'd/rung by the user).
+	 */
+	mutex_lock(&priv->io_lock);
+	ret = aie4_doorbell_setup(hwctx, &resp);
+	if (ret) {
+		mutex_unlock(&priv->io_lock);
+		aie4_put_cert_comp(cert_comp);
+		aie4_msg_destroy_context(ndev, resp.hw_context_id);
+		priv->hw_ctx_id = CTX_INVALID_ID;
+		hwctx->fw_ctx_id = -1;
+		return ret;
+	}
+	WRITE_ONCE(priv->cert_comp, cert_comp);
+	mutex_unlock(&priv->io_lock);
+	hwctx->doorbell_offset = CTX_INVALID_DOORBELL;
+	wake_up_all(&priv->job_list_wq);
 
 	return 0;
 }
 
-static void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx)
+void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx, enum aie4_hwctx_flags flags)
 {
 	struct amdxdna_client *client = hwctx->client;
 	struct amdxdna_hwctx_priv *priv = hwctx->priv;
 	struct amdxdna_dev *xdna = client->xdna;
 	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct cert_comp *cert_comp;
 
 	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
 
-	aie4_msg_destroy_context(ndev, priv->hw_ctx_id);
-	aie4_put_cert_comp(priv->cert_comp);
+	mutex_lock(&priv->io_lock);
+	cert_comp = priv->cert_comp;
+	WRITE_ONCE(priv->cert_comp, NULL);
+	mutex_unlock(&priv->io_lock);
+
+	if (cert_comp) {
+		wake_up_all(&cert_comp->waitq);
+		aie4_put_cert_comp(cert_comp);
+	}
+
+	if (flags != AIE4_HWCTX_DISCONNECT)
+		aie4_msg_destroy_context(ndev, priv->hw_ctx_id);
+
+	priv->hw_ctx_id = CTX_INVALID_ID;
+	hwctx->fw_ctx_id = -1;
+	hwctx->doorbell_offset = CTX_INVALID_DOORBELL;
+
+	cancel_work_sync(&priv->job_work);
 }
 
 static void aie4_hwctx_umq_fini(struct amdxdna_hwctx *hwctx)
 {
 	if (hwctx->priv && hwctx->priv->umq_bo)
-		amdxdna_gem_put_obj(hwctx->priv->umq_bo);
+		drm_gem_object_put(to_gobj(hwctx->priv->umq_bo));
 }
 
 static int aie4_hwctx_umq_init(struct amdxdna_hwctx *hwctx)
 {
+	const size_t indir_pkts_sz = CTX_MAX_CMDS * HSA_MAX_LEVEL1_INDIRECT_ENTRIES *
+				     sizeof(struct host_indirect_packet_data);
+	const size_t pkts_sz = CTX_MAX_CMDS * sizeof(struct host_queue_packet);
 	struct amdxdna_hwctx_priv *priv = hwctx->priv;
 	struct amdxdna_dev *xdna = hwctx->client->xdna;
 	struct amdxdna_gem_obj *umq_bo;
 	struct host_queue_header *qhdr;
+	u64 data_dev_addr;
+	void *umq_va;
 	int ret;
+	int i;
 
+	/*
+	 * The HSA queue lives in a user-allocated BO (umq_bo_hdl) in both user- and
+	 * kernel-mode submission; the driver does not allocate it privately. Under
+	 * PASID/SVA the device reaches the queue through the submitting process's
+	 * own page tables, so it must have a user virtual address - a kernel-private
+	 * buffer would be unreachable by the device.
+	 */
 	umq_bo = amdxdna_gem_get_obj(hwctx->client, hwctx->umq_bo_hdl, AMDXDNA_BO_SHARE);
 	if (!umq_bo) {
 		XDNA_ERR(xdna, "cannot find umq_bo handle %d", hwctx->umq_bo_hdl);
 		return -ENOENT;
 	}
-	if (umq_bo->mem.size < sizeof(*qhdr)) {
-		XDNA_ERR(xdna, "umq_bo size is too small");
+
+	/*
+	 * Kernel-mode submission: the driver fills the host queue and rings the
+	 * doorbell, so the user umq_bo must hold the header plus the direct and
+	 * level-1 indirect packet arrays.
+	 */
+	if (umq_bo->mem.size < sizeof(*qhdr) ||
+	    (umq_bo->mem.size < sizeof(*qhdr) + pkts_sz + indir_pkts_sz)) {
+		XDNA_ERR(xdna, "umq_bo size %zu is too small",
+			 (size_t)umq_bo->mem.size);
 		ret = -EINVAL;
 		goto put_umq_bo;
 	}
 
-	/* get kva address for host queue read index and write index */
-	qhdr = amdxdna_gem_vmap(umq_bo);
-	if (!qhdr) {
+	umq_va = amdxdna_gem_vmap(umq_bo);
+	if (!umq_va) {
 		ret = -ENOMEM;
 		goto put_umq_bo;
 	}
+	qhdr = umq_va;
 
 	priv->umq_bo = umq_bo;
 	priv->umq_read_index = &qhdr->read_index;
 	priv->umq_write_index = &qhdr->write_index;
 
+	/*
+	 * The queue content is driver-owned and never trusted from user space
+	 * (only read_index is read back to detect completion). Lay out the
+	 * direct packets right after the header and the indirect packets after
+	 * them, and publish the same base via data_address for CERT.
+	 */
+	data_dev_addr = amdxdna_gem_dev_addr(umq_bo) + sizeof(*qhdr);
+	priv->umq_pkts = umq_va + sizeof(*qhdr);
+	priv->umq_indirect_pkts = umq_va + sizeof(*qhdr) + pkts_sz;
+	priv->umq_indirect_pkts_dev_addr = data_dev_addr + pkts_sz;
+
+	/*
+	 * Only the header + direct/indirect packet regions are driver-owned and
+	 * used for kernel submission; the size check above guarantees they fit.
+	 * Clear just that range, not the whole user-sized BO, so an oversized
+	 * umq_bo cannot force a huge memset (and page faults) under dev_lock.
+	 */
+	memset(umq_va, 0, sizeof(*qhdr) + pkts_sz + indir_pkts_sz);
+	priv->write_index = QUEUE_INDEX_START;
+	qhdr->read_index = QUEUE_INDEX_START;
+	qhdr->write_index = QUEUE_INDEX_START;
+	qhdr->version.major = HOST_QUEUE_MAJOR_VERSION;
+	qhdr->version.minor = HOST_QUEUE_MINOR_VERSION;
+	qhdr->capacity = CTX_MAX_CMDS;
+	qhdr->data_address = data_dev_addr;
+	for (i = 0; i < CTX_MAX_CMDS; i++)
+		priv->umq_pkts[i].pkt_header.common_header.opcode = OPCODE_EXEC_BUF;
+	for (i = 0; i < CTX_MAX_CMDS * HSA_MAX_LEVEL1_INDIRECT_ENTRIES; i++) {
+		priv->umq_indirect_pkts[i].header.opcode = OPCODE_EXEC_BUF;
+		priv->umq_indirect_pkts[i].header.count = sizeof(struct exec_buf);
+		priv->umq_indirect_pkts[i].header.distribute = 1;
+	}
+
 	return 0;
 
 put_umq_bo:
@@ -227,28 +348,56 @@ int aie4_hwctx_init(struct amdxdna_hwctx *hwctx)
 {
 	struct amdxdna_client *client = hwctx->client;
 	struct amdxdna_dev *xdna = client->xdna;
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
 	struct amdxdna_hwctx_priv *priv;
 	int ret;
 
+	if (!AIE_FEATURE_ON(&ndev->aie, AIE4_HSA_COMMAND))
+		return -EOPNOTSUPP;
+
 	priv = kzalloc_obj(*priv);
 	if (!priv)
 		return -ENOMEM;
 	hwctx->priv = priv;
+	priv->hwctx = hwctx;
+
+	/*
+	 * io_lock guards the per-hwctx cert_comp binding (the connected sentinel)
+	 * for every ctx, so initialize it unconditionally.  kzalloc left cert_comp
+	 * NULL: disconnected until create links it.
+	 */
+	mutex_init(&priv->io_lock);
+
+	INIT_LIST_HEAD(&priv->pending_job_list);
+	INIT_LIST_HEAD(&priv->running_job_list);
+	init_waitqueue_head(&priv->job_list_wq);
+	INIT_WORK(&priv->job_work, job_worker);
 
 	ret = aie4_hwctx_umq_init(hwctx);
 	if (ret)
-		goto free_priv;
+		goto destroy_lock;
 
 	ret = aie4_hwctx_create(hwctx);
 	if (ret)
 		goto umq_fini;
 
-	XDNA_DBG(xdna, "hwctx %s init completed", hwctx->name);
+	priv->job_work_q = alloc_ordered_workqueue("aie4_job_%d_%d", 0,
+						   client->pid, hwctx->fw_ctx_id);
+	if (!priv->job_work_q) {
+		XDNA_ERR(xdna, "Create job_work_q failed");
+		ret = -ENOMEM;
+		goto destroy_ctx;
+	}
+
+	XDNA_DBG(xdna, "hwctx %d.%d init completed", client->pid, hwctx->fw_ctx_id);
 	return 0;
 
+destroy_ctx:
+	aie4_hwctx_destroy(hwctx, AIE4_HWCTX_NORMAL);
 umq_fini:
 	aie4_hwctx_umq_fini(hwctx);
-free_priv:
+destroy_lock:
+	mutex_destroy(&priv->io_lock);
 	kfree(priv);
 	hwctx->priv = NULL;
 	return ret;
@@ -256,8 +405,14 @@ int aie4_hwctx_init(struct amdxdna_hwctx *hwctx)
 
 void aie4_hwctx_fini(struct amdxdna_hwctx *hwctx)
 {
-	aie4_hwctx_destroy(hwctx);
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+
+	aie4_hwctx_destroy(hwctx, AIE4_HWCTX_ERROR);
+	cancel_work_sync(&priv->job_work);
+	if (priv->job_work_q)
+		destroy_workqueue(priv->job_work_q);
 	aie4_hwctx_umq_fini(hwctx);
+	mutex_destroy(&priv->io_lock);
 	kfree(hwctx->priv);
 }
 
@@ -301,10 +456,12 @@ static inline bool check_cmd_done(struct amdxdna_hwctx *hwctx, u64 seq)
 int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout)
 {
 	unsigned long wait_jifs = MAX_SCHEDULE_TIMEOUT;
-	struct amdxdna_hwctx_priv *priv = hwctx->priv;
-	struct cert_comp *cert_comp = priv->cert_comp;
+	struct cert_comp *cert_comp = aie4_get_cert_comp(hwctx);
 	long ret;
 
+	if (!cert_comp)
+		return -EAGAIN;
+
 	if (timeout)
 		wait_jifs = msecs_to_jiffies(timeout);
 
@@ -315,5 +472,7 @@ int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout)
 	if (!ret)
 		ret = -ETIME;
 
+	aie4_put_cert_comp(cert_comp);
+
 	return ret <= 0 ? ret : 0;
 }
diff --git a/drivers/accel/amdxdna/aie4_host_queue.h b/drivers/accel/amdxdna/aie4_host_queue.h
index 97e535939b32..6876811b05f2 100644
--- a/drivers/accel/amdxdna/aie4_host_queue.h
+++ b/drivers/accel/amdxdna/aie4_host_queue.h
@@ -6,9 +6,14 @@
 #ifndef _AIE4_HOST_QUEUE_H_
 #define _AIE4_HOST_QUEUE_H_
 
+#include <linux/bits.h>
 #include <linux/types.h>
 
 #define CTX_MAX_CMDS                    32
+#define HSA_MAX_LEVEL1_INDIRECT_ENTRIES	6
+#define QUEUE_INDEX_START		0
+#define HOST_QUEUE_MAJOR_VERSION	1
+#define HOST_QUEUE_MINOR_VERSION	0
 
 /*
  * Host queue header layout.
@@ -31,4 +36,64 @@ struct host_queue_header {
 	__u64 data_address; /* The xdna dev addr for payload. */
 } __packed;
 
+/* Payload for an OPCODE_EXEC_BUF host-queue packet (single command). */
+struct exec_buf {
+	u32 dtrace_buf_host_addr_low;
+	u32 dpu_control_code_host_addr_low;
+	u32 dpu_control_code_host_addr_high;
+	u16 args_len;
+	u16 dtrace_buf_host_addr_high;
+	u32 args_host_addr_low;
+	u32 args_host_addr_high;
+} __packed;
+
+#define OPCODE_EXEC_BUF		1
+#define CHAIN_FLG_LAST_CMD	0
+#define CHAIN_FLG_NOT_LAST_CMD	1
+struct common_header {
+	u16 reserved; /* MBZ. */
+	u8 opcode;
+	u8 chain_flag;
+	u16 count;
+	u8 distribute;
+	u8 indirect;
+} __packed;
+
+struct host_queue_packet_header {
+	struct common_header common_header;
+	u64 completion_signal;
+} __packed;
+
+struct host_queue_packet {
+	struct host_queue_packet_header pkt_header;
+	u32 data[12]; /* total 64-byte packet */
+} __packed;
+
+struct host_indirect_packet_entry {
+	u32 host_addr_low;
+	u32 host_addr_high_uc_index;
+} __packed;
+
+#define HIPE_HOST_ADDR_HIGH_SHIFT	0
+#define HIPE_HOST_ADDR_HIGH_MASK	GENMASK(24, 0)
+#define HIPE_UC_INDEX_SHIFT		25
+#define HIPE_UC_INDEX_MASK		GENMASK(31, 25)
+
+static inline void hipe_set_host_addr_high(u32 *val, u32 addr_hi)
+{
+	*val &= ~HIPE_HOST_ADDR_HIGH_MASK;
+	*val |= (addr_hi << HIPE_HOST_ADDR_HIGH_SHIFT) & HIPE_HOST_ADDR_HIGH_MASK;
+}
+
+static inline void hipe_set_uc_index(u32 *val, u32 uc_idx)
+{
+	*val &= ~HIPE_UC_INDEX_MASK;
+	*val |= (uc_idx << HIPE_UC_INDEX_SHIFT) & HIPE_UC_INDEX_MASK;
+}
+
+struct host_indirect_packet_data {
+	struct common_header header;
+	struct exec_buf payload;
+} __packed;
+
 #endif /* _AIE4_HOST_QUEUE_H_ */
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index c6e7f6a80f69..f549d9e69d41 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -8,7 +8,10 @@
 
 #include <linux/device.h>
 #include <linux/iopoll.h>
+#include <linux/list.h>
 #include <linux/pci.h>
+#include <linux/wait.h>
+#include <linux/workqueue.h>
 
 #include "aie.h"
 #include "aie4_msg_priv.h"
@@ -25,15 +28,57 @@ struct cert_comp {
 	wait_queue_head_t               waitq;
 };
 
+/*
+ * aie4 kernel-submission job states (stored in amdxdna_sched_job priv.aie4.state).
+ * Anonymous enum - the aie4_job_state identifier is already a field-access macro.
+ */
+enum {
+	AIE4_JOB_STATE_INIT,
+	AIE4_JOB_STATE_PENDING,
+	AIE4_JOB_STATE_SUBMITTING,
+	AIE4_JOB_STATE_SUBMITTED,
+	AIE4_JOB_STATE_DONE,
+};
+
 struct amdxdna_hwctx_priv {
+	struct amdxdna_hwctx		*hwctx;
 	struct amdxdna_gem_obj          *umq_bo;
 	u64                             *umq_read_index;
 	u64                             *umq_write_index;
+	/* Last valid read_index, returned when a sampled index looks invalid. */
+	u64                             last_read_index;
 
 	struct cert_comp                *cert_comp;
 	u32                             hw_ctx_id;
 
+	/* Kernel-mode submission: driver fills the user HSA queue and rings
+	 * the doorbell.  umq_pkts/umq_indirect_pkts alias the user umq_bo;
+	 * their content is driver-owned, only read_index is trusted from the
+	 * shared queue.
+	 */
+	u64                             write_index;
+	struct host_queue_packet        *umq_pkts;
+	struct host_indirect_packet_data *umq_indirect_pkts;
+	u64                             umq_indirect_pkts_dev_addr;
+	/*
+	 * Transport-private doorbell kick target.  On PCI this is doorbell_base +
+	 * doorbell_off + firmware offset, set by aie4_doorbell_setup() and
+	 * dereferenced only by aie4_doorbell_ring() in aie4_pci.c.  Never touched
+	 * by aie4_ctx.c (unused on the platform build).  Gated by the cert_comp
+	 * connected sentinel, so it needs no INVALID poison.
+	 */
 	void                    __iomem *doorbell_addr;
+
+	struct mutex                    io_lock; /* serialize submit, protect job lists */
+	struct list_head                pending_job_list;
+	/* Head of pending_job_list, updated under io_lock; read locklessly by the
+	 * submit wait condition so it never takes a lock inside wait_event().
+	 */
+	struct amdxdna_sched_job        *pending_head;
+	struct list_head                running_job_list;
+	wait_queue_head_t               job_list_wq;
+	struct work_struct              job_work;
+	struct workqueue_struct         *job_work_q;
 };
 
 struct amdxdna_dev_priv {
@@ -110,9 +155,18 @@ int aie4_set_ctx_hysteresis(struct amdxdna_dev_hdl *ndev, u32 timeout_us);
 u32 aie4_msg_pasid(struct amdxdna_client *client);
 
 /* aie4_ctx.c */
+enum aie4_hwctx_flags {
+	AIE4_HWCTX_NORMAL = 0,
+	AIE4_HWCTX_GRACEFUL,
+	AIE4_HWCTX_DISCONNECT, /* sets has_reset, do not destroy context */
+	AIE4_HWCTX_ERROR, /* sets has_reset, destroy context */
+};
+
 int aie4_hwctx_init(struct amdxdna_hwctx *hwctx);
 void aie4_hwctx_fini(struct amdxdna_hwctx *hwctx);
 int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout);
+int aie4_hwctx_create(struct amdxdna_hwctx *hwctx);
+void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx, enum aie4_hwctx_flags);
 
 /* aie4_pci.c */
 int aie4_restore_power_mode(struct amdxdna_dev_hdl *ndev);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 13/20] accel/amdxdna: Prepare for AIE4 command submission
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (11 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 12/20] accel/amdxdna: Implement AIE4 kernel queue lifecycle and memory layout David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  4:00   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 14/20] accel/amdxdna: Implement AIE4 command packet building and submission David Zhang
                   ` (6 subsequent siblings)
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello, Wendy Liang

Prepare the data structures and completion wait helpers required for
AIE4 command submission:
- Define struct amdxdna_cmd_start_dpu in amdxdna_ctx.h for the
  ERT_START_DPU payload.
- Extend union amdxdna_job_priv with an aie4 member for queue list
  linkage and job state tracking.
- Implement smp_rmb() ordering and non-sleeping retry in get_read_index(),
  returning the cached last_read_index on read tearing to prevent false
  timeouts.
- Update check_cmd_done() and aie4_cmd_wait() to detect asynchronous
  device disconnect and reset via check_cert_comp_linked().

Co-developed-by: Max Zhen <max.zhen@amd.com>
Signed-off-by: Max Zhen <max.zhen@amd.com>
Co-developed-by: Wendy Liang <wendy.liang@amd.com>
Signed-off-by: Wendy Liang <wendy.liang@amd.com>
Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_ctx.c    | 89 ++++++++++++++++++++++++-----
 drivers/accel/amdxdna/amdxdna_ctx.h | 20 +++++++
 2 files changed, 95 insertions(+), 14 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
index 3927c9fef05f..59d37bd5c5a0 100644
--- a/drivers/accel/amdxdna/aie4_ctx.c
+++ b/drivers/accel/amdxdna/aie4_ctx.c
@@ -423,34 +423,92 @@ static inline bool valid_queue_index(u64 read, u64 write, u32 capacity)
 
 static u64 get_read_index(struct amdxdna_hwctx *hwctx)
 {
-	u64 wi = READ_ONCE(*hwctx->priv->umq_write_index);
-	u64 ri = READ_ONCE(*hwctx->priv->umq_read_index);
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
 	struct amdxdna_dev *xdna = hwctx->client->xdna;
+	u64 ri, wi;
+
+	/*
+	 * Sample read_index (written by CERT) before write_index. CERT can
+	 * never complete more than has been published, so a write_index sampled
+	 * after read_index always satisfies wi >= ri; sampling write_index
+	 * first races the submit path / CERT and yields a bogus ri > wi.
+	 *
+	 * In kernel-mode submission write_index is the driver's host-owned copy
+	 * in coherent kernel memory (always >= the value mirrored into the UMQ,
+	 * and the device never writes it).
+	 *
+	 * Security: read_index lives in the umq_bo, which the owning process can
+	 * map. Under PASID/SVA the device reaches the queue through that process's
+	 * own page tables, preventing access to other contexts' unshared memory.
+	 * A forged read_index only completes the process's own command early and
+	 * corrupts or hangs itself; however, if BOs are exported via dma-buf,
+	 * premature fence signaling can cause an importing process or device to
+	 * observe DMA completion before CERT has finished processing.
+	 */
+	ri = READ_ONCE(*priv->umq_read_index);
+	/* Order the read_index sample before the write_index sample. */
+	smp_rmb();
+	wi = READ_ONCE(priv->write_index);
 
 	/*
 	 * CERT cannot update read index as uint64 atomically. Driver may read
-	 * half-updated read index when it has bits in high 32bit. In case read
-	 * index is not valid, wait for some time and retry once. It should
-	 * allow CERT to complete the read index update.
+	 * a half-updated read index when it has bits in the high 32 bits. If it
+	 * looks invalid, re-sample once -- WITHOUT sleeping, since this can run as
+	 * a wait_event() condition. If still invalid, report not-advanced; the
+	 * waiter re-checks on the next completion wake or timeout.
 	 */
 	if (!valid_queue_index(ri, wi, CTX_MAX_CMDS)) {
-		XDNA_WARN(xdna, "Invalid index, ri %llu, wi %llu", ri, wi);
-		usleep_range(100, 200);
-		ri = READ_ONCE(*hwctx->priv->umq_read_index);
+		ri = READ_ONCE(*priv->umq_read_index);
+		/* Order the read_index sample before the write_index sample. */
+		smp_rmb();
+		wi = READ_ONCE(priv->write_index);
 		if (!valid_queue_index(ri, wi, CTX_MAX_CMDS)) {
-			XDNA_ERR(xdna, "Invalid index after retry, ri %llu, wi %llu", ri, wi);
-			ri = 0;
+			/*
+			 * Still invalid (torn 64-bit read, or a transient
+			 * accounting skew). Return the last valid read_index
+			 * instead of 0: read_index only advances, so the cached
+			 * value is a safe lower bound -- it never reports a
+			 * command complete that isn't, and never regresses the
+			 * worker into falsely timing out a finished job.
+			 */
+			XDNA_DBG(xdna, "Invalid index, ri %llu, wi %llu", ri, wi);
+			return READ_ONCE(priv->last_read_index);
 		}
 	}
 
+	WRITE_ONCE(priv->last_read_index, ri);
 	return ri;
 }
 
-static inline bool check_cmd_done(struct amdxdna_hwctx *hwctx, u64 seq)
+/*
+ * The ctx is "connected" as long as @comp is still the cert_comp linked to it.
+ * A disconnect (teardown/reset) unlinks (and may re-link a fresh) cert_comp, so
+ * a changed pointer means the caller must retry (-EAGAIN). This runs as a
+ * wait_event() condition on the completion hot path (cert_comp->waitq is shared
+ * per MSI-X), so keep it lockless: the caller pins @comp with a kref, making
+ * this a pure pointer-identity compare - never a dereference, ABA-safe - and
+ * READ_ONCE pairs with the WRITE_ONCE in aie4_hwctx_create()/
+ * aie4_hwctx_destroy().
+ */
+static bool check_cert_comp_linked(struct amdxdna_hwctx *hwctx, struct cert_comp *comp)
 {
-	u64 read_idx = get_read_index(hwctx);
+	/* READ_ONCE pairs with the link/unlink WRITE_ONCE. */
+	return comp == READ_ONCE(hwctx->priv->cert_comp);
+}
+
+static inline bool check_cmd_done(struct amdxdna_hwctx *hwctx, u64 seq, struct cert_comp *comp)
+{
+	/*
+	 * Runs as a wait_event() condition, so it must not sleep.
+	 * check_cert_comp_linked() is lockless (a READ_ONCE pointer compare); a
+	 * disconnect (teardown/reset) unlinks @comp and breaks the wait, and the
+	 * caller then confirms real completion by re-reading read_index, so a
+	 * disconnect wake is not mistaken for success.
+	 */
+	if (!check_cert_comp_linked(hwctx, comp))
+		return true;
 
-	return read_idx > seq;
+	return get_read_index(hwctx) > seq;
 }
 
 int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout)
@@ -466,11 +524,14 @@ int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout)
 		wait_jifs = msecs_to_jiffies(timeout);
 
 	ret = wait_event_interruptible_timeout(cert_comp->waitq,
-					       (check_cmd_done(hwctx, seq)),
+					       check_cmd_done(hwctx, seq, cert_comp),
 					       wait_jifs);
 
 	if (!ret)
 		ret = -ETIME;
+	else if (ret > 0 && get_read_index(hwctx) <= seq)
+		/* Woke on disconnect/reset, not on real completion. */
+		ret = -EAGAIN;
 
 	aie4_put_cert_comp(cert_comp);
 
diff --git a/drivers/accel/amdxdna/amdxdna_ctx.h b/drivers/accel/amdxdna/amdxdna_ctx.h
index 9bbc3db4ebde..b3677851d1c5 100644
--- a/drivers/accel/amdxdna/amdxdna_ctx.h
+++ b/drivers/accel/amdxdna/amdxdna_ctx.h
@@ -48,6 +48,18 @@ struct amdxdna_cmd_start_npu {
 	u32 prop_args[];  /* properties and regular kernel arguments */
 };
 
+/*
+ * struct amdxdna_cmd_start_dpu - interpretation of data payload for
+ * ERT_START_DPU in amdxdna_cmd.
+ */
+struct amdxdna_cmd_start_dpu {
+	u64 dtrace_buffer;		/* dtrace buffer address 2 words */
+	u64 instruction_buffer;		/* buffer address 2 words */
+	u32 instruction_buffer_size;	/* size of buffer in bytes */
+	u16 uc_index;			/* microblaze controller index */
+	u16 chained;			/* number of following amdxdna_cmd_start_dpu elements */
+};
+
 /*
  * Interpretation of the beginning of data payload for ERT_CMD_CHAIN in
  * amdxdna_cmd. The rest of the payload in amdxdna_cmd is cmd BO handles.
@@ -138,8 +150,14 @@ struct amdxdna_drv_cmd {
 };
 
 struct app_health_report;
+
 union amdxdna_job_priv {
 	struct app_health_report *aie2_health;
+	/* aie4 kernel submission: queue linkage + job state */
+	struct {
+		struct list_head	list;
+		u32			state;
+	} aie4;
 };
 
 struct amdxdna_sched_job {
@@ -162,6 +180,8 @@ struct amdxdna_sched_job {
 };
 
 #define aie2_job_health priv.aie2_health
+#define aie4_job_list	priv.aie4.list
+#define aie4_job_state	priv.aie4.state
 
 static inline u32
 amdxdna_cmd_get_op(struct amdxdna_gem_obj *abo)
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 14/20] accel/amdxdna: Implement AIE4 command packet building and submission
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (12 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 13/20] accel/amdxdna: Prepare for AIE4 command submission David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  4:04   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 15/20] accel/amdxdna: Finalize runtime PM before acquiring dev_lock on removal David Zhang
                   ` (5 subsequent siblings)
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello, Wendy Liang

Implement kernel-mode command submission and hardware queue packet
assembly for AIE4:
- Add aie4_cmd_submit() to validate incoming command buffers, reserve
  GEM fences, serialize submissions via the context pending list, and
  dispatch to the hardware queue.
- Lock all job BOs and attach job->fence directly to their reservation
  objects.
- In amdxdna_fence_create(), allocate a unique timeline context per job.
- In amdxdna_fence_get_timeline_name(), use dev_name() backed by the
  struct device rather than hwctx->name.
- Implement packet encoders: fill_direct_pkt() and fill_indirect_pkt().
- Advance the hardware write index, ring the doorbell, and register
  in-flight jobs on the running list for completion tracking.
- Add aie4_hwctx_wait_for_running() with timeout to safely quiesce worker
  threads.

Co-developed-by: Max Zhen <max.zhen@amd.com>
Signed-off-by: Max Zhen <max.zhen@amd.com>
Co-developed-by: Wendy Liang <wendy.liang@amd.com>
Signed-off-by: Wendy Liang <wendy.liang@amd.com>
Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_ctx.c    | 718 +++++++++++++++++++++++++++-
 drivers/accel/amdxdna/aie4_pci.c    |   2 +
 drivers/accel/amdxdna/aie4_pci.h    |   3 +
 drivers/accel/amdxdna/amdxdna_ctx.c |  20 +-
 4 files changed, 733 insertions(+), 10 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
index 59d37bd5c5a0..e9ba1ed93997 100644
--- a/drivers/accel/amdxdna/aie4_ctx.c
+++ b/drivers/accel/amdxdna/aie4_ctx.c
@@ -25,9 +25,7 @@
 #define CTX_INVALID_ID			(~0U)
 #define CTX_INVALID_DOORBELL		AMDXDNA_INVALID_DOORBELL_OFFSET
 
-static void job_worker(struct work_struct *work)
-{
-}
+static void job_worker(struct work_struct *work);
 
 static struct cert_comp *aie4_lookup_cert_comp(struct amdxdna_dev_hdl *ndev, u32 msix_idx)
 {
@@ -209,6 +207,7 @@ int aie4_hwctx_create(struct amdxdna_hwctx *hwctx)
 		hwctx->fw_ctx_id = -1;
 		return ret;
 	}
+	WRITE_ONCE(priv->has_reset, false);
 	WRITE_ONCE(priv->cert_comp, cert_comp);
 	mutex_unlock(&priv->io_lock);
 	hwctx->doorbell_offset = CTX_INVALID_DOORBELL;
@@ -217,6 +216,23 @@ int aie4_hwctx_create(struct amdxdna_hwctx *hwctx)
 	return 0;
 }
 
+/*
+ * The connected sentinel for the submit wait gate: a linked cert_comp means the
+ * ctx is created and (for kernel submission) its doorbell is set up. Read
+ * locklessly (a pointer null-check, never a dereference); create publishes it as
+ * the last store under io_lock and destroy clears it first, so "connected"
+ * implies a valid doorbell.
+ */
+static bool aie4_hwctx_connected(struct amdxdna_hwctx *hwctx)
+{
+	return !!READ_ONCE(hwctx->priv->cert_comp);
+}
+
+static bool aie4_hwctx_has_reset(struct amdxdna_hwctx *hwctx)
+{
+	return READ_ONCE(hwctx->priv->has_reset);
+}
+
 void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx, enum aie4_hwctx_flags flags)
 {
 	struct amdxdna_client *client = hwctx->client;
@@ -224,10 +240,16 @@ void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx, enum aie4_hwctx_flags flags
 	struct amdxdna_dev *xdna = client->xdna;
 	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
 	struct cert_comp *cert_comp;
+	bool has_reset = false;
 
 	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
 
+	if (flags == AIE4_HWCTX_DISCONNECT || flags == AIE4_HWCTX_ERROR)
+		has_reset = true;
+
 	mutex_lock(&priv->io_lock);
+	if (has_reset)
+		WRITE_ONCE(priv->has_reset, true);
 	cert_comp = priv->cert_comp;
 	WRITE_ONCE(priv->cert_comp, NULL);
 	mutex_unlock(&priv->io_lock);
@@ -237,6 +259,9 @@ void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx, enum aie4_hwctx_flags flags
 		aie4_put_cert_comp(cert_comp);
 	}
 
+	if (has_reset)
+		wake_up_all(&priv->job_list_wq);
+
 	if (flags != AIE4_HWCTX_DISCONNECT)
 		aie4_msg_destroy_context(ndev, priv->hw_ctx_id);
 
@@ -244,7 +269,15 @@ void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx, enum aie4_hwctx_flags flags
 	hwctx->fw_ctx_id = -1;
 	hwctx->doorbell_offset = CTX_INVALID_DOORBELL;
 
-	cancel_work_sync(&priv->job_work);
+	/*
+	 * When has_reset is true (AIE4_HWCTX_DISCONNECT or ERROR), cancel_work_sync()
+	 * is skipped so the worker can wake up and abort in-flight jobs. Callers
+	 * that re-create the context after DISCONNECT (e.g. reset recovery) must
+	 * synchronize the worker (via aie4_hwctx_wait_for_running()) before calling
+	 * aie4_hwctx_create().
+	 */
+	if (!has_reset)
+		cancel_work_sync(&priv->job_work);
 }
 
 static void aie4_hwctx_umq_fini(struct amdxdna_hwctx *hwctx)
@@ -407,10 +440,18 @@ void aie4_hwctx_fini(struct amdxdna_hwctx *hwctx)
 {
 	struct amdxdna_hwctx_priv *priv = hwctx->priv;
 
+	/*
+	 * ERROR sets has_reset (like a TDR reset) so the worker drains the running
+	 * list - the ctx is gone, so in-flight jobs are reaped - and stays live
+	 * for aie4_hwctx_wait_for_running() to wait on. Submitters are already gone
+	 * (amdxdna_hwctx_destroy_rcu() synchronize_srcu'd them out), then the
+	 * queue is torn down.
+	 */
 	aie4_hwctx_destroy(hwctx, AIE4_HWCTX_ERROR);
-	cancel_work_sync(&priv->job_work);
-	if (priv->job_work_q)
+	if (priv->job_work_q) {
+		aie4_hwctx_wait_for_running(hwctx);
 		destroy_workqueue(priv->job_work_q);
+	}
 	aie4_hwctx_umq_fini(hwctx);
 	mutex_destroy(&priv->io_lock);
 	kfree(hwctx->priv);
@@ -537,3 +578,668 @@ int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout)
 
 	return ret <= 0 ? ret : 0;
 }
+
+/* ---- kernel-mode submission (driver fills the queue and rings doorbell) ---- */
+
+/* Publish a command to CERT and return the assigned command sequence (slot). */
+static u64 publish_cmd(struct amdxdna_hwctx *hwctx)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+	u64 wi = priv->write_index;
+
+	/* Paired with the lockless READ_ONCE() readers of write_index. */
+	WRITE_ONCE(priv->write_index, wi + 1);
+	/* Order the packet-slot writes before CERT sees the new write_index. */
+	wmb();
+	WRITE_ONCE(*priv->umq_write_index, wi + 1);
+	return wi;
+}
+
+static int wait_till_seq_completed(struct amdxdna_hwctx *hwctx, u64 seq)
+{
+	struct cert_comp *cert_comp;
+	int ret;
+
+	/*
+	 * Freezable + interruptible: the submit path (wait_till_connected_hsa_not_full)
+	 * reaches here while holding hwctx_srcu, and ctx teardown blocks on
+	 * synchronize_srcu(), so a signal (e.g. the app being killed) must be
+	 * able to unwind the wait - otherwise a full queue with a silent CERT
+	 * would hang the submitter in D state and stall teardown forever.
+	 * TASK_FREEZABLE lets the freezer suspend this wait in place during
+	 * S3/S4 instead of aborting the suspend. Harmless for the job worker
+	 * kthread (never gets a signal; simply freezes/thaws around it).
+	 */
+	cert_comp = aie4_get_cert_comp(hwctx);
+	if (!cert_comp)
+		return -EAGAIN;
+
+	ret = wait_event_freezable(cert_comp->waitq,
+				   check_cmd_done(hwctx, seq, cert_comp));
+	if (ret) {
+		aie4_put_cert_comp(cert_comp);
+		return ret;	/* -ERESTARTSYS: signal on the submit path */
+	}
+
+	if (check_cert_comp_linked(hwctx, cert_comp))
+		ret = 0;			/* real completion */
+	else
+		ret = -EAGAIN;			/* disconnect (suspend or TDR) */
+
+	aie4_put_cert_comp(cert_comp);
+	return ret;
+}
+
+static int wait_till_connected_hsa_not_full(struct amdxdna_hwctx *hwctx,
+					    bool wait_through_reset)
+{
+	struct amdxdna_dev *xdna = hwctx->client->xdna;
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+	u64 wi = READ_ONCE(priv->write_index);
+	bool hsa_not_full = !!(wi < CTX_MAX_CMDS);
+	int ret;
+
+	do {
+		mutex_unlock(&priv->io_lock);
+		if (!hsa_not_full) {
+			ret = wait_till_seq_completed(hwctx, wi - CTX_MAX_CMDS);
+			if (ret && ret != -EAGAIN) {
+				mutex_lock(&priv->io_lock);
+				return ret;
+			}
+			if (!ret)
+				hsa_not_full = true;
+		}
+		ret = wait_event_freezable(priv->job_list_wq,
+					   aie4_hwctx_connected(hwctx) ||
+					   (!wait_through_reset &&
+					    aie4_hwctx_has_reset(hwctx)));
+		mutex_lock(&priv->io_lock);
+		if (ret)
+			return ret;
+		if (!wait_through_reset && aie4_hwctx_has_reset(hwctx)) {
+			XDNA_DBG(xdna, "ctx %s reset while submitting; unwinding -ECONNRESET",
+				 hwctx->name);
+			return -ECONNRESET;
+		}
+	} while (!hsa_not_full || !aie4_hwctx_connected(hwctx));
+
+	return 0;
+}
+
+static int fill_indirect_pkt(struct amdxdna_hwctx_priv *priv, u64 slot_idx,
+			     u32 total_slots, struct amdxdna_cmd_start_dpu *dpu,
+			     u16 entries)
+{
+	struct host_queue_packet *pkt = &priv->umq_pkts[slot_idx];
+	struct host_indirect_packet_entry *hipe =
+		(struct host_indirect_packet_entry *)(pkt->data);
+	u16 i;
+
+	for (i = 0; i < entries; i++, dpu++, hipe++) {
+		struct host_indirect_packet_data *hipd;
+		u64 indirect_pkt_dev_addr;
+		u32 uci = dpu->uc_index;
+		u32 idx;
+
+		/*
+		 * dpu is the user-shared cmd_abo payload, so uc_index is read at
+		 * use time here and indexes priv->umq_indirect_pkts[]. Reject an
+		 * out-of-range value: the slot is reused, so skipping the entry
+		 * would leave a stale one that count still advertises to CERT.
+		 * Abort before the packet is published.
+		 */
+		if (uci >= HSA_MAX_LEVEL1_INDIRECT_ENTRIES) {
+			XDNA_ERR(priv->hwctx->client->xdna, "Invalid uc index %d", uci);
+			return -EINVAL;
+		}
+		idx = uci * total_slots + slot_idx;
+		hipd = &priv->umq_indirect_pkts[idx];
+		indirect_pkt_dev_addr = priv->umq_indirect_pkts_dev_addr +
+			sizeof(struct host_indirect_packet_data) * idx;
+
+		/* Point the indirect entry at the indirect packet. */
+		hipe->host_addr_low = lower_32_bits(indirect_pkt_dev_addr);
+		hipe_set_host_addr_high(&hipe->host_addr_high_uc_index,
+					upper_32_bits(indirect_pkt_dev_addr));
+		hipe_set_uc_index(&hipe->host_addr_high_uc_index, uci);
+
+		/* Fill in the indirect packet. */
+		hipd->payload.dpu_control_code_host_addr_low =
+			lower_32_bits(dpu->instruction_buffer);
+		hipd->payload.dpu_control_code_host_addr_high =
+			upper_32_bits(dpu->instruction_buffer);
+		hipd->payload.dtrace_buf_host_addr_low =
+			lower_32_bits(dpu->dtrace_buffer);
+		hipd->payload.dtrace_buf_host_addr_high =
+			lower_16_bits(upper_32_bits(dpu->dtrace_buffer));
+	}
+	pkt->pkt_header.common_header.distribute = 1;
+	pkt->pkt_header.common_header.indirect = 1;
+	pkt->pkt_header.common_header.count = entries * sizeof(*hipe);
+	return 0;
+}
+
+static void fill_direct_pkt(struct amdxdna_hwctx_priv *priv, u64 slot_idx,
+			    struct amdxdna_cmd_start_dpu *dpu)
+{
+	struct host_queue_packet *pkt = &priv->umq_pkts[slot_idx];
+	struct exec_buf *ebuf = (struct exec_buf *)(pkt->data);
+
+	memset(pkt->data, 0, sizeof(pkt->data));
+	ebuf->dpu_control_code_host_addr_low = lower_32_bits(dpu->instruction_buffer);
+	ebuf->dpu_control_code_host_addr_high = upper_32_bits(dpu->instruction_buffer);
+	ebuf->dtrace_buf_host_addr_low = lower_32_bits(dpu->dtrace_buffer);
+	ebuf->dtrace_buf_host_addr_high = lower_16_bits(upper_32_bits(dpu->dtrace_buffer));
+	pkt->pkt_header.common_header.distribute = 0;
+	pkt->pkt_header.common_header.indirect = 0;
+	pkt->pkt_header.common_header.count = sizeof(*ebuf);
+}
+
+/*
+ * Build and submit one HSA command for @cmd_abo into the user host queue and
+ * ring the doorbell. Called with io_lock held.
+ *
+ * Security: cmd_abo is shared with user space; cache and validate its fields
+ * before use and never trust the queue content (only read_index is read back).
+ */
+static int submit_one_cmd(struct amdxdna_hwctx *hwctx,
+			  struct amdxdna_gem_obj *cmd_abo, bool last_of_chain,
+			  bool first_cmd, u64 *seq)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+	struct amdxdna_dev *xdna = hwctx->client->xdna;
+	struct amdxdna_cmd_start_dpu *dpu;
+	struct host_queue_packet *pkt;
+	u32 payload_size;
+	u64 slot_idx;
+	u16 chained;
+	int ret;
+	u32 op;
+
+	op = amdxdna_cmd_get_op(cmd_abo);
+	if (op != ERT_START_DPU) {
+		XDNA_ERR(xdna, "Invalid exec buf op, %d", op);
+		return -EINVAL;
+	}
+
+	dpu = amdxdna_cmd_get_payload(cmd_abo, &payload_size);
+	if (!dpu) {
+		XDNA_ERR(xdna, "Invalid DPU payload");
+		return -EINVAL;
+	}
+	/*
+	 * cmd_abo is shared with user space; validate the cached chained count
+	 * against the actual payload size before dereferencing chained+1 DPU
+	 * entries, so a bogus count cannot drive an out-of-bounds read.
+	 */
+	chained = dpu->chained;
+	if (chained >= HSA_MAX_LEVEL1_INDIRECT_ENTRIES) {
+		XDNA_ERR(xdna, "Invalid DPU data");
+		return -EINVAL;
+	}
+	if (payload_size < (u32)(chained + 1) * sizeof(*dpu)) {
+		XDNA_ERR(xdna, "DPU payload %u too small for %u entries",
+			 payload_size, chained + 1);
+		return -EINVAL;
+	}
+
+	/*
+	 * Block until a queue slot is free and the ctx is connected (io_lock is
+	 * dropped across the sleeps inside and re-acquired). The only failure is a
+	 * signal interrupting the wait (-ERESTARTSYS), e.g. the app being killed.
+	 */
+	ret = wait_till_connected_hsa_not_full(hwctx, first_cmd);
+	if (ret) {
+		XDNA_DBG(xdna, "Wait for queue slot / ctx reconnect interrupted, ret %d", ret);
+		return ret;
+	}
+
+	slot_idx = priv->write_index & (CTX_MAX_CMDS - 1);
+	if (chained) {
+		ret = fill_indirect_pkt(priv, slot_idx, CTX_MAX_CMDS, dpu, chained + 1);
+		if (ret)
+			return ret;
+	} else {
+		fill_direct_pkt(priv, slot_idx, dpu);
+	}
+
+	pkt = &priv->umq_pkts[slot_idx];
+	pkt->pkt_header.common_header.opcode = OPCODE_EXEC_BUF;
+	pkt->pkt_header.common_header.chain_flag =
+		last_of_chain ? CHAIN_FLG_LAST_CMD : CHAIN_FLG_NOT_LAST_CMD;
+	pkt->pkt_header.common_header.reserved = 0x0;
+	pkt->pkt_header.completion_signal = amdxdna_gem_dev_addr(cmd_abo) +
+					    offsetof(struct amdxdna_cmd, header);
+	*seq = publish_cmd(hwctx);
+	aie4_doorbell_ring(hwctx);
+	XDNA_DBG(xdna, "Submitted one cmd, %s seq %lld", hwctx->name, *seq);
+	return 0;
+}
+
+/*
+ * Return the head running job without removing it. The job worker keeps the
+ * in-flight job on the list while it waits so that a disconnect (suspend) can
+ * just leave it there for resume - no dequeue/requeue - and running_job_list is
+ * never transiently empty while a job is in flight.
+ */
+static struct amdxdna_sched_job *peek_running_job(struct amdxdna_hwctx *hwctx)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+	struct amdxdna_sched_job *job;
+
+	mutex_lock(&priv->io_lock);
+	job = list_first_entry_or_null(&priv->running_job_list,
+				       struct amdxdna_sched_job, aie4_job_list);
+	mutex_unlock(&priv->io_lock);
+	return job;
+}
+
+/* Remove a job from the running list once it is completed or reaped. */
+static void dequeue_running_job(struct amdxdna_hwctx *hwctx, struct amdxdna_sched_job *job)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+
+	mutex_lock(&priv->io_lock);
+	list_del(&job->aie4_job_list);
+	mutex_unlock(&priv->io_lock);
+}
+
+static void aie4_job_release(struct kref *ref)
+{
+	struct amdxdna_sched_job *job =
+		container_of(ref, struct amdxdna_sched_job, refcnt);
+
+	amdxdna_sched_job_cleanup(job);
+	if (job->out_fence)
+		dma_fence_put(job->out_fence);
+	kfree(job);
+}
+
+static void job_done(struct amdxdna_sched_job *job)
+{
+	job->aie4_job_state = AIE4_JOB_STATE_DONE;
+	dma_fence_signal(job->fence);
+	/*
+	 * Release the address-space reference taken at submit. On SVA/IOMMU
+	 * platforms the device walks the submitter's page tables while the job
+	 * runs, so its mm must stay alive until completion.
+	 */
+	mmput_async(job->mm);
+	kref_put(&job->refcnt, aie4_job_release);
+}
+
+static void job_complete(struct amdxdna_sched_job *job)
+{
+	job_done(job);
+}
+
+/*
+ * When CERT cannot complete a command (context teardown), the driver advances
+ * read_index so any waiter observes the command as finished. Only valid while
+ * the context is disconnected -- never race CERT's own read_index updates.
+ */
+static void update_read_index(struct amdxdna_hwctx *hwctx, u64 idx)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+
+	/* Order cmd-bo state write before the waiter observes completion. */
+	wmb();
+	WRITE_ONCE(*priv->umq_read_index, idx);
+}
+
+static void job_abort(struct amdxdna_sched_job *job)
+{
+	struct amdxdna_hwctx *hwctx = job->hwctx;
+
+	XDNA_WARN(hwctx->client->xdna, "aborting %s job %lld", hwctx->name, job->seq);
+	amdxdna_cmd_set_state(job->cmd_bo, ERT_CMD_STATE_ABORT);
+	dma_fence_set_error(job->fence, -ECANCELED);
+	/*
+	 * Only force read_index forward when CERT has not already moved it past this
+	 * job. On the reset drain read_index is still <= job->seq (CERT stopped), so
+	 * advance it here to release waiters. For a partial chain closed by a later
+	 * command's LAST_CMD, CERT has already advanced read_index past this job -
+	 * do not clobber it.
+	 */
+	if (get_read_index(hwctx) <= job->seq)
+		update_read_index(hwctx, job->seq + 1);
+	job_done(job);
+}
+
+static void job_worker(struct work_struct *work)
+{
+	struct amdxdna_hwctx_priv *priv =
+		container_of(work, struct amdxdna_hwctx_priv, job_work);
+	struct amdxdna_hwctx *hwctx = priv->hwctx;
+	struct amdxdna_sched_job *job;
+
+	while ((job = peek_running_job(hwctx))) {
+		wait_till_seq_completed(hwctx, job->seq);
+		if (get_read_index(hwctx) > job->seq) {
+			dequeue_running_job(hwctx, job);
+			/*
+			 * read_index advanced past this job. A fully published
+			 * job (SUBMITTED) ran to completion. A partial chain
+			 * (SUBMITTING: a later sub-command failed to publish, so
+			 * the chain never got CHAIN_FLG_LAST_CMD) only reaches
+			 * here once a *later* command's LAST_CMD closes the
+			 * dangling runlist and advances read_index past it - so
+			 * report it ABORT, not a false completion. If no such
+			 * command follows, read_index never advances and we stay
+			 * parked in wait_till_seq_completed() above until the
+			 * user's wait_command() times out and breaks the wait
+			 * (or ctx teardown reaps it).
+			 */
+			if (job->aie4_job_state != AIE4_JOB_STATE_SUBMITTED)
+				job_abort(job);
+			else
+				job_complete(job);
+		} else if (aie4_hwctx_has_reset(hwctx)) {
+			dequeue_running_job(hwctx, job);
+			job_abort(job);
+		} else {
+			/* suspend/resume */
+			break;
+		}
+	}
+}
+
+int aie4_hwctx_wait_for_running(struct amdxdna_hwctx *hwctx)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+	struct amdxdna_dev *xdna = hwctx->client->xdna;
+	struct amdxdna_sched_job *job;
+	long error;
+	int ret = 0;
+
+	mutex_lock(&priv->io_lock);
+	job = READ_ONCE(priv->pending_head);
+	if (job && job->aie4_job_state == AIE4_JOB_STATE_SUBMITTING) {
+		mutex_unlock(&priv->io_lock);
+		error = wait_event_timeout(priv->job_list_wq,
+					   READ_ONCE(priv->pending_head) != job,
+					   msecs_to_jiffies(2000));
+		if (!error) {
+			XDNA_WARN(xdna, "hwctx %s wait for submitting job timed out",
+				  hwctx->name);
+			ret = -ETIMEDOUT;
+		}
+	} else {
+		mutex_unlock(&priv->io_lock);
+	}
+
+	queue_work(priv->job_work_q, &priv->job_work);
+	flush_work(&priv->job_work);
+	return ret;
+}
+
+/*
+ * Submit the command(s) carried by @job into the host queue. Called with
+ * io_lock held. A single ERT_START_DPU maps to one queue entry; an
+ * ERT_CMD_CHAIN expands to one entry per sub-command, only the last of which
+ * carries CHAIN_FLG_LAST_CMD so CERT runs the whole chain back to back.
+ *
+ * job->seq tracks the last published sequence; the worker waits on it to reap
+ * the entire chain. job->aie4_job_state advances past PENDING as soon as any
+ * sub-command is published, so the caller knows whether in-flight commands must
+ * still be reaped even when a later sub-command fails to enqueue.
+ *
+ * Security: the chain payload and its BO handles come from user space; cache
+ * command_count and validate it against the payload size before walking the
+ * handle array so a bogus count cannot drive an out-of-bounds read.
+ */
+static int submit_job_cmds(struct amdxdna_hwctx *hwctx,
+			   struct amdxdna_sched_job *job, u32 op)
+{
+	struct amdxdna_gem_obj *cmd_abo = job->cmd_bo;
+	struct amdxdna_dev *xdna = hwctx->client->xdna;
+	struct amdxdna_cmd_chain *payload;
+	u32 payload_len, ccnt;
+	int ret;
+	u32 i;
+
+	/* Single cmd. */
+	if (op == ERT_START_DPU) {
+		ret = submit_one_cmd(hwctx, cmd_abo, true, true, &job->seq);
+		if (!ret)
+			job->aie4_job_state = AIE4_JOB_STATE_SUBMITTED;
+		return ret;
+	}
+
+	/* Cmd chain. */
+	payload = amdxdna_cmd_get_payload(cmd_abo, &payload_len);
+	if (!payload) {
+		XDNA_ERR(xdna, "Invalid cmd payload for chained cmd");
+		return -EINVAL;
+	}
+	ccnt = payload->command_count;
+	/*
+	 * A chain (runlist) must fit within the queue. CERT advances the host-visible
+	 * read_index only once per runlist - at the last sub-command (CHAIN_FLG_LAST_CMD),
+	 * not per sub-command - so a chain's own entries never free a queue slot until
+	 * the whole chain has been published and run. A chain longer than the queue
+	 * could therefore never publish its tail: it would block forever in
+	 * wait_till_connected_hsa_not_full() waiting for a slot that only frees at chain end.
+	 * Reject ccnt > CTX_MAX_CMDS. Also validate against the payload size before
+	 * walking the handle array so a bogus count cannot drive an out-of-bounds read.
+	 */
+	if (!ccnt || ccnt > CTX_MAX_CMDS ||
+	    payload_len < struct_size(payload, data, ccnt)) {
+		XDNA_ERR(xdna, "Invalid command count %u", ccnt);
+		return -EINVAL;
+	}
+
+	for (i = 0; i < ccnt; i++) {
+		u32 boh = (u32)(payload->data[i]);
+		struct amdxdna_gem_obj *abo;
+
+		abo = amdxdna_gem_get_obj(hwctx->client, boh, AMDXDNA_BO_SHARE);
+		if (!abo) {
+			XDNA_ERR(xdna, "Failed to find cmd BO %u", boh);
+			ret = -ENOENT;
+			break;
+		}
+
+		/*
+		 * submit_one_cmd() blocks in wait_till_connected_hsa_not_full() until the
+		 * ctx is connected and a slot is free, so a concurrent suspend/disconnect
+		 * is waited out inline rather than returned here. The first sub-command
+		 * (i == 0, nothing published yet) waits through a TDR reset and runs on
+		 * the recreated ctx; a later sub-command returns -ECONNRESET if a reset
+		 * landed while waiting for a slot, so the published prefix is not split
+		 * across the reset. It also returns -ERESTARTSYS on a signal, or a
+		 * validation error. Break on any; a published prefix is then reaped by
+		 * the job worker's reset drain (see below).
+		 */
+		ret = submit_one_cmd(hwctx, abo, i + 1 == ccnt, i == 0, &job->seq);
+		amdxdna_gem_put_obj(abo);
+		if (ret)
+			break;
+		job->aie4_job_state = AIE4_JOB_STATE_SUBMITTING;
+	}
+	if (i == ccnt)
+		job->aie4_job_state = AIE4_JOB_STATE_SUBMITTED;
+
+	/*
+	 * As long as at least one sub-command was published, return success so the
+	 * caller enqueues the job on the running list; the job worker then reaps the
+	 * published prefix and reports the partial chain as failed (ABORT). Only when
+	 * nothing was published (i == 0) is the error returned to the caller.
+	 */
+	if (i > 0)
+		return 0;
+
+	return ret;
+}
+
+/*
+ * Whole-job submission is serialized across submitters that share a ctx via the
+ * pending list: a job is appended on entry and only the head of the list is
+ * allowed to publish its command(s). Because the head stays on the list for the
+ * entire duration of submit_job_cmds() -- which may drop io_lock to wait for
+ * free queue slots -- no other submitter can interleave its commands into the
+ * middle of the head job's command chain. io_lock protects the lists; the
+ * job_list_wq waitqueue notifies parked submitters when the head changes.
+ */
+/* Publish the current pending-list head for the lockless submit wait condition.
+ * Caller holds io_lock.
+ */
+static void update_pending_head(struct amdxdna_hwctx_priv *priv)
+{
+	WRITE_ONCE(priv->pending_head,
+		   list_first_entry_or_null(&priv->pending_job_list,
+					    struct amdxdna_sched_job, aie4_job_list));
+}
+
+static void enqueue_pending_job(struct amdxdna_hwctx *hwctx,
+				struct amdxdna_sched_job *job)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+
+	mutex_lock(&priv->io_lock);
+	list_add_tail(&job->aie4_job_list, &priv->pending_job_list);
+	job->aie4_job_state = AIE4_JOB_STATE_PENDING;
+	update_pending_head(priv);
+	mutex_unlock(&priv->io_lock);
+
+	/* Let the next pending submitter re-check whether it is now first. */
+	wake_up_all(&priv->job_list_wq);
+}
+
+static void cancel_pending_job(struct amdxdna_hwctx *hwctx,
+			       struct amdxdna_sched_job *job)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+
+	mutex_lock(&priv->io_lock);
+	list_del(&job->aie4_job_list);
+	job->aie4_job_state = AIE4_JOB_STATE_INIT;
+	update_pending_head(priv);
+	mutex_unlock(&priv->io_lock);
+	/* Let the next pending submitter re-check whether it is now first. */
+	wake_up_all(&priv->job_list_wq);
+}
+
+int aie4_cmd_submit(struct amdxdna_hwctx *hwctx, struct amdxdna_sched_job *job, u64 *seq)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+	struct amdxdna_dev *xdna = hwctx->client->xdna;
+	struct ww_acquire_ctx acquire_ctx;
+	struct amdxdna_gem_obj *abo;
+	u32 op;
+	int i;
+	int ret;
+
+	XDNA_DBG(xdna, "ctx %s job %p received", hwctx->name, job);
+
+	if (!job->cmd_bo) {
+		XDNA_ERR(xdna, "No command BO in job");
+		return -EINVAL;
+	}
+
+	op = amdxdna_cmd_get_op(job->cmd_bo);
+	if (op != ERT_START_DPU && op != ERT_CMD_CHAIN) {
+		XDNA_ERR(xdna, "Invalid cmd opcode %d", op);
+		return -EINVAL;
+	}
+
+	INIT_LIST_HEAD(&job->aie4_job_list);
+
+	/*
+	 * Hold a reference on the submitter's address space until the job
+	 * completes (job_done): on SVA/IOMMU platforms the device walks the
+	 * submitter's page tables while the command runs. Balanced with the
+	 * mmput_async() in job_done() and the mmput() on the failure paths below.
+	 */
+	if (!mmget_not_zero(job->mm)) {
+		XDNA_ERR(xdna, "Failed to get mm reference");
+		return -ESRCH;
+	}
+
+	/*
+	 * Lock all job BOs and reserve fences. This attaches job->out_fence
+	 * to each BO's reservation object, ensuring concurrent invalidation
+	 * waits for the job to complete.
+	 */
+	ret = drm_gem_lock_reservations(job->bos, job->bo_cnt, &acquire_ctx);
+	if (ret) {
+		XDNA_WARN(xdna, "Failed to lock BOs, ret %d", ret);
+		goto put_mm;
+	}
+
+	for (i = 0; i < job->bo_cnt; i++) {
+		ret = dma_resv_reserve_fences(job->bos[i]->resv, 1);
+		if (ret) {
+			XDNA_WARN(xdna, "Failed to reserve fences %d", ret);
+			drm_gem_unlock_reservations(job->bos, job->bo_cnt, &acquire_ctx);
+			goto put_mm;
+		}
+	}
+
+	down_read(&xdna->notifier_lock);
+	for (i = 0; i < job->bo_cnt; i++) {
+		abo = to_xdna_obj(job->bos[i]);
+		if (abo->mem.map_invalid) {
+			up_read(&xdna->notifier_lock);
+			drm_gem_unlock_reservations(job->bos, job->bo_cnt, &acquire_ctx);
+			ret = -EINVAL;
+			goto put_mm;
+		}
+	}
+
+	job->out_fence = dma_fence_get(job->fence);
+	for (i = 0; i < job->bo_cnt; i++)
+		dma_resv_add_fence(job->bos[i]->resv, job->out_fence, DMA_RESV_USAGE_WRITE);
+
+	up_read(&xdna->notifier_lock);
+	drm_gem_unlock_reservations(job->bos, job->bo_cnt, &acquire_ctx);
+
+	/*
+	 * Wait until this job is at the head of the pending list before touching
+	 * the queue (see enqueue_pending_job). Freezable so the freezer can
+	 * suspend a parked submitter in place across S3/S4 rather than aborting
+	 * the suspend; still interruptible so a signal (app exit/kill/^C) unwinds
+	 * it and does not keep ctx teardown (synchronize_srcu) blocked.
+	 */
+	enqueue_pending_job(hwctx, job);
+	ret = wait_event_freezable(priv->job_list_wq,
+				   READ_ONCE(priv->pending_head) == job);
+	if (ret) {
+		cancel_pending_job(hwctx, job);
+		goto signal_fence;
+	}
+
+	mutex_lock(&priv->io_lock);
+	ret = submit_job_cmds(hwctx, job, op);
+	if (ret) {
+		mutex_unlock(&priv->io_lock);
+		cancel_pending_job(hwctx, job);
+		goto signal_fence;
+	}
+
+	list_move_tail(&job->aie4_job_list, &priv->running_job_list);
+	update_pending_head(priv);
+	*seq = job->seq;
+	mutex_unlock(&priv->io_lock);
+
+	/* Release the next pending submitter and kick the reaper. */
+	wake_up_all(&priv->job_list_wq);
+	atomic64_inc(&hwctx->job_submit_cnt);
+	queue_work(priv->job_work_q, &priv->job_work);
+	return 0;
+
+signal_fence:
+	/*
+	 * Map internal -ERESTARTSYS to -ECANCELED for the fence so downstream
+	 * consumers (e.g. sync_file, dma-buf importers) do not observe internal
+	 * signal restart codes, while preserving ret for the syscall return.
+	 */
+	dma_fence_set_error(job->fence, ret == -ERESTARTSYS ? -ECANCELED : ret);
+	dma_fence_signal(job->fence);
+	dma_fence_put(job->out_fence);
+	job->out_fence = NULL;
+put_mm:
+	mmput(job->mm);
+	return ret;
+}
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index f5fdc24689f8..f180983a692d 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -1109,6 +1109,7 @@ const struct amdxdna_dev_ops aie4_vf_ops = {
 	.debugfs_init		= aie4_debugfs_init,
 	.hwctx_init		= aie4_hwctx_init,
 	.hwctx_fini		= aie4_hwctx_fini,
+	.cmd_submit		= aie4_cmd_submit,
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
 	.set_aie_state		= aie4_set_state,
@@ -1120,6 +1121,7 @@ const struct amdxdna_dev_ops aie4_classic_ops = {
 	.debugfs_init		= aie4_debugfs_init,
 	.hwctx_init		= aie4_hwctx_init,
 	.hwctx_fini		= aie4_hwctx_fini,
+	.cmd_submit		= aie4_cmd_submit,
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
 	.set_aie_state		= aie4_set_state,
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index f549d9e69d41..6b67c9da560e 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -50,6 +50,7 @@ struct amdxdna_hwctx_priv {
 
 	struct cert_comp                *cert_comp;
 	u32                             hw_ctx_id;
+	bool                            has_reset;
 
 	/* Kernel-mode submission: driver fills the user HSA queue and rings
 	 * the doorbell.  umq_pkts/umq_indirect_pkts alias the user umq_bo;
@@ -165,8 +166,10 @@ enum aie4_hwctx_flags {
 int aie4_hwctx_init(struct amdxdna_hwctx *hwctx);
 void aie4_hwctx_fini(struct amdxdna_hwctx *hwctx);
 int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout);
+int aie4_cmd_submit(struct amdxdna_hwctx *hwctx, struct amdxdna_sched_job *job, u64 *seq);
 int aie4_hwctx_create(struct amdxdna_hwctx *hwctx);
 void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx, enum aie4_hwctx_flags);
+int aie4_hwctx_wait_for_running(struct amdxdna_hwctx *hwctx);
 
 /* aie4_pci.c */
 int aie4_restore_power_mode(struct amdxdna_dev_hdl *ndev);
diff --git a/drivers/accel/amdxdna/amdxdna_ctx.c b/drivers/accel/amdxdna/amdxdna_ctx.c
index 888e857ec558..2a635eab1e55 100644
--- a/drivers/accel/amdxdna/amdxdna_ctx.c
+++ b/drivers/accel/amdxdna/amdxdna_ctx.c
@@ -25,7 +25,7 @@
 struct amdxdna_fence {
 	struct dma_fence	base;
 	spinlock_t		lock; /* for base */
-	struct amdxdna_hwctx	*hwctx;
+	struct device		*dev;
 };
 
 static const char *amdxdna_fence_get_driver_name(struct dma_fence *fence)
@@ -39,7 +39,14 @@ static const char *amdxdna_fence_get_timeline_name(struct dma_fence *fence)
 
 	xdna_fence = container_of(fence, struct amdxdna_fence, base);
 
-	return xdna_fence->hwctx->name;
+	/*
+	 * Use device name rather than hwctx name: the fence is published into
+	 * BO reservation objects via dma_resv_add_fence() and can outlive the
+	 * hwctx (e.g. when a BO is exported as a dma-buf and imported by
+	 * another process). The device outlives any individual context, so
+	 * dev_name() is safe to call at any point during the fence's lifetime.
+	 */
+	return dev_name(xdna_fence->dev);
 }
 
 static const struct dma_fence_ops fence_ops = {
@@ -55,9 +62,14 @@ static struct dma_fence *amdxdna_fence_create(struct amdxdna_hwctx *hwctx)
 	if (!fence)
 		return NULL;
 
-	fence->hwctx = hwctx;
+	fence->dev = hwctx->client->xdna->ddev.dev;
 	spin_lock_init(&fence->lock);
-	dma_fence_init(&fence->base, &fence_ops, &fence->lock, hwctx->id, 0);
+	/*
+	 * Each job fence needs a unique timeline context so dma_resv_add_fence()
+	 * does not evict a prior in-flight job's fence from a shared BO's
+	 * reservation object when concurrent submissions access the same BO.
+	 */
+	dma_fence_init(&fence->base, &fence_ops, &fence->lock, dma_fence_context_alloc(1), 0);
 	return &fence->base;
 }
 
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 15/20] accel/amdxdna: Finalize runtime PM before acquiring dev_lock on removal
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (13 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 14/20] accel/amdxdna: Implement AIE4 command packet building and submission David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:32 ` [PATCH V1 16/20] accel/amdxdna: Implement AIE4 suspend and resume David Zhang
                   ` (4 subsequent siblings)
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

When the device is runtime-suspended, pm_runtime_forbid() synchronously
resumes the device via rpm_resume(), which invokes
amdxdna_pm_runtime_resume(). Because amdxdna_pm_runtime_resume()
acquires dev_lock, calling amdxdna_pm_fini() inside ops->fini() while
holding dev_lock in amdxdna_remove() causes a deadlock.

Move amdxdna_pm_fini() out of ops->fini() and invoke it before acquiring
dev_lock in amdxdna_remove() as well as the probe failure unwind path.
Also call pm_runtime_dont_use_autosuspend() in amdxdna_pm_fini() to
disable autosuspend upon teardown.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie2_pci.c        | 2 --
 drivers/accel/amdxdna/amdxdna_pci_drv.c | 4 ++++
 drivers/accel/amdxdna/amdxdna_pm.c      | 1 +
 3 files changed, 5 insertions(+), 2 deletions(-)

diff --git a/drivers/accel/amdxdna/aie2_pci.c b/drivers/accel/amdxdna/aie2_pci.c
index b70af1923643..0d209b7b6484 100644
--- a/drivers/accel/amdxdna/aie2_pci.c
+++ b/drivers/accel/amdxdna/aie2_pci.c
@@ -623,7 +623,6 @@ static int aie2_init(struct amdxdna_dev *xdna)
 	release_firmware(fw);
 	aie2_msg_init(ndev);
 	amdxdna_vbnv_init(xdna);
-	amdxdna_pm_init(xdna);
 	return 0;
 
 stop_hw:
@@ -638,7 +637,6 @@ static int aie2_init(struct amdxdna_dev *xdna)
 
 static void aie2_fini(struct amdxdna_dev *xdna)
 {
-	amdxdna_pm_fini(xdna);
 	aie2_hw_stop(xdna);
 	aie2_hwctx_sched_fini(xdna->dev_handle);
 }
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.c b/drivers/accel/amdxdna/amdxdna_pci_drv.c
index 8b6e7283e057..1d0b91e73260 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.c
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.c
@@ -413,6 +413,8 @@ static int amdxdna_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 		goto iommu_fini;
 	}
 
+	amdxdna_pm_init(xdna);
+
 	ret = amdxdna_sysfs_init(xdna);
 	if (ret) {
 		XDNA_ERR(xdna, "Create amdxdna attrs failed: %d", ret);
@@ -431,6 +433,7 @@ static int amdxdna_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 failed_sysfs_fini:
 	amdxdna_sysfs_fini(xdna);
 failed_dev_fini:
+	amdxdna_pm_fini(xdna);
 	mutex_lock(&xdna->dev_lock);
 	xdna->dev_info->ops->fini(xdna);
 	mutex_unlock(&xdna->dev_lock);
@@ -446,6 +449,7 @@ static void amdxdna_remove(struct pci_dev *pdev)
 
 	drm_dev_unplug(&xdna->ddev);
 	amdxdna_sysfs_fini(xdna);
+	amdxdna_pm_fini(xdna);
 
 	mutex_lock(&xdna->client_lock);
 	mutex_lock(&xdna->dev_lock);
diff --git a/drivers/accel/amdxdna/amdxdna_pm.c b/drivers/accel/amdxdna/amdxdna_pm.c
index b1fafddd7ad5..9c030b7836fb 100644
--- a/drivers/accel/amdxdna/amdxdna_pm.c
+++ b/drivers/accel/amdxdna/amdxdna_pm.c
@@ -75,4 +75,5 @@ void amdxdna_pm_fini(struct amdxdna_dev *xdna)
 
 	pm_runtime_get_noresume(dev);
 	pm_runtime_forbid(dev);
+	pm_runtime_dont_use_autosuspend(dev);
 }
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 16/20] accel/amdxdna: Implement AIE4 suspend and resume
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (14 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 15/20] accel/amdxdna: Finalize runtime PM before acquiring dev_lock on removal David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  4:07   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 17/20] accel/amdxdna: Link SR-IOV VFs for power management sequencing David Zhang
                   ` (3 subsequent siblings)
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

Implement suspend and resume callbacks for AIE4 Physical Function (PF),
Virtual Function (VF), and Classic device types:
- Add .suspend and .resume hooks in amdxdna_dev_ops for aie4_pf_ops,
  aie4_vf_ops, and aie4_classic_ops.
- Implement aie4_hwctx_suspend_all() to destroy or drain contexts across
  all registered clients and wait for in-flight jobs.
- Implement aie4_hwctx_resume_all() to recreate firmware contexts and
  kick doorbells via aie4_hwctx_resume_jobs() to resume hardware queue
  consumption.
- Guard firmware destroy message in aie4_hwctx_destroy() when context
  ID is invalid.
- Restore SR-IOV virtual functions on PF resume via aie4_restore_sriov().

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_ctx.c        |  18 +-
 drivers/accel/amdxdna/aie4_pci.c        | 243 ++++++++++++++++++++++++
 drivers/accel/amdxdna/aie4_pci.h        |   9 +
 drivers/accel/amdxdna/aie4_sriov.c      |   4 +-
 drivers/accel/amdxdna/amdxdna_pci_drv.h |   3 +
 5 files changed, 275 insertions(+), 2 deletions(-)

diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
index e9ba1ed93997..c55138a754fd 100644
--- a/drivers/accel/amdxdna/aie4_ctx.c
+++ b/drivers/accel/amdxdna/aie4_ctx.c
@@ -262,7 +262,7 @@ void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx, enum aie4_hwctx_flags flags
 	if (has_reset)
 		wake_up_all(&priv->job_list_wq);
 
-	if (flags != AIE4_HWCTX_DISCONNECT)
+	if (flags != AIE4_HWCTX_DISCONNECT && priv->hw_ctx_id != CTX_INVALID_ID)
 		aie4_msg_destroy_context(ndev, priv->hw_ctx_id);
 
 	priv->hw_ctx_id = CTX_INVALID_ID;
@@ -393,6 +393,7 @@ int aie4_hwctx_init(struct amdxdna_hwctx *hwctx)
 		return -ENOMEM;
 	hwctx->priv = priv;
 	priv->hwctx = hwctx;
+	priv->hw_ctx_id = CTX_INVALID_ID;
 
 	/*
 	 * io_lock guards the per-hwctx cert_comp binding (the connected sentinel)
@@ -974,6 +975,21 @@ int aie4_hwctx_wait_for_running(struct amdxdna_hwctx *hwctx)
 	return ret;
 }
 
+void aie4_hwctx_resume_jobs(struct amdxdna_hwctx *hwctx)
+{
+	struct amdxdna_hwctx_priv *priv = hwctx->priv;
+
+	mutex_lock(&priv->io_lock);
+	if (list_empty(&priv->running_job_list)) {
+		mutex_unlock(&priv->io_lock);
+		return;
+	}
+	aie4_doorbell_ring(hwctx);
+	mutex_unlock(&priv->io_lock);
+
+	queue_work(priv->job_work_q, &priv->job_work);
+}
+
 /*
  * Submit the command(s) carried by @job into the host queue. Called with
  * io_lock held. A single ERT_START_DPU maps to one queue entry; an
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index f180983a692d..6a50c1499ec9 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -1096,11 +1096,250 @@ static void aie4_debugfs_init(struct amdxdna_dev *xdna)
 					   &aie4_ctx_hysteresis_fops);
 }
 
+void aie4_hwctx_suspend_all(struct amdxdna_dev_hdl *ndev, int clean_jobs)
+{
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	struct amdxdna_client *client;
+	struct amdxdna_hwctx *hwctx;
+	unsigned long hwctx_id;
+	int idx;
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+
+	amdxdna_for_each_client(xdna, client) {
+		idx = srcu_read_lock(&client->hwctx_srcu);
+		amdxdna_for_each_hwctx(client, hwctx_id, hwctx) {
+			/* clean up workers and drain running jobs */
+			if (clean_jobs) {
+				int ret;
+
+				aie4_hwctx_destroy(hwctx, AIE4_HWCTX_ERROR);
+				ret = aie4_hwctx_wait_for_running(hwctx);
+				if (ret)
+					XDNA_WARN(xdna, "hwctx %s wait for running failed %d",
+						  hwctx->name, ret);
+			} else {
+				aie4_hwctx_destroy(hwctx, AIE4_HWCTX_NORMAL);
+			}
+		}
+		srcu_read_unlock(&client->hwctx_srcu, idx);
+	}
+
+	XDNA_DBG(xdna, "Finished hwctx suspend");
+}
+
+int aie4_hwctx_resume_all(struct amdxdna_dev_hdl *ndev)
+{
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	struct amdxdna_client *client;
+	struct amdxdna_hwctx *hwctx;
+	unsigned long hwctx_id;
+	int ret, idx;
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+
+	amdxdna_for_each_client(xdna, client) {
+		idx = srcu_read_lock(&client->hwctx_srcu);
+		amdxdna_for_each_hwctx(client, hwctx_id, hwctx) {
+			ret = aie4_hwctx_create(hwctx);
+			if (ret)
+				goto error;
+			aie4_hwctx_resume_jobs(hwctx);
+		}
+		srcu_read_unlock(&client->hwctx_srcu, idx);
+	}
+
+	XDNA_DBG(xdna, "Finished hwctx resume");
+	return 0;
+error:
+	srcu_read_unlock(&client->hwctx_srcu, idx);
+	XDNA_DBG(xdna, "Failed hwctx resume");
+	return ret;
+}
+
+static int aie4_restore_sriov(struct amdxdna_dev_hdl *ndev)
+{
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
+	int ret;
+
+	if (ndev->num_vfs) {
+		if (pci_num_vf(pdev) != ndev->num_vfs) {
+			XDNA_ERR(xdna, "inconsistent vf number");
+			return -EINVAL;
+		}
+		ret = aie4_create_vfs(ndev, ndev->num_vfs);
+		if (ret) {
+			XDNA_ERR(xdna, "create vfs failed, %d", ret);
+			return ret;
+		}
+		XDNA_DBG(xdna, "restored num_vfs %d", ndev->num_vfs);
+	}
+
+	return 0;
+}
+
+static int aie4_pf_suspend(struct amdxdna_dev *xdna)
+{
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+	aie4_pf_hw_stop(ndev);
+	pci_disable_device(pdev);
+
+	XDNA_DBG(xdna, "pf suspend done");
+	return 0;
+}
+
+static int aie4_pf_resume(struct amdxdna_dev *xdna)
+{
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
+	int ret;
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+
+	ret = pci_enable_device(pdev);
+	if (ret) {
+		XDNA_ERR(xdna, "enable pci device failed %d", ret);
+		return ret;
+	}
+	pci_set_master(pdev);
+
+	ret = aie4_pf_hw_start(ndev);
+	if (ret) {
+		XDNA_ERR(xdna, "hw_start failed %d", ret);
+		goto pci_disable;
+	}
+
+	ret = aie4_restore_sriov(ndev);
+	if (ret)
+		goto hw_stop;
+
+	XDNA_DBG(xdna, "pf resume done");
+	return 0;
+hw_stop:
+	aie4_pf_hw_stop(ndev);
+pci_disable:
+	pci_disable_device(pdev);
+	return ret;
+}
+
+static int aie4_vf_suspend(struct amdxdna_dev *xdna)
+{
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+	aie4_hwctx_suspend_all(ndev, false);
+	/*
+	 * partition_fini and mailbox messages should not be called here
+	 * because PF suspend will do the cleanup for all VFs.
+	 */
+	aie4_mailbox_fini(ndev);
+	pci_disable_device(pdev);
+
+	XDNA_DBG(xdna, "vf suspend done");
+	return 0;
+}
+
+static int aie4_vf_resume(struct amdxdna_dev *xdna)
+{
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
+	int ret;
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+
+	ret = pci_enable_device(pdev);
+	if (ret) {
+		XDNA_ERR(xdna, "enable pci device failed %d", ret);
+		return ret;
+	}
+	pci_set_master(pdev);
+
+	ret = aie4_vf_hw_start(ndev);
+	if (ret) {
+		XDNA_ERR(xdna, "hw_start failed %d", ret);
+		goto pci_disable;
+	}
+
+	ret = aie4_hwctx_resume_all(ndev);
+	if (ret) {
+		XDNA_ERR(xdna, "hwctx_resume failed %d", ret);
+		goto hw_clear;
+	}
+
+	XDNA_DBG(xdna, "vf resume done");
+	return 0;
+
+hw_clear:
+	aie4_hwctx_suspend_all(ndev, true);
+	aie4_vf_hw_stop(ndev);
+pci_disable:
+	pci_disable_device(pdev);
+	return ret;
+}
+
+static int aie4_classic_suspend(struct amdxdna_dev *xdna)
+{
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+	aie4_hwctx_suspend_all(ndev, false);
+	aie4_classic_hw_stop(ndev);
+	pci_disable_device(pdev);
+
+	XDNA_DBG(xdna, "classic suspend done");
+	return 0;
+}
+
+static int aie4_classic_resume(struct amdxdna_dev *xdna)
+{
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
+	int ret;
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+
+	ret = pci_enable_device(pdev);
+	if (ret) {
+		XDNA_ERR(xdna, "enable pci device failed %d", ret);
+		return ret;
+	}
+	pci_set_master(pdev);
+
+	ret = aie4_classic_hw_start(ndev);
+	if (ret) {
+		XDNA_ERR(xdna, "hw_start failed %d", ret);
+		goto pci_disable;
+	}
+
+	ret = aie4_hwctx_resume_all(ndev);
+	if (ret) {
+		XDNA_ERR(xdna, "hwctx_resume failed %d", ret);
+		goto hw_clear;
+	}
+
+	XDNA_DBG(xdna, "classic resume done");
+	return 0;
+hw_clear:
+	aie4_hwctx_suspend_all(ndev, true);
+	aie4_classic_hw_stop(ndev);
+pci_disable:
+	pci_disable_device(pdev);
+	return ret;
+}
+
 const struct amdxdna_dev_ops aie4_pf_ops = {
 	.init			= aie4_pf_init,
 	.fini			= aie4_pf_fini,
 	.debugfs_init		= aie4_debugfs_init,
 	.sriov_configure        = aie4_sriov_configure,
+	.resume			= aie4_pf_resume,
+	.suspend		= aie4_pf_suspend,
 };
 
 const struct amdxdna_dev_ops aie4_vf_ops = {
@@ -1113,6 +1352,8 @@ const struct amdxdna_dev_ops aie4_vf_ops = {
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
 	.set_aie_state		= aie4_set_state,
+	.resume			= aie4_vf_resume,
+	.suspend		= aie4_vf_suspend,
 };
 
 const struct amdxdna_dev_ops aie4_classic_ops = {
@@ -1125,4 +1366,6 @@ const struct amdxdna_dev_ops aie4_classic_ops = {
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
 	.set_aie_state		= aie4_set_state,
+	.resume			= aie4_classic_resume,
+	.suspend		= aie4_classic_suspend,
 };
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index 6b67c9da560e..df15c63317b1 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -109,6 +109,7 @@ struct amdxdna_dev_hdl {
 	u32				total_col;
 	u32				max_aieclk_level;
 	u32				max_npuhclk_level;
+	u32				num_vfs;
 
 	struct dpm_clk_freq		dpm_clk_tbl[AIE4_MAX_DPM_LEVEL_COUNT];
 
@@ -170,8 +171,11 @@ int aie4_cmd_submit(struct amdxdna_hwctx *hwctx, struct amdxdna_sched_job *job,
 int aie4_hwctx_create(struct amdxdna_hwctx *hwctx);
 void aie4_hwctx_destroy(struct amdxdna_hwctx *hwctx, enum aie4_hwctx_flags);
 int aie4_hwctx_wait_for_running(struct amdxdna_hwctx *hwctx);
+void aie4_hwctx_resume_jobs(struct amdxdna_hwctx *hwctx);
 
 /* aie4_pci.c */
+void aie4_hwctx_suspend_all(struct amdxdna_dev_hdl *ndev, int clean_jobs);
+int aie4_hwctx_resume_all(struct amdxdna_dev_hdl *ndev);
 int aie4_restore_power_mode(struct amdxdna_dev_hdl *ndev);
 
 /*
@@ -191,9 +195,14 @@ void aie4_free_notification(struct cert_comp *comp);
 
 /* aie4_sriov.c */
 #if IS_ENABLED(CONFIG_PCI_IOV)
+int aie4_create_vfs(struct amdxdna_dev_hdl *ndev, int num_vfs);
 int aie4_sriov_configure(struct amdxdna_dev *xdna, int num_vfs);
 int aie4_sriov_stop(struct amdxdna_dev_hdl *ndev);
 #else
+static inline int aie4_create_vfs(struct amdxdna_dev_hdl *ndev, int num_vfs)
+{
+	return 0;
+}
 #define aie4_sriov_configure NULL
 static inline int aie4_sriov_stop(struct amdxdna_dev_hdl *ndev)
 {
diff --git a/drivers/accel/amdxdna/aie4_sriov.c b/drivers/accel/amdxdna/aie4_sriov.c
index e1ce633768a5..0eea28f62676 100644
--- a/drivers/accel/amdxdna/aie4_sriov.c
+++ b/drivers/accel/amdxdna/aie4_sriov.c
@@ -26,7 +26,7 @@ static int aie4_destroy_vfs(struct amdxdna_dev_hdl *ndev)
 	return ret;
 }
 
-static int aie4_create_vfs(struct amdxdna_dev_hdl *ndev, int num_vfs)
+int aie4_create_vfs(struct amdxdna_dev_hdl *ndev, int num_vfs)
 {
 	DECLARE_AIE_MSG(aie4_msg_create_vfs, AIE4_MSG_OP_CREATE_VFS);
 	int ret;
@@ -55,6 +55,7 @@ int aie4_sriov_stop(struct amdxdna_dev_hdl *ndev)
 	}
 
 	pci_disable_sriov(pdev);
+	ndev->num_vfs = 0;
 	return aie4_destroy_vfs(ndev);
 }
 
@@ -75,6 +76,7 @@ static int aie4_sriov_start(struct amdxdna_dev_hdl *ndev, int num_vfs)
 		return ret;
 	}
 
+	ndev->num_vfs = num_vfs;
 	return num_vfs;
 }
 
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.h b/drivers/accel/amdxdna/amdxdna_pci_drv.h
index 11f46ec738d7..45170c945d71 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.h
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.h
@@ -168,6 +168,9 @@ struct amdxdna_client {
 #define amdxdna_for_each_hwctx(client, hwctx_id, entry)		\
 	xa_for_each(&(client)->hwctx_xa, hwctx_id, entry)
 
+#define amdxdna_for_each_client(xdna, client)			\
+	list_for_each_entry(client, &(xdna)->client_list, node)
+
 /* Add device info below */
 extern const struct amdxdna_dev_info dev_npu1_info;
 extern const struct amdxdna_dev_info dev_npu3_classic_info;
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 17/20] accel/amdxdna: Link SR-IOV VFs for power management sequencing
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (15 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 16/20] accel/amdxdna: Implement AIE4 suspend and resume David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:59   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 18/20] accel/amdxdna: Implement runtime suspend and resume support David Zhang
                   ` (2 subsequent siblings)
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

Add PM device links between Physical Function (PF) supplier and Virtual
Function (VF) consumers via device_link_add() upon SR-IOV enablement.
This ensures the PM core enforces the proper power management sequence:
suspending VFs before the PF, and resuming the PF before VFs. If linking
fails, roll back SR-IOV initialization.

Also add a comment clarifying that pci_disable_sriov() removes VF drivers
before firmware VF contexts are destroyed.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_sriov.c | 74 ++++++++++++++++++++++++++++++
 1 file changed, 74 insertions(+)

diff --git a/drivers/accel/amdxdna/aie4_sriov.c b/drivers/accel/amdxdna/aie4_sriov.c
index 0eea28f62676..bfea6ff00ec0 100644
--- a/drivers/accel/amdxdna/aie4_sriov.c
+++ b/drivers/accel/amdxdna/aie4_sriov.c
@@ -56,9 +56,75 @@ int aie4_sriov_stop(struct amdxdna_dev_hdl *ndev)
 
 	pci_disable_sriov(pdev);
 	ndev->num_vfs = 0;
+
+	/*
+	 * pci_disable_sriov() removes VF drivers first; call destroy_vfs after
+	 * so firmware VF contexts are not cleared before VF drivers finish cleanup.
+	 */
 	return aie4_destroy_vfs(ndev);
 }
 
+static int aie4_for_each_vfs(struct amdxdna_dev *xdna,
+			     int (*cb)(struct amdxdna_dev *, struct pci_dev *))
+{
+	struct pci_dev *pdev_pf = to_pci_dev(xdna->ddev.dev);
+	struct pci_dev *pdev_vf;
+	int pos, ret;
+	u16 vf_did;
+
+	pos = pci_find_ext_capability(pdev_pf, PCI_EXT_CAP_ID_SRIOV);
+	if (!pos)
+		return 0;
+	ret = pci_read_config_word(pdev_pf, pos + PCI_SRIOV_VF_DID, &vf_did);
+	if (ret) {
+		XDNA_ERR(xdna, "read VF Device ID failed %d", ret);
+		return -ENODEV;
+	}
+
+	for (pdev_vf = pci_get_device(pdev_pf->vendor, vf_did, NULL);
+	     pdev_vf;
+	     pdev_vf = pci_get_device(pdev_pf->vendor, vf_did, pdev_vf)) {
+		if (!pdev_vf->is_virtfn || pdev_vf->physfn != pdev_pf)
+			continue;
+
+		ret = cb(xdna, pdev_vf);
+		if (ret) {
+			/*
+			 * On early return the next iteration never runs, so
+			 * release the current device's ref manually.
+			 * On normal loop exit pci_get_device() returning NULL
+			 * already releases the last device's ref internally.
+			 */
+			pci_dev_put(pdev_vf);
+			return ret;
+		}
+	}
+
+	return 0;
+}
+
+static int aie4_link_vf(struct amdxdna_dev *xdna, struct pci_dev *pdev_vf)
+{
+	struct pci_dev *pdev_pf = to_pci_dev(xdna->ddev.dev);
+	struct device_link *link;
+
+	link = device_link_add(&pdev_vf->dev,   /* consumer = VF */
+			       &pdev_pf->dev,   /* supplier = PF */
+			       DL_FLAG_PM_RUNTIME | DL_FLAG_AUTOREMOVE_CONSUMER);
+	if (!link) {
+		XDNA_ERR(xdna, "Failed to link VF %s", pci_name(pdev_vf));
+		return -EINVAL;
+	}
+
+	XDNA_DBG(xdna, "Linked VF %s", pci_name(pdev_vf));
+	return 0;
+}
+
+static int aie4_link_vfs(struct amdxdna_dev *xdna)
+{
+	return aie4_for_each_vfs(xdna, aie4_link_vf);
+}
+
 static int aie4_sriov_start(struct amdxdna_dev_hdl *ndev, int num_vfs)
 {
 	struct amdxdna_dev *xdna = ndev->aie.xdna;
@@ -76,6 +142,14 @@ static int aie4_sriov_start(struct amdxdna_dev_hdl *ndev, int num_vfs)
 		return ret;
 	}
 
+	ret = aie4_link_vfs(xdna);
+	if (ret) {
+		XDNA_ERR(xdna, "link VFs failed, ret: %d", ret);
+		pci_disable_sriov(pdev);
+		aie4_destroy_vfs(ndev);
+		return ret;
+	}
+
 	ndev->num_vfs = num_vfs;
 	return num_vfs;
 }
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 18/20] accel/amdxdna: Implement runtime suspend and resume support
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (16 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 17/20] accel/amdxdna: Link SR-IOV VFs for power management sequencing David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  4:00   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 19/20] accel/amdxdna: Add stub hwctx_config for AIE4 David Zhang
  2026-09-30  3:32 ` [PATCH V1 20/20] accel/amdxdna: Enable AIE4 firmware logging to DRAM David Zhang
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

Add runtime suspend/resume for AIE4 driver.

Wake and acquire an RPM reference across amdxdna_sriov_configure() so
that SR-IOV management commands execute with the device active.

Update amdxdna_pm.c to implement amdxdna_pm_runtime_suspend() and
amdxdna_pm_runtime_resume().

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie2_pci.c        |  2 ++
 drivers/accel/amdxdna/aie4_pci.c        | 31 ++++++++++++++++++++++
 drivers/accel/amdxdna/aie4_pci.h        |  8 ++++++
 drivers/accel/amdxdna/aie4_sriov.c      | 26 +++++++++++++++++++
 drivers/accel/amdxdna/amdxdna_pci_drv.c | 15 ++++++++---
 drivers/accel/amdxdna/amdxdna_pci_drv.h |  2 ++
 drivers/accel/amdxdna/amdxdna_pm.c      | 34 +++++++++++++++++++++++++
 drivers/accel/amdxdna/amdxdna_pm.h      |  4 ++-
 8 files changed, 118 insertions(+), 4 deletions(-)

diff --git a/drivers/accel/amdxdna/aie2_pci.c b/drivers/accel/amdxdna/aie2_pci.c
index 0d209b7b6484..f90435e1f65e 100644
--- a/drivers/accel/amdxdna/aie2_pci.c
+++ b/drivers/accel/amdxdna/aie2_pci.c
@@ -1207,6 +1207,8 @@ const struct amdxdna_dev_ops aie2_ops = {
 	.fini = aie2_fini,
 	.resume = aie2_hw_resume,
 	.suspend = aie2_hw_suspend,
+	.runtime_resume = aie2_hw_resume,
+	.runtime_suspend = aie2_hw_suspend,
 	.get_aie_info = aie2_get_info,
 	.set_aie_state = aie2_set_state,
 	.hwctx_init = aie2_hwctx_init,
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index 6a50c1499ec9..007b14be5245 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -1192,6 +1192,17 @@ static int aie4_pf_suspend(struct amdxdna_dev *xdna)
 	return 0;
 }
 
+static int aie4_pf_runtime_suspend(struct amdxdna_dev *xdna)
+{
+	int ret;
+
+	ret = aie4_vfs_alive(xdna);
+	if (ret)
+		return ret;
+
+	return aie4_pf_suspend(xdna);
+}
+
 static int aie4_pf_resume(struct amdxdna_dev *xdna)
 {
 	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
@@ -1244,6 +1255,20 @@ static int aie4_vf_suspend(struct amdxdna_dev *xdna)
 	return 0;
 }
 
+static int aie4_vf_runtime_suspend(struct amdxdna_dev *xdna)
+{
+	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
+	struct pci_dev *pdev = to_pci_dev(xdna->ddev.dev);
+
+	drm_WARN_ON(&xdna->ddev, !mutex_is_locked(&xdna->dev_lock));
+	aie4_hwctx_suspend_all(ndev, false);
+	aie4_vf_hw_stop(ndev);
+	pci_disable_device(pdev);
+
+	XDNA_DBG(xdna, "vf runtime suspend done");
+	return 0;
+}
+
 static int aie4_vf_resume(struct amdxdna_dev *xdna)
 {
 	struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
@@ -1340,6 +1365,8 @@ const struct amdxdna_dev_ops aie4_pf_ops = {
 	.sriov_configure        = aie4_sriov_configure,
 	.resume			= aie4_pf_resume,
 	.suspend		= aie4_pf_suspend,
+	.runtime_resume		= aie4_pf_resume,
+	.runtime_suspend	= aie4_pf_runtime_suspend,
 };
 
 const struct amdxdna_dev_ops aie4_vf_ops = {
@@ -1354,6 +1381,8 @@ const struct amdxdna_dev_ops aie4_vf_ops = {
 	.set_aie_state		= aie4_set_state,
 	.resume			= aie4_vf_resume,
 	.suspend		= aie4_vf_suspend,
+	.runtime_resume		= aie4_vf_resume,
+	.runtime_suspend	= aie4_vf_runtime_suspend,
 };
 
 const struct amdxdna_dev_ops aie4_classic_ops = {
@@ -1368,4 +1397,6 @@ const struct amdxdna_dev_ops aie4_classic_ops = {
 	.set_aie_state		= aie4_set_state,
 	.resume			= aie4_classic_resume,
 	.suspend		= aie4_classic_suspend,
+	.runtime_resume		= aie4_classic_resume,
+	.runtime_suspend	= aie4_classic_suspend,
 };
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index df15c63317b1..062275be7ee7 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -198,16 +198,24 @@ void aie4_free_notification(struct cert_comp *comp);
 int aie4_create_vfs(struct amdxdna_dev_hdl *ndev, int num_vfs);
 int aie4_sriov_configure(struct amdxdna_dev *xdna, int num_vfs);
 int aie4_sriov_stop(struct amdxdna_dev_hdl *ndev);
+int aie4_vfs_alive(struct amdxdna_dev *xdna);
 #else
 static inline int aie4_create_vfs(struct amdxdna_dev_hdl *ndev, int num_vfs)
 {
 	return 0;
 }
+
 #define aie4_sriov_configure NULL
+
 static inline int aie4_sriov_stop(struct amdxdna_dev_hdl *ndev)
 {
 	return 0;
 }
+
+static inline int aie4_vfs_alive(struct amdxdna_dev *xdna)
+{
+	return 0;
+}
 #endif
 
 extern const struct amdxdna_dev_ops aie4_pf_ops;
diff --git a/drivers/accel/amdxdna/aie4_sriov.c b/drivers/accel/amdxdna/aie4_sriov.c
index bfea6ff00ec0..10c8a2648e23 100644
--- a/drivers/accel/amdxdna/aie4_sriov.c
+++ b/drivers/accel/amdxdna/aie4_sriov.c
@@ -6,6 +6,7 @@
 #include <drm/amdxdna_accel.h>
 #include <drm/drm_print.h>
 #include <linux/pci.h>
+#include <linux/pm_runtime.h>
 
 #include "aie.h"
 #include "aie4_msg_priv.h"
@@ -120,6 +121,31 @@ static int aie4_link_vf(struct amdxdna_dev *xdna, struct pci_dev *pdev_vf)
 	return 0;
 }
 
+static int aie4_check_vf_alive(struct amdxdna_dev *xdna, struct pci_dev *pdev_vf)
+{
+	struct device_driver *drv = READ_ONCE(pdev_vf->dev.driver);
+
+	if (drv && drv->owner != THIS_MODULE) {
+		XDNA_WARN(xdna, "VF:%s is in passthrough", pci_name(pdev_vf));
+		return -EBUSY;
+	}
+
+	if (!pm_runtime_suspended(&pdev_vf->dev)) {
+		XDNA_WARN(xdna, "VF:%s is busy", pci_name(pdev_vf));
+		return -EBUSY;
+	}
+	return 0;
+}
+
+int aie4_vfs_alive(struct amdxdna_dev *xdna)
+{
+	if (pci_vfs_assigned(to_pci_dev(xdna->ddev.dev))) {
+		XDNA_WARN(xdna, "VF devices are being used in VMs, cannot suspend");
+		return -EBUSY;
+	}
+	return aie4_for_each_vfs(xdna, aie4_check_vf_alive);
+}
+
 static int aie4_link_vfs(struct amdxdna_dev *xdna)
 {
 	return aie4_for_each_vfs(xdna, aie4_link_vf);
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.c b/drivers/accel/amdxdna/amdxdna_pci_drv.c
index 1d0b91e73260..4933962844f9 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.c
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.c
@@ -467,18 +467,27 @@ static void amdxdna_remove(struct pci_dev *pdev)
 
 static const struct dev_pm_ops amdxdna_pm_ops = {
 	SYSTEM_SLEEP_PM_OPS(amdxdna_pm_suspend, amdxdna_pm_resume)
-	RUNTIME_PM_OPS(amdxdna_pm_suspend, amdxdna_pm_resume, NULL)
+	RUNTIME_PM_OPS(amdxdna_pm_runtime_suspend, amdxdna_pm_runtime_resume, NULL)
 };
 
 static int amdxdna_sriov_configure(struct pci_dev *pdev, int num_vfs)
 {
 	struct amdxdna_dev *xdna = pci_get_drvdata(pdev);
+	int ret;
 
 	guard(mutex)(&xdna->dev_lock);
+
+	ret = amdxdna_pm_resume_get_locked(xdna);
+	if (ret)
+		return ret;
+
 	if (xdna->dev_info->ops->sriov_configure)
-		return xdna->dev_info->ops->sriov_configure(xdna, num_vfs);
+		ret = xdna->dev_info->ops->sriov_configure(xdna, num_vfs);
+	else
+		ret = -EOPNOTSUPP;
 
-	return -ENOENT;
+	amdxdna_pm_suspend_put(xdna);
+	return ret;
 }
 
 static struct pci_driver amdxdna_pci_driver = {
diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.h b/drivers/accel/amdxdna/amdxdna_pci_drv.h
index 45170c945d71..8c735f184f00 100644
--- a/drivers/accel/amdxdna/amdxdna_pci_drv.h
+++ b/drivers/accel/amdxdna/amdxdna_pci_drv.h
@@ -57,6 +57,8 @@ struct amdxdna_dev_ops {
 	void (*debugfs_init)(struct amdxdna_dev *xdna);
 	int (*resume)(struct amdxdna_dev *xdna);
 	int (*suspend)(struct amdxdna_dev *xdna);
+	int (*runtime_resume)(struct amdxdna_dev *xdna);
+	int (*runtime_suspend)(struct amdxdna_dev *xdna);
 	int (*sriov_configure)(struct amdxdna_dev *xdna, int num_vfs);
 	int (*hwctx_init)(struct amdxdna_hwctx *hwctx);
 	void (*hwctx_fini)(struct amdxdna_hwctx *hwctx);
diff --git a/drivers/accel/amdxdna/amdxdna_pm.c b/drivers/accel/amdxdna/amdxdna_pm.c
index 9c030b7836fb..2b4ce301bb96 100644
--- a/drivers/accel/amdxdna/amdxdna_pm.c
+++ b/drivers/accel/amdxdna/amdxdna_pm.c
@@ -37,11 +37,40 @@ int amdxdna_pm_resume(struct device *dev)
 	return ret;
 }
 
+int amdxdna_pm_runtime_suspend(struct device *dev)
+{
+	struct amdxdna_dev *xdna = to_xdna_dev(dev_get_drvdata(dev));
+	int ret = -EOPNOTSUPP;
+
+	guard(mutex)(&xdna->dev_lock);
+	if (xdna->dev_info->ops->runtime_suspend)
+		ret = xdna->dev_info->ops->runtime_suspend(xdna);
+
+	XDNA_DBG(xdna, "Runtime suspend done ret %d", ret);
+	return ret;
+}
+
+int amdxdna_pm_runtime_resume(struct device *dev)
+{
+	struct amdxdna_dev *xdna = to_xdna_dev(dev_get_drvdata(dev));
+	int ret = -EOPNOTSUPP;
+
+	guard(mutex)(&xdna->dev_lock);
+	if (xdna->dev_info->ops->runtime_resume)
+		ret = xdna->dev_info->ops->runtime_resume(xdna);
+
+	XDNA_DBG(xdna, "Runtime resume done ret %d", ret);
+	return ret;
+}
+
 int amdxdna_pm_resume_get(struct amdxdna_dev *xdna)
 {
 	struct device *dev = xdna->ddev.dev;
 	int ret;
 
+	if (!pm_runtime_enabled(dev))
+		return 0;
+
 	ret = pm_runtime_resume_and_get(dev);
 	if (ret) {
 		XDNA_ERR(xdna, "Resume failed: %d", ret);
@@ -55,6 +84,10 @@ void amdxdna_pm_suspend_put(struct amdxdna_dev *xdna)
 {
 	struct device *dev = xdna->ddev.dev;
 
+	if (!pm_runtime_enabled(dev))
+		return;
+
+	pm_runtime_mark_last_busy(dev);
 	pm_runtime_put_autosuspend(dev);
 }
 
@@ -66,6 +99,7 @@ void amdxdna_pm_init(struct amdxdna_dev *xdna)
 	pm_runtime_set_autosuspend_delay(dev, AMDXDNA_AUTOSUSPEND_DELAY);
 	pm_runtime_use_autosuspend(dev);
 	pm_runtime_allow(dev);
+	pm_runtime_mark_last_busy(dev);
 	pm_runtime_put_autosuspend(dev);
 }
 
diff --git a/drivers/accel/amdxdna/amdxdna_pm.h b/drivers/accel/amdxdna/amdxdna_pm.h
index 3d26b973e0e3..26df3d50d8ea 100644
--- a/drivers/accel/amdxdna/amdxdna_pm.h
+++ b/drivers/accel/amdxdna/amdxdna_pm.h
@@ -9,7 +9,9 @@
 #include "amdxdna_pci_drv.h"
 
 int amdxdna_pm_suspend(struct device *dev);
-int amdxdna_pm_resume(struct device  *dev);
+int amdxdna_pm_resume(struct device *dev);
+int amdxdna_pm_runtime_suspend(struct device *dev);
+int amdxdna_pm_runtime_resume(struct device *dev);
 int amdxdna_pm_resume_get(struct amdxdna_dev *xdna);
 void amdxdna_pm_suspend_put(struct amdxdna_dev *xdna);
 void amdxdna_pm_init(struct amdxdna_dev *xdna);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 19/20] accel/amdxdna: Add stub hwctx_config for AIE4
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (17 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 18/20] accel/amdxdna: Implement runtime suspend and resume support David Zhang
@ 2026-09-30  3:32 ` David Zhang
  2026-09-30  3:59   ` sashiko-bot
  2026-09-30  3:32 ` [PATCH V1 20/20] accel/amdxdna: Enable AIE4 firmware logging to DRAM David Zhang
  19 siblings, 1 reply; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

XRT issues DRM_AMDXDNA_CONFIG_HWCTX during hardware context
initialization. If hwctx_config is NULL, the ioctl returns -EOPNOTSUPP,
causing userspace validation tests like GEMM to fail.

Add a stub aie4_hwctx_config() returning 0 and wire it to aie4_vf_ops
and aie4_classic_ops.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_pci.c | 8 ++++++++
 1 file changed, 8 insertions(+)

diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index 007b14be5245..62ee7dfc7bd3 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -1369,12 +1369,19 @@ const struct amdxdna_dev_ops aie4_pf_ops = {
 	.runtime_suspend	= aie4_pf_runtime_suspend,
 };
 
+static int aie4_hwctx_config(struct amdxdna_hwctx *hwctx, u32 type, u64 value,
+			     void *buf, u32 size)
+{
+	return 0;
+}
+
 const struct amdxdna_dev_ops aie4_vf_ops = {
 	.init			= aie4_vf_init,
 	.fini			= aie4_vf_fini,
 	.debugfs_init		= aie4_debugfs_init,
 	.hwctx_init		= aie4_hwctx_init,
 	.hwctx_fini		= aie4_hwctx_fini,
+	.hwctx_config		= aie4_hwctx_config,
 	.cmd_submit		= aie4_cmd_submit,
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
@@ -1391,6 +1398,7 @@ const struct amdxdna_dev_ops aie4_classic_ops = {
 	.debugfs_init		= aie4_debugfs_init,
 	.hwctx_init		= aie4_hwctx_init,
 	.hwctx_fini		= aie4_hwctx_fini,
+	.hwctx_config		= aie4_hwctx_config,
 	.cmd_submit		= aie4_cmd_submit,
 	.cmd_wait		= aie4_cmd_wait,
 	.get_aie_info		= aie4_get_info,
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* [PATCH V1 20/20] accel/amdxdna: Enable AIE4 firmware logging to DRAM
  2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
                   ` (18 preceding siblings ...)
  2026-09-30  3:32 ` [PATCH V1 19/20] accel/amdxdna: Add stub hwctx_config for AIE4 David Zhang
@ 2026-09-30  3:32 ` David Zhang
  19 siblings, 0 replies; 33+ messages in thread
From: David Zhang @ 2026-09-30  3:32 UTC (permalink / raw)
  To: quic_jhugo, karol.wachowski, max.zhen, lizhi.hou, ogabbay,
	dri-devel, linux-kernel
  Cc: David Zhang, sonal.santan, mario.limonciello

Allocate a DRAM buffer for firmware logging, and also set the log level
to less verbose. This will give the best performance numbers.

Signed-off-by: David Zhang <yidong.zhang@amd.com>
---
 drivers/accel/amdxdna/aie4_message.c  | 20 ++++++++++
 drivers/accel/amdxdna/aie4_msg_priv.h | 24 ++++++++++++
 drivers/accel/amdxdna/aie4_pci.c      | 54 +++++++++++++++++++++++++++
 drivers/accel/amdxdna/aie4_pci.h      |  5 +++
 4 files changed, 103 insertions(+)

diff --git a/drivers/accel/amdxdna/aie4_message.c b/drivers/accel/amdxdna/aie4_message.c
index 1bddcb183db6..936fa2f43d91 100644
--- a/drivers/accel/amdxdna/aie4_message.c
+++ b/drivers/accel/amdxdna/aie4_message.c
@@ -232,6 +232,26 @@ int aie4_attach_work_buffer(struct amdxdna_dev_hdl *ndev)
 	return ret;
 }
 
+int aie4_start_fw_log(struct amdxdna_dev_hdl *ndev, u32 level)
+{
+	DECLARE_AIE_MSG(aie4_msg_start_fw_log, AIE4_MSG_OP_START_FW_LOG);
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	int ret;
+
+	req.buff_addr = ndev->fw_log_buf_addr;
+	req.buff_size = ndev->fw_log_buf_size;
+	req.log_level = level;
+
+	ret = aie_send_mgmt_msg_wait(&ndev->aie, &msg);
+	if (ret)
+		XDNA_WARN(xdna, "Failed to start fw log, ret %d", ret);
+	else
+		XDNA_DBG(xdna, "Started fw log, level %u size 0x%x",
+			 level, ndev->fw_log_buf_size);
+
+	return ret;
+}
+
 int aie4_msg_set_power_mode(struct amdxdna_dev_hdl *ndev, u8 power_mode)
 {
 	DECLARE_AIE_MSG(aie4_msg_power_override, AIE4_MSG_OP_POWER_OVERRIDE);
diff --git a/drivers/accel/amdxdna/aie4_msg_priv.h b/drivers/accel/amdxdna/aie4_msg_priv.h
index 77984683a7b6..51e7f4b5ed53 100644
--- a/drivers/accel/amdxdna/aie4_msg_priv.h
+++ b/drivers/accel/amdxdna/aie4_msg_priv.h
@@ -29,6 +29,7 @@ enum aie4_msg_opcode {
 	AIE4_MSG_OP_GET_CURRENT_DPM_LEVEL            = 0x30013,
 
 	AIE4_MSG_OP_ATTACH_WORK_BUFFER               = 0x40001,
+	AIE4_MSG_OP_START_FW_LOG                     = 0x40003,
 };
 
 enum aie4_msg_status {
@@ -275,4 +276,27 @@ struct aie4_msg_attach_work_buffer_resp {
 	enum aie4_msg_status status;
 } __packed;
 
+/* Dynamic firmware log levels. */
+enum aie4_fw_log_level {
+	AIE4_FW_LOG_LEVEL_OFF,
+	AIE4_FW_LOG_LEVEL_ERR,
+	AIE4_FW_LOG_LEVEL_WRN,
+	AIE4_FW_LOG_LEVEL_INF,
+	AIE4_FW_LOG_LEVEL_DBG,
+	AIE4_FW_LOG_LEVEL_MAX,
+};
+
+#define AIE4_FW_LOG_BUF_SIZE      SZ_1M
+
+struct aie4_msg_start_fw_log_req {
+	__u64 buff_addr;
+	__u32 buff_size;
+	__u32 log_level;
+	__u32 reserved;
+} __packed;
+
+struct aie4_msg_start_fw_log_resp {
+	enum aie4_msg_status status;
+} __packed;
+
 #endif /* _AIE4_MSG_PRIV_H_ */
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index 62ee7dfc7bd3..480bd64b0020 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
@@ -404,6 +404,13 @@ static int aie4_config_fw(struct amdxdna_dev_hdl *ndev)
 	/* Best-effort tuning knob; failure is warned inside and does not fail hw start */
 	aie4_set_ctx_hysteresis(ndev, ndev->ctx_switch_hysteresis_us);
 
+	/*
+	 * Give the firmware high performance DRAM to log into with less
+	 * verbose level.
+	 */
+	if (ndev->fw_log_buf)
+		aie4_start_fw_log(ndev, AIE4_FW_LOG_LEVEL_ERR);
+
 	return 0;
 }
 
@@ -960,6 +967,45 @@ static void aie4_free_work_buffer(struct amdxdna_dev_hdl *ndev)
 	ndev->work_buf = NULL;
 }
 
+/*
+ * Firmware logging is best effort: a device that cannot spare the buffer still
+ * runs, it just does not log. Never fail hw start or probe on this.
+ */
+static void aie4_alloc_fw_log_buffer(struct amdxdna_dev_hdl *ndev)
+{
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+	u32 buf_size = AIE4_FW_LOG_BUF_SIZE;
+
+	ndev->fw_log_buf = amdxdna_alloc_msg_buffer(xdna, &buf_size,
+						    &ndev->fw_log_buf_addr);
+	if (IS_ERR(ndev->fw_log_buf)) {
+		XDNA_WARN(xdna, "Failed to alloc fw log buffer, size 0x%x",
+			  AIE4_FW_LOG_BUF_SIZE);
+		ndev->fw_log_buf = NULL;
+		return;
+	}
+
+	ndev->fw_log_buf_size = buf_size;
+	XDNA_DBG(xdna, "FW log buffer allocated: size 0x%x", buf_size);
+}
+
+/*
+ * Only called from the fini paths, after hw stop has already stopped the
+ * firmware, so the firmware cannot still be writing into the buffer. This
+ * mirrors the work buffer, which has no detach message either.
+ */
+static void aie4_free_fw_log_buffer(struct amdxdna_dev_hdl *ndev)
+{
+	struct amdxdna_dev *xdna = ndev->aie.xdna;
+
+	if (!ndev->fw_log_buf)
+		return;
+
+	amdxdna_free_msg_buffer(xdna, ndev->fw_log_buf_size, ndev->fw_log_buf,
+				ndev->fw_log_buf_addr);
+	ndev->fw_log_buf = NULL;
+}
+
 static int aie4_pf_init(struct amdxdna_dev *xdna)
 {
 	int ret;
@@ -972,6 +1018,8 @@ static int aie4_pf_init(struct amdxdna_dev *xdna)
 	if (ret)
 		return ret;
 
+	aie4_alloc_fw_log_buffer(xdna->dev_handle);
+
 	ret = aie4_pf_hw_start(xdna->dev_handle);
 	if (ret)
 		goto free_work_buf;
@@ -979,6 +1027,7 @@ static int aie4_pf_init(struct amdxdna_dev *xdna)
 	return 0;
 
 free_work_buf:
+	aie4_free_fw_log_buffer(xdna->dev_handle);
 	aie4_free_work_buffer(xdna->dev_handle);
 	return ret;
 }
@@ -1006,6 +1055,8 @@ static int aie4_classic_init(struct amdxdna_dev *xdna)
 	if (ret)
 		return ret;
 
+	aie4_alloc_fw_log_buffer(xdna->dev_handle);
+
 	ret = aie4_classic_hw_start(xdna->dev_handle);
 	if (ret)
 		goto free_work_buf;
@@ -1013,6 +1064,7 @@ static int aie4_classic_init(struct amdxdna_dev *xdna)
 	return 0;
 
 free_work_buf:
+	aie4_free_fw_log_buffer(xdna->dev_handle);
 	aie4_free_work_buffer(xdna->dev_handle);
 	return ret;
 }
@@ -1021,6 +1073,7 @@ static void aie4_pf_fini(struct amdxdna_dev *xdna)
 {
 	aie4_sriov_stop(xdna->dev_handle);
 	aie4_pf_hw_stop(xdna->dev_handle);
+	aie4_free_fw_log_buffer(xdna->dev_handle);
 	aie4_free_work_buffer(xdna->dev_handle);
 }
 
@@ -1032,6 +1085,7 @@ static void aie4_vf_fini(struct amdxdna_dev *xdna)
 static void aie4_classic_fini(struct amdxdna_dev *xdna)
 {
 	aie4_classic_hw_stop(xdna->dev_handle);
+	aie4_free_fw_log_buffer(xdna->dev_handle);
 	aie4_free_work_buffer(xdna->dev_handle);
 }
 
diff --git a/drivers/accel/amdxdna/aie4_pci.h b/drivers/accel/amdxdna/aie4_pci.h
index 062275be7ee7..18d1f03247d3 100644
--- a/drivers/accel/amdxdna/aie4_pci.h
+++ b/drivers/accel/amdxdna/aie4_pci.h
@@ -120,6 +120,10 @@ struct amdxdna_dev_hdl {
 	dma_addr_t			work_buf_addr;
 	u32				work_buf_size;
 
+	void				*fw_log_buf;
+	dma_addr_t			fw_log_buf_addr;
+	u32				fw_log_buf_size;
+
 	u8				pw_mode;
 
 	/*
@@ -143,6 +147,7 @@ int aie4_query_aie_version(struct amdxdna_dev_hdl *ndev,
 			   struct amdxdna_drm_query_aie_version *version);
 int aie4_suspend_fw(struct amdxdna_dev_hdl *ndev);
 int aie4_attach_work_buffer(struct amdxdna_dev_hdl *ndev);
+int aie4_start_fw_log(struct amdxdna_dev_hdl *ndev, u32 level);
 int aie4_query_npu_firmware_version(struct amdxdna_dev_hdl *ndev,
 				    struct amdxdna_drm_query_firmware_version *fw_version);
 int aie4_query_cert_firmware_version(struct amdxdna_dev_hdl *ndev,
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 09/20] accel/amdxdna: Add context switch hysteresis with debugfs control
  2026-09-30  3:32 ` [PATCH V1 09/20] accel/amdxdna: Add context switch hysteresis with debugfs control David Zhang
@ 2026-09-30  3:52   ` sashiko-bot
  0 siblings, 0 replies; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  3:52 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] amdxdna: ignoring timeout in aie4_config_fw() leaves device wedged
--

commit 22f93820c29f525444ddc497eafd8728678db445
Author: David Zhang <yidong.zhang@amd.com>
accel/amdxdna: Add context switch hysteresis with debugfs control

This commit adds aie4_set_ctx_hysteresis() to configure the AIE4 context
switch hysteresis timeout via the SET_RUNTIME_CONFIG message, and it exposes a
debugfs node to change the timeout at runtime.

> diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
> index 95e682a3a4b75..60348ec5bc53b 100644
> --- a/drivers/accel/amdxdna/aie4_pci.c
> +++ b/drivers/accel/amdxdna/aie4_pci.c
[ ... ]
> @@ -325,6 +327,20 @@ int aie4_restore_power_mode(struct amdxdna_dev_hdl *ndev)
>  	return aie4_msg_set_power_mode(ndev, ndev->pw_mode);
>  }
>  
> +static int aie4_config_fw(struct amdxdna_dev_hdl *ndev)
> +{
> +	int ret;
> +
> +	ret = aie4_attach_work_buffer(ndev);
> +	if (ret)
> +		return ret;
> +
> +	/* Best-effort tuning knob; failure is warned inside and does not fail hw start */
> +	aie4_set_ctx_hysteresis(ndev, ndev->ctx_switch_hysteresis_us);

[Severity: Medium]
Does ignoring the return value of aie4_set_ctx_hysteresis() leave the device
in a wedged state if a timeout occurs?

If the firmware hangs or times out while processing the context hysteresis
configuration message during initialization or resume, aie_send_mgmt_msg_wait()
returns -ETIME and destroys the management channel.

Because aie4_config_fw() explicitly ignores all errors to treat the knob as
"best-effort", initialization will falsely succeed even if the management
channel has been destroyed. This leaves the device silently unusable. Is there
a way to distinguish between non-fatal errors and fatal timeouts to prevent
the probe from succeeding with a destroyed management channel?

> +
> +	return 0;
> +}
> +
>  static int aie4_pf_hw_start(struct amdxdna_dev_hdl *ndev)
>  {
>  	int ret;

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=9

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 03/20] accel/amdxdna: Add CERT firmware version support
  2026-09-30  3:32 ` [PATCH V1 03/20] accel/amdxdna: Add CERT firmware version support David Zhang
@ 2026-09-30  3:53   ` sashiko-bot
  0 siblings, 0 replies; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  3:53 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] amdxdna: redundant entry in npu3_cert_feature_table
--

commit 3d4c843c0c43dded98465c96cc3cb5ac8f46a7fb
Author: David Zhang <yidong.zhang@amd.com>

accel/amdxdna: Add CERT firmware version support

This commit adds support for querying and validating CERT firmware versions
on AIE4 platforms. It introduces mailbox opcodes, unifies firmware version
storage, and provides host queue protocol validation against driver
capabilities.

> diff --git a/drivers/accel/amdxdna/npu3_regs.c b/drivers/accel/amdxdna/npu3_regs.c
> index 8d287ef32fff7..31208c42ad5ff 100644
> --- a/drivers/accel/amdxdna/npu3_regs.c
> +++ b/drivers/accel/amdxdna/npu3_regs.c
> @@ -42,6 +42,12 @@ static const struct amdxdna_fw_feature_tbl npu3_fw_feature_table[] = {
>  	{ 0 }
>  };
>  
> +static const struct amdxdna_fw_feature_tbl npu3_cert_feature_table[] = {
> +	{ .major = 1, .min_minor = 0 },
> +	{ .features = BIT_U64(AIE4_HSA_COMMAND), .major = 1, .min_minor = 0 },

[Severity: Low]
Is this first entry without any features intentional?

When aie_check_protocol_impl() iterates through this table, both entries 
match major = 1 and min_minor >= 0, so it will process both. The first 
entry will perform a no-op bitwise OR on the feature mask 
(aie->feature_mask |= 0) before the second entry applies the 
AIE4_HSA_COMMAND feature.

This isn't a functional bug, but seems like a redundant configuration that 
could be simplified by removing the first entry.

> +	{ 0 }
> +};
> +
>  static const struct amdxdna_dev_priv npu3_dev_priv = {
>  	.npufw_path             = "npu.sbin",
>  	.certfw_path            = "cert.sbin",

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=3

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 07/20] accel/amdxdna: Add get and set power_mode for AIE4
  2026-09-30  3:32 ` [PATCH V1 07/20] accel/amdxdna: Add get and set power_mode for AIE4 David Zhang
@ 2026-09-30  3:56   ` sashiko-bot
  0 siblings, 0 replies; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  3:56 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] amdxdna: partition leak in aie4_vf_hw_start() error path
--

commit bc6bbbe04622b7e463790d57007bb4c50fb33683
Author: David Zhang <yidong.zhang@amd.com>

    accel/amdxdna: Add get and set power_mode for AIE4

This commit adds power mode support for AIE4 devices via DRM_AMDXDNA_GET_INFO
and DRM_AMDXDNA_SET_STATE. It implements aie4_set_power_mode() and restores
the cached user power mode override whenever the hardware starts.

> diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
> index aea3edd51b4f5..d11bdbf168816 100644
> --- a/drivers/accel/amdxdna/aie4_pci.c
> +++ b/drivers/accel/amdxdna/aie4_pci.c
[ ... ]
> @@ -346,6 +373,10 @@ static int aie4_vf_hw_start(struct amdxdna_dev_hdl *ndev)
>  	if (ret)
>  		goto mailbox_fini;
>  
> +	ret = aie4_restore_power_mode(ndev);
> +	if (ret)
> +		goto mailbox_fini;
> +

[Severity: Medium]
Does this error path leak the firmware hardware partition?

If aie4_restore_power_mode() fails here, we jump directly to the mailbox_fini
label. This bypasses aie4_partition_fini(), leaving the partition that was
successfully initialized just prior to this call permanently leaked.

>  	return 0;
>  
>  mailbox_fini:
>  	aie4_mailbox_fini(ndev);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=7

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 17/20] accel/amdxdna: Link SR-IOV VFs for power management sequencing
  2026-09-30  3:32 ` [PATCH V1 17/20] accel/amdxdna: Link SR-IOV VFs for power management sequencing David Zhang
@ 2026-09-30  3:59   ` sashiko-bot
  0 siblings, 0 replies; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  3:59 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] pm: device link between PF and VF permanently lost on first VF unbind
--

commit 2785e7a8d2421554408eee0ed2983e10f0090091
Author: David Zhang <yidong.zhang@amd.com>

accel/amdxdna: Link SR-IOV VFs for power management sequencing

This commit adds PM device links between the Physical Function (PF) and
Virtual Functions (VFs) when SR-IOV is enabled. This ensures proper PM
sequencing so that VFs are suspended before the PF, and the PF is
resumed before VFs.

> diff --git a/drivers/accel/amdxdna/aie4_sriov.c b/drivers/accel/amdxdna/aie4_sriov.c
> index 0eea28f626763..bfea6ff00ec04 100644
> --- a/drivers/accel/amdxdna/aie4_sriov.c
> +++ b/drivers/accel/amdxdna/aie4_sriov.c

[ ... ]

> +static int aie4_link_vf(struct amdxdna_dev *xdna, struct pci_dev *pdev_vf)
> +{
> +	struct pci_dev *pdev_pf = to_pci_dev(xdna->ddev.dev);
> +	struct device_link *link;
> +
> +	link = device_link_add(&pdev_vf->dev,   /* consumer = VF */
> +			       &pdev_pf->dev,   /* supplier = PF */
> +			       DL_FLAG_PM_RUNTIME | DL_FLAG_AUTOREMOVE_CONSUMER);

[Severity: High]
Could this device link be permanently lost on the first VF driver unbind?

The DL_FLAG_AUTOREMOVE_CONSUMER flag automatically deletes the device link
when the consumer (VF) driver unbinds. Since the link is only created by the
PF driver during aie4_sriov_start() when SR-IOV is first enabled, it appears
it won't be recreated if a driver is subsequently re-bound to the VF.

If an administrator unbinds and rebinds a VF driver, would the missing
PM device link allow the PM core to runtime-suspend the PF while the
newly-bound VF is actively using it? Removing DL_FLAG_AUTOREMOVE_CONSUMER
might ensure the link remains tied to the device lifecycle instead.

> +	if (!link) {
> +		XDNA_ERR(xdna, "Failed to link VF %s", pci_name(pdev_vf));
> +		return -EINVAL;
> +	}
> +
> +	XDNA_DBG(xdna, "Linked VF %s", pci_name(pdev_vf));
> +	return 0;
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=17

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 19/20] accel/amdxdna: Add stub hwctx_config for AIE4
  2026-09-30  3:32 ` [PATCH V1 19/20] accel/amdxdna: Add stub hwctx_config for AIE4 David Zhang
@ 2026-09-30  3:59   ` sashiko-bot
  0 siblings, 0 replies; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  3:59 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] amdxdna: aie4_hwctx_config unconditionally returns success for unsupported types
--

commit 2fb62e44c1bbca14b0a25248c16851b8cb0c7aa6
Author: David Zhang <yidong.zhang@amd.com>

accel/amdxdna: Add stub hwctx_config for AIE4

XRT issues DRM_AMDXDNA_CONFIG_HWCTX during hardware context
initialization. If hwctx_config is NULL, the ioctl returns -EOPNOTSUPP,
causing userspace validation tests like GEMM to fail. Add a stub 
aie4_hwctx_config() returning 0 and wire it to aie4_vf_ops and
aie4_classic_ops.

> diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
> index 007b14be5245f..62ee7dfc7bd37 100644
> --- a/drivers/accel/amdxdna/aie4_pci.c
> +++ b/drivers/accel/amdxdna/aie4_pci.c
> @@ -1369,12 +1369,19 @@ const struct amdxdna_dev_ops aie4_pf_ops = {
>  	.runtime_suspend	= aie4_pf_runtime_suspend,
>  };
>  
> +static int aie4_hwctx_config(struct amdxdna_hwctx *hwctx, u32 type, u64 value,
> +			     void *buf, u32 size)
> +{
> +	return 0;
> +}

[Severity: Medium]
Does returning 0 unconditionally here falsely report success for unsupported
hardware context configuration commands?

If a userspace library like XRT calls the AMDXDNA_CONFIG_HWCTX ioctl with an
unsupported type (such as DRM_AMDXDNA_HWCTX_ASSIGN_DBG_BUF), returning 0
could cause userspace to assume the operation was successful. Could this
cause userspace to read uninitialized or garbage data from the buffer since
it relies on standard kernel error codes like -EOPNOTSUPP or -EINVAL to
detect hardware capabilities?

> +
>  const struct amdxdna_dev_ops aie4_vf_ops = {
>  	.init			= aie4_vf_init,
>  	.fini			= aie4_vf_fini,
>  	.debugfs_init		= aie4_debugfs_init,
>  	.hwctx_init		= aie4_hwctx_init,
>  	.hwctx_fini		= aie4_hwctx_fini,
> +	.hwctx_config		= aie4_hwctx_config,
>  	.cmd_submit		= aie4_cmd_submit,
>  	.cmd_wait		= aie4_cmd_wait,
>  	.get_aie_info		= aie4_get_info,

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=19

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 13/20] accel/amdxdna: Prepare for AIE4 command submission
  2026-09-30  3:32 ` [PATCH V1 13/20] accel/amdxdna: Prepare for AIE4 command submission David Zhang
@ 2026-09-30  4:00   ` sashiko-bot
  0 siblings, 0 replies; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  4:00 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Race in get_read_index() allows last_read_index to regress
--

commit b18a81613bb4ba1a66541441dd19a7be5abf5cef
Author: David Zhang <yidong.zhang@amd.com>

accel/amdxdna: Prepare for AIE4 command submission

This commit prepares the data structures and completion wait helpers required
for AIE4 command submission. It defines struct amdxdna_cmd_start_dpu, extends
union amdxdna_job_priv, and updates completion wait helpers to handle read
tearing and asynchronous device disconnects.

> diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
> index 3927c9fef05f9..59d37bd5c5a08 100644
> --- a/drivers/accel/amdxdna/aie4_ctx.c
> +++ b/drivers/accel/amdxdna/aie4_ctx.c

[ ... ]

> @@ -423,34 +423,92 @@ static inline bool valid_queue_index(u64 read, u64 write, u32 capacity)
>  
>  static u64 get_read_index(struct amdxdna_hwctx *hwctx)
>  {

[ ... ]

>  	if (!valid_queue_index(ri, wi, CTX_MAX_CMDS)) {
> -		XDNA_WARN(xdna, "Invalid index, ri %llu, wi %llu", ri, wi);
> -		usleep_range(100, 200);
> -		ri = READ_ONCE(*hwctx->priv->umq_read_index);
> +		ri = READ_ONCE(*priv->umq_read_index);
> +		/* Order the read_index sample before the write_index sample. */
> +		smp_rmb();
> +		wi = READ_ONCE(priv->write_index);
>  		if (!valid_queue_index(ri, wi, CTX_MAX_CMDS)) {
> -			XDNA_ERR(xdna, "Invalid index after retry, ri %llu, wi %llu", ri, wi);
> -			ri = 0;
> +			/*
> +			 * Still invalid (torn 64-bit read, or a transient
> +			 * accounting skew). Return the last valid read_index
> +			 * instead of 0: read_index only advances, so the cached
> +			 * value is a safe lower bound -- it never reports a
> +			 * command complete that isn't, and never regresses the
> +			 * worker into falsely timing out a finished job.
> +			 */
> +			XDNA_DBG(xdna, "Invalid index, ri %llu, wi %llu", ri, wi);
> +			return READ_ONCE(priv->last_read_index);
>  		}
>  	}
>  
> +	WRITE_ONCE(priv->last_read_index, ri);

[Severity: High]
Can this lockless blind write allow an older read index to overwrite a newer
one when executed concurrently by multiple waiters?

If Thread A reads `ri = 10` from hardware and Thread B reads `ri = 12`, Thread B
could execute `WRITE_ONCE(priv->last_read_index, 12)` before Thread A executes
`WRITE_ONCE(priv->last_read_index, 10)`, causing last_read_index to regress.

This breaks the monotonicity required to safely prevent false timeouts.

>  	return ri;
>  }

[ ... ]

> @@ -466,11 +524,14 @@ int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout)
>  		wait_jifs = msecs_to_jiffies(timeout);
>  
>  	ret = wait_event_interruptible_timeout(cert_comp->waitq,
> -					       (check_cmd_done(hwctx, seq)),
> +					       check_cmd_done(hwctx, seq, cert_comp),
>  					       wait_jifs);
>  
>  	if (!ret)
>  		ret = -ETIME;
> +	else if (ret > 0 && get_read_index(hwctx) <= seq)
> +		/* Woke on disconnect/reset, not on real completion. */
> +		ret = -EAGAIN;

[Severity: High]
Could this disconnect-detection logic incorrectly trigger when get_read_index()
returns the corrupted, backward-jumped last_read_index during a torn read?

If last_read_index regresses due to the race condition above, a subsequent call
to get_read_index() experiencing a torn read would fall back to this stale
value. 

This check would then falsely evaluate to true, causing successfully completed
jobs to falsely return -EAGAIN to userspace.

>  
>  	aie4_put_cert_comp(cert_comp);
>

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=13

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 12/20] accel/amdxdna: Implement AIE4 kernel queue lifecycle and memory layout
  2026-09-30  3:32 ` [PATCH V1 12/20] accel/amdxdna: Implement AIE4 kernel queue lifecycle and memory layout David Zhang
@ 2026-09-30  4:00   ` sashiko-bot
  0 siblings, 0 replies; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  4:00 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] amdxdna: System-wide DoS via dev_lock when sleeping with SRCU read lock held
--

commit 132c3fc1d7de0e87a5c560f0114f2517ab24e40c
Author: David Zhang <yidong.zhang@amd.com>

accel/amdxdna: Implement AIE4 kernel queue lifecycle and memory layout

This commit initializes the kernel-mode submission required buffers, workqueue,
and hardware contexts for AIE4. It updates the queue definition, adds a job
workqueue for pending and running jobs, initializes direct and indirect packets,
and implements hardware context lifecycle operations.

> diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
> index 5a2fc19bad206..3927c9fef05f9 100644
> --- a/drivers/accel/amdxdna/aie4_ctx.c
> +++ b/drivers/accel/amdxdna/aie4_ctx.c

[ ... ]

> @@ -301,10 +456,12 @@ static inline bool check_cmd_done(struct amdxdna_hwctx *hwctx, u64 seq)
>  int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout)
>  {
>  	unsigned long wait_jifs = MAX_SCHEDULE_TIMEOUT;
> -	struct amdxdna_hwctx_priv *priv = hwctx->priv;
> -	struct cert_comp *cert_comp = priv->cert_comp;
> +	struct cert_comp *cert_comp = aie4_get_cert_comp(hwctx);
>  	long ret;
>  
> +	if (!cert_comp)
> +		return -EAGAIN;
> +
>  	if (timeout)
>  		wait_jifs = msecs_to_jiffies(timeout);
>  
>  	ret = wait_event_interruptible_timeout(cert_comp->waitq,
>  					       (check_cmd_done(hwctx, seq)),
>  					       wait_jifs);

[Severity: High]
Could this indefinite sleep cause a system-wide denial of service by deadlocking
the device?

If a user calls the wait ioctl with an infinite timeout, the caller acquires
srcu_read_lock(&client->hwctx_srcu) and sleeps indefinitely in
aie4_cmd_wait() waiting for hardware completion.

Concurrently, if another thread attempts to destroy the same hardware context,
it will acquire the global xdna->dev_lock and call amdxdna_hwctx_destroy_rcu().
This function blocks on synchronize_srcu(&client->hwctx_srcu), waiting for the
wait ioctl to drop the SRCU read lock.

However, the hardware destruction that would wake up the wait queue is called 
after synchronize_srcu():

amdxdna_hwctx_destroy_rcu() {
    ...
    synchronize_srcu(&client->hwctx_srcu);
    
    /* At this point, user is not able to submit new commands */
    xdna->dev_info->ops->hwctx_fini(hwctx);
}

This creates a circular dependency where aie4_cmd_wait() will never wake up
because hwctx_fini() cannot run, leaving xdna->dev_lock held forever. Could we
handle the destruction order differently or abort the wait before calling
synchronize_srcu()?

>  	if (!ret)
>  		ret = -ETIME;
>  
> +	aie4_put_cert_comp(cert_comp);
> +
>  	return ret <= 0 ? ret : 0;
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=12

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 18/20] accel/amdxdna: Implement runtime suspend and resume support
  2026-09-30  3:32 ` [PATCH V1 18/20] accel/amdxdna: Implement runtime suspend and resume support David Zhang
@ 2026-09-30  4:00   ` sashiko-bot
  0 siblings, 0 replies; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  4:00 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] pm: dev_pm_ops and runtime callbacks lack pm_ptr() wrappers
--

commit 43e22dab09d226fdbeb4b68267732652cfce0172
Author: David Zhang <yidong.zhang@amd.com>

accel/amdxdna: Implement runtime suspend and resume support

Add runtime suspend/resume for AIE4 driver. This updates amdxdna_pm.c
to implement amdxdna_pm_runtime_suspend() and amdxdna_pm_runtime_resume(),
and acquires an RPM reference across amdxdna_sriov_configure().

> diff --git a/drivers/accel/amdxdna/amdxdna_pci_drv.c b/drivers/accel/amdxdna/amdxdna_pci_drv.c
> index 1d0b91e732602..4933962844f94 100644
> --- a/drivers/accel/amdxdna/amdxdna_pci_drv.c
> +++ b/drivers/accel/amdxdna/amdxdna_pci_drv.c
> @@ -467,18 +467,27 @@ static void amdxdna_remove(struct pci_dev *pdev)
>  
>  static const struct dev_pm_ops amdxdna_pm_ops = {
>  	SYSTEM_SLEEP_PM_OPS(amdxdna_pm_suspend, amdxdna_pm_resume)
> -	RUNTIME_PM_OPS(amdxdna_pm_suspend, amdxdna_pm_resume, NULL)
> +	RUNTIME_PM_OPS(amdxdna_pm_runtime_suspend, amdxdna_pm_runtime_resume, NULL)
>  };

[Severity: Low]
Does this code introduce a regression by preventing the dev_pm_ops structure
and its callbacks from being optimized out?

When building with CONFIG_PM=n, assigning the dev_pm_ops structure directly
without the pm_ptr() wrapper in the pci_driver definition:

static struct pci_driver amdxdna_pci_driver = {
	...
	.driver.pm = &amdxdna_pm_ops,
};

means the amdxdna_pm_ops structure and the newly added callbacks
(amdxdna_pm_runtime_suspend and amdxdna_pm_runtime_resume referenced via
RUNTIME_PM_OPS) will be linked into the kernel as dead code.

Should this assignment be updated to use pm_ptr(&amdxdna_pm_ops) to ensure
the dead code is eliminated when PM support is disabled?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=18

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 14/20] accel/amdxdna: Implement AIE4 command packet building and submission
  2026-09-30  3:32 ` [PATCH V1 14/20] accel/amdxdna: Implement AIE4 command packet building and submission David Zhang
@ 2026-09-30  4:04   ` sashiko-bot
  2026-10-05 21:52     ` Zhang, Yidong (David)
  0 siblings, 1 reply; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  4:04 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
- [Critical] accel/amdxdna: TOCTOU in submit_job_cmds allows arbitrary DMA to freed memory (UAF)
- [High] accel/amdxdna: Unhandled mid-chain submission failure permanently hangs dma_fence
- [High] accel/amdxdna: dtrace_buffer address truncated to 48 bits, corrupting memory
- [High] accel/amdxdna: dma_fence uses unreferenced struct device pointer causing UAF
- [High] accel/amdxdna: Submitter holding SRCU lock waits forever for reconnect, deadlocking teardown
- [High] accel/amdxdna: Missing READ_ONCE allows double-fetch of user-shared control fields
--

commit 3eb430c5785965c7805a7d4c133aa9e24f11869b
Author: David Zhang <yidong.zhang@amd.com>

accel/amdxdna: Implement AIE4 command packet building and submission

Implement kernel-mode command submission and hardware queue packet
assembly for AIE4.

> diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
> --- a/drivers/accel/amdxdna/aie4_ctx.c
> +++ b/drivers/accel/amdxdna/aie4_ctx.c

[ ... ]

> @@ -622,7 +622,10 @@ static int wait_till_connected_hsa_not_full(struct amdxdna_hwctx *hwctx,
>  			if (!ret)
>  				hsa_not_full = true;
>  		}
>  		ret = wait_event_freezable(priv->job_list_wq,
>  					   aie4_hwctx_connected(hwctx) ||
>  					   (!wait_through_reset &&
>  					    aie4_hwctx_has_reset(hwctx)));

[Severity: High]
Can this block indefinitely and cause an SRCU deadlock?

When wait_through_reset is true for the first command, the wait condition
completely ignores the reset or error state. Since the submitter thread holds
the hwctx_srcu read lock, if a device unplug occurs and the teardown thread
calls synchronize_srcu(), the teardown will wait forever for this lock to be
released, while this code waits for a reconnect that will never happen.

>  		mutex_lock(&priv->io_lock);

[ ... ]

> @@ -696,7 +699,8 @@ static int fill_indirect_pkt(struct amdxdna_hwctx_priv *priv, u64 slot_idx,
>  	for (i = 0; i < entries; i++, dpu++, hipe++) {
>  		struct host_indirect_packet_data *hipd;
>  		u64 indirect_pkt_dev_addr;
>  		u32 uci = dpu->uc_index;

[Severity: High]
Can the compiler double-fetch dpu->uc_index from shared memory here?

Because READ_ONCE() is missing, a concurrent userspace thread could modify
the value immediately after the bounds check against
HSA_MAX_LEVEL1_INDIRECT_ENTRIES, allowing an out-of-bounds array index into
priv->umq_indirect_pkts[idx] later in this function.

>  		u32 idx;
>  
>  		/*

[ ... ]

> @@ -715,7 +718,8 @@ static int fill_indirect_pkt(struct amdxdna_hwctx_priv *priv, u64 slot_idx,
>  			upper_32_bits(dpu->instruction_buffer);
>  		hipd->payload.dtrace_buf_host_addr_low =
>  			lower_32_bits(dpu->dtrace_buffer);
>  		hipd->payload.dtrace_buf_host_addr_high =
>  			lower_16_bits(upper_32_bits(dpu->dtrace_buffer));

[Severity: High]
Does this truncate the 64-bit user address to 48 bits?

By using lower_16_bits(upper_32_bits()), bits 48-63 of the virtual address
are silently dropped. On SVA/IOMMU platforms with 5-level paging, user
virtual addresses can reach 57 bits, which might cause the hardware to write
trace data to unintended memory locations.

>  	}
>  	pkt->pkt_header.common_header.distribute = 1;

[ ... ]

> @@ -730,7 +732,8 @@ static void fill_direct_pkt(struct amdxdna_hwctx_priv *priv, u64 slot_idx,
>  	ebuf->dpu_control_code_host_addr_high = upper_32_bits(dpu->instruction_buffer);
>  	ebuf->dtrace_buf_host_addr_low = lower_32_bits(dpu->dtrace_buffer);
>  	ebuf->dtrace_buf_host_addr_high = lower_16_bits(upper_32_bits(dpu->dtrace_buffer));

[Severity: High]
Similar to fill_indirect_pkt(), can this truncate the user address to 48
bits, potentially corrupting memory on systems with 5-level paging?

>  	pkt->pkt_header.common_header.distribute = 0;

[ ... ]

> @@ -772,7 +774,8 @@ static int submit_one_cmd(struct amdxdna_hwctx *hwctx,
>  	 * against the actual payload size before dereferencing chained+1 DPU
>  	 * entries, so a bogus count cannot drive an out-of-bounds read.
>  	 */
>  	chained = dpu->chained;

[Severity: High]
Could the lack of READ_ONCE() here allow a double-fetch of dpu->chained?

If userspace modifies this value after the bounds check, it could bypass the
validation.

>  	if (chained >= HSA_MAX_LEVEL1_INDIRECT_ENTRIES) {
>  		XDNA_ERR(xdna, "Invalid DPU data");

[ ... ]

> @@ -1012,7 +1014,8 @@ static int submit_job_cmds(struct amdxdna_hwctx *hwctx,
>  		XDNA_ERR(xdna, "Invalid cmd payload for chained cmd");
>  		return -EINVAL;
>  	}
>  	ccnt = payload->command_count;

[Severity: High]
Since payload points to a GEM object mapped into user memory, can the compiler
optimize this and re-fetch payload->command_count after the bounds validation?

Without READ_ONCE(), userspace could potentially modify the count concurrently
to bypass the checks and trigger out-of-bounds reads.

>  	/*
>  	 * A chain (runlist) must fit within the queue. CERT advances the host-visible

[ ... ]

> @@ -1032,15 +1034,22 @@ static int submit_job_cmds(struct amdxdna_hwctx *hwctx,
>  	for (i = 0; i < ccnt; i++) {
>  		u32 boh = (u32)(payload->data[i]);
>  		struct amdxdna_gem_obj *abo;
>  
>  		abo = amdxdna_gem_get_obj(hwctx->client, boh, AMDXDNA_BO_SHARE);
>  		if (!abo) {
>  			XDNA_ERR(xdna, "Failed to find cmd BO %u", boh);
>  			ret = -ENOENT;
>  			break;
>  		}
>  
>  		/*
>  		 * submit_one_cmd() blocks in wait_till_connected_hsa_not_full() until the
>  		 * ctx is connected and a slot is free, so a concurrent suspend/disconnect
>  		 * is waited out inline rather than returned here. The first sub-command
>  		 * (i == 0, nothing published yet) waits through a TDR reset and runs on
>  		 * the recreated ctx; a later sub-command returns -ECONNRESET if a reset
>  		 * landed while waiting for a slot, so the published prefix is not split
>  		 * across the reset. It also returns -ERESTARTSYS on a signal, or a
>  		 * validation error. Break on any; a published prefix is then reaped by
>  		 * the job worker's reset drain (see below).
>  		 */
>  		ret = submit_one_cmd(hwctx, abo, i + 1 == ccnt, i == 0, &job->seq);
>  		amdxdna_gem_put_obj(abo);

[Severity: Critical]
Does this allow arbitrary DMA to freed memory?

The driver reads the BO handle (boh) directly from the user-shared GEM object.
Userspace can modify this handle concurrently after the DRM scheduler
validates job->bos.

Because this substituted BO is not part of job->bos, it lacks a tracking
fence. By immediately dropping the kernel reference via amdxdna_gem_put_obj(),
userspace can close the handle and free the memory while the hardware
asynchronously executes DMA operations on the freed physical memory.

>  		if (ret)
>  			break;

[ ... ]

> @@ -1054,6 +1056,7 @@ static int submit_job_cmds(struct amdxdna_hwctx *hwctx,
>  	 * published prefix and reports the partial chain as failed (ABORT). Only when
>  	 * nothing was published (i == 0) is the error returned to the caller.
>  	 */
>  	if (i > 0)
>  		return 0;

[Severity: High]
If submit_one_cmd() fails for a sub-command mid-chain (e.g., due to an
invalid payload) after the first command is published, this masks the failure
and returns success.

Since the final command is never published, CHAIN_FLG_LAST_CMD is never sent
to the hardware. Will this cause the hardware to wait indefinitely, preventing
job_done() from ever signaling the dma_fence?

>  
>  	return ret;
>  }

[ ... ]

> @@ -1182,7 +1184,8 @@ int aie4_cmd_submit(struct amdxdna_hwctx *hwctx, struct amdxdna_sched_job *job,
>  
>  	job->out_fence = dma_fence_get(job->fence);
>  	for (i = 0; i < job->bo_cnt; i++)
>  		dma_resv_add_fence(job->bos[i]->resv, job->out_fence, DMA_RESV_USAGE_WRITE);

[Severity: High]
Following up on the mid-chain failure issue in submit_job_cmds(), since the
fence is attached to shared BOs here, if it never signals due to an abandoned
hardware queue, won't any future access to those BOs by other processes hang
forever, creating a denial of service?

>  
>  	up_read(&xdna->notifier_lock);

> diff --git a/drivers/accel/amdxdna/amdxdna_ctx.c b/drivers/accel/amdxdna/amdxdna_ctx.c
> --- a/drivers/accel/amdxdna/amdxdna_ctx.c
> +++ b/drivers/accel/amdxdna/amdxdna_ctx.c

[ ... ]

> @@ -47,7 +49,8 @@ static const char *amdxdna_fence_get_timeline_name(struct dma_fence *fence)
>  	 * another process). The device outlives any individual context, so
>  	 * dev_name() is safe to call at any point during the fence's lifetime.
>  	 */
>  	return dev_name(xdna_fence->dev);

[Severity: High]
Can this trigger a use-after-free if the sync_file outlives the DRM device?

Userspace querying the timeline after the device has been torn down or
hot-unplugged would access the freed struct device pointer here.

>  }
>  

[ ... ]

> @@ -62,7 +65,8 @@ static struct dma_fence *amdxdna_fence_create(struct amdxdna_hwctx *hwctx)
>  	if (!fence)
>  		return NULL;
>  
>  	fence->dev = hwctx->client->xdna->ddev.dev;

[Severity: High]
Does this store the struct device pointer without taking a reference via
get_device()?

This allows the use-after-free mentioned above when the fence outlives the
device.

>  	spin_lock_init(&fence->lock);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=14

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 16/20] accel/amdxdna: Implement AIE4 suspend and resume
  2026-09-30  3:32 ` [PATCH V1 16/20] accel/amdxdna: Implement AIE4 suspend and resume David Zhang
@ 2026-09-30  4:07   ` sashiko-bot
  2026-10-05 21:49     ` Zhang, Yidong (David)
  0 siblings, 1 reply; 33+ messages in thread
From: sashiko-bot @ 2026-09-30  4:07 UTC (permalink / raw)
  To: David Zhang; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] amdxdna: submitter thread hangs forever on failed device resume
--

commit 3c6d084a73c19b1b3a172be33a5eb12a9d7f8889
Author: David Zhang <yidong.zhang@amd.com>

    accel/amdxdna: Implement AIE4 suspend and resume
    
    Implement suspend and resume callbacks for AIE4 Physical Function (PF),
    Virtual Function (VF), and Classic device types.

> diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
> index f180983a692dc..6a50c1499ec93 100644
> --- a/drivers/accel/amdxdna/aie4_pci.c
> +++ b/drivers/accel/amdxdna/aie4_pci.c

[ ... ]

> @@ -1096,11 +1096,250 @@ static void aie4_debugfs_init(struct amdxdna_dev *xdna)

[ ... ]

> +static int aie4_vf_resume(struct amdxdna_dev *xdna)
> +{

[ ... ]

> +	ret = aie4_hwctx_resume_all(ndev);
> +	if (ret) {
> +		XDNA_ERR(xdna, "hwctx_resume failed %d", ret);
> +		goto hw_clear;
> +	}
> +
> +	XDNA_DBG(xdna, "vf resume done");
> +	return 0;
> +
> +hw_clear:
> +	aie4_hwctx_suspend_all(ndev, true);
> +	aie4_vf_hw_stop(ndev);
> +pci_disable:
> +	pci_disable_device(pdev);

[Severity: High]
Could this error path leave submitter threads permanently hung in an
uninterruptible wait?

When aie4_vf_resume() or aie4_classic_resume() fails, they call
aie4_hwctx_suspend_all(ndev, true) which destroys the context with
AIE4_HWCTX_ERROR (setting has_reset = true). 

However, a submitter thread waiting in wait_till_connected_hsa_not_full() in
drivers/accel/amdxdna/aie4_ctx.c with wait_through_reset = true explicitly
ignores has_reset:

    ret = wait_event_freezable(priv->job_list_wq,
                               aie4_hwctx_connected(hwctx) ||
                               (!wait_through_reset &&
                                aie4_hwctx_has_reset(hwctx)));

Since the hardware resume failed, the context will never be recreated and
connected, causing it to block forever. This holds the srcu read lock and
prevents device removal (causing a deadlock in synchronize_srcu()).

> +	return ret;
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=16

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 16/20] accel/amdxdna: Implement AIE4 suspend and resume
  2026-09-30  4:07   ` sashiko-bot
@ 2026-10-05 21:49     ` Zhang, Yidong (David)
  0 siblings, 0 replies; 33+ messages in thread
From: Zhang, Yidong (David) @ 2026-10-05 21:49 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: dri-devel


On 9/29/2026 9:07 PM, sashiko-bot@kernel.org wrote:
> [You don't often get email from sashiko-bot@kernel.org. Learn why this is important at https://aka.ms/LearnAboutSenderIdentification ]
>
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] amdxdna: submitter thread hangs forever on failed device resume
> --
>
> commit 3c6d084a73c19b1b3a172be33a5eb12a9d7f8889
> Author: David Zhang <yidong.zhang@amd.com>
>
>      accel/amdxdna: Implement AIE4 suspend and resume
>
>      Implement suspend and resume callbacks for AIE4 Physical Function (PF),
>      Virtual Function (VF), and Classic device types.
>
>> diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
>> index f180983a692dc..6a50c1499ec93 100644
>> --- a/drivers/accel/amdxdna/aie4_pci.c
>> +++ b/drivers/accel/amdxdna/aie4_pci.c
> [ ... ]
>
>> @@ -1096,11 +1096,250 @@ static void aie4_debugfs_init(struct amdxdna_dev *xdna)
> [ ... ]
>
>> +static int aie4_vf_resume(struct amdxdna_dev *xdna)
>> +{
> [ ... ]
>
>> +     ret = aie4_hwctx_resume_all(ndev);
>> +     if (ret) {
>> +             XDNA_ERR(xdna, "hwctx_resume failed %d", ret);
>> +             goto hw_clear;
>> +     }
>> +
>> +     XDNA_DBG(xdna, "vf resume done");
>> +     return 0;
>> +
>> +hw_clear:
>> +     aie4_hwctx_suspend_all(ndev, true);
>> +     aie4_vf_hw_stop(ndev);
>> +pci_disable:
>> +     pci_disable_device(pdev);
> [Severity: High]
> Could this error path leave submitter threads permanently hung in an
> uninterruptible wait?
>
> When aie4_vf_resume() or aie4_classic_resume() fails, they call
> aie4_hwctx_suspend_all(ndev, true) which destroys the context with
> AIE4_HWCTX_ERROR (setting has_reset = true).
>
> However, a submitter thread waiting in wait_till_connected_hsa_not_full() in
> drivers/accel/amdxdna/aie4_ctx.c with wait_through_reset = true explicitly
> ignores has_reset:
>
>      ret = wait_event_freezable(priv->job_list_wq,
>                                 aie4_hwctx_connected(hwctx) ||
>                                 (!wait_through_reset &&
>                                  aie4_hwctx_has_reset(hwctx)));
>
> Since the hardware resume failed, the context will never be recreated and
> connected, causing it to block forever. This holds the srcu read lock and
> prevents device removal (causing a deadlock in synchronize_srcu()).

I will fix the synchronize_srcu() deadlock in context teardown by stopping
the hardware context and waking all waiters before destroying it.
The fix will be in upcoming PATCH V2.

/David

>> +     return ret;
>> +}
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=16

^ permalink raw reply	[flat|nested] 33+ messages in thread

* Re: [PATCH V1 14/20] accel/amdxdna: Implement AIE4 command packet building and submission
  2026-09-30  4:04   ` sashiko-bot
@ 2026-10-05 21:52     ` Zhang, Yidong (David)
  0 siblings, 0 replies; 33+ messages in thread
From: Zhang, Yidong (David) @ 2026-10-05 21:52 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: dri-devel


On 9/29/2026 9:04 PM, sashiko-bot@kernel.org wrote:
> [You don't often get email from sashiko-bot@kernel.org. Learn why this is important at https://aka.ms/LearnAboutSenderIdentification ]
>
> Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
> - [Critical] accel/amdxdna: TOCTOU in submit_job_cmds allows arbitrary DMA to freed memory (UAF)
> - [High] accel/amdxdna: Unhandled mid-chain submission failure permanently hangs dma_fence
> - [High] accel/amdxdna: dtrace_buffer address truncated to 48 bits, corrupting memory
> - [High] accel/amdxdna: dma_fence uses unreferenced struct device pointer causing UAF
> - [High] accel/amdxdna: Submitter holding SRCU lock waits forever for reconnect, deadlocking teardown
> - [High] accel/amdxdna: Missing READ_ONCE allows double-fetch of user-shared control fields
> --
>
> commit 3eb430c5785965c7805a7d4c133aa9e24f11869b
> Author: David Zhang <yidong.zhang@amd.com>
>
> accel/amdxdna: Implement AIE4 command packet building and submission
>
> Implement kernel-mode command submission and hardware queue packet
> assembly for AIE4.
>
>> diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
>> --- a/drivers/accel/amdxdna/aie4_ctx.c
>> +++ b/drivers/accel/amdxdna/aie4_ctx.c
> [ ... ]
>
>> @@ -622,7 +622,10 @@ static int wait_till_connected_hsa_not_full(struct amdxdna_hwctx *hwctx,
>>                        if (!ret)
>>                                hsa_not_full = true;
>>                }
>>                ret = wait_event_freezable(priv->job_list_wq,
>>                                           aie4_hwctx_connected(hwctx) ||
>>                                           (!wait_through_reset &&
>>                                            aie4_hwctx_has_reset(hwctx)));
> [Severity: High]
> Can this block indefinitely and cause an SRCU deadlock?
>
> When wait_through_reset is true for the first command, the wait condition
> completely ignores the reset or error state. Since the submitter thread holds
> the hwctx_srcu read lock, if a device unplug occurs and the teardown thread
> calls synchronize_srcu(), the teardown will wait forever for this lock to be
> released, while this code waits for a reconnect that will never happen.
>
>>                mutex_lock(&priv->io_lock);
> [ ... ]
>
>> @@ -696,7 +699,8 @@ static int fill_indirect_pkt(struct amdxdna_hwctx_priv *priv, u64 slot_idx,
>>        for (i = 0; i < entries; i++, dpu++, hipe++) {
>>                struct host_indirect_packet_data *hipd;
>>                u64 indirect_pkt_dev_addr;
>>                u32 uci = dpu->uc_index;
> [Severity: High]
> Can the compiler double-fetch dpu->uc_index from shared memory here?
>
> Because READ_ONCE() is missing, a concurrent userspace thread could modify
> the value immediately after the bounds check against
> HSA_MAX_LEVEL1_INDIRECT_ENTRIES, allowing an out-of-bounds array index into
> priv->umq_indirect_pkts[idx] later in this function.
>
>>                u32 idx;
>>
>>                /*
> [ ... ]
>
>> @@ -715,7 +718,8 @@ static int fill_indirect_pkt(struct amdxdna_hwctx_priv *priv, u64 slot_idx,
>>                        upper_32_bits(dpu->instruction_buffer);
>>                hipd->payload.dtrace_buf_host_addr_low =
>>                        lower_32_bits(dpu->dtrace_buffer);
>>                hipd->payload.dtrace_buf_host_addr_high =
>>                        lower_16_bits(upper_32_bits(dpu->dtrace_buffer));
> [Severity: High]
> Does this truncate the 64-bit user address to 48 bits?
>
> By using lower_16_bits(upper_32_bits()), bits 48-63 of the virtual address
> are silently dropped. On SVA/IOMMU platforms with 5-level paging, user
> virtual addresses can reach 57 bits, which might cause the hardware to write
> trace data to unintended memory locations.
>
>>        }
>>        pkt->pkt_header.common_header.distribute = 1;
> [ ... ]
>
>> @@ -730,7 +732,8 @@ static void fill_direct_pkt(struct amdxdna_hwctx_priv *priv, u64 slot_idx,
>>        ebuf->dpu_control_code_host_addr_high = upper_32_bits(dpu->instruction_buffer);
>>        ebuf->dtrace_buf_host_addr_low = lower_32_bits(dpu->dtrace_buffer);
>>        ebuf->dtrace_buf_host_addr_high = lower_16_bits(upper_32_bits(dpu->dtrace_buffer));
> [Severity: High]
> Similar to fill_indirect_pkt(), can this truncate the user address to 48
> bits, potentially corrupting memory on systems with 5-level paging?
>
>>        pkt->pkt_header.common_header.distribute = 0;
> [ ... ]
>
>> @@ -772,7 +774,8 @@ static int submit_one_cmd(struct amdxdna_hwctx *hwctx,
>>         * against the actual payload size before dereferencing chained+1 DPU
>>         * entries, so a bogus count cannot drive an out-of-bounds read.
>>         */
>>        chained = dpu->chained;
> [Severity: High]
> Could the lack of READ_ONCE() here allow a double-fetch of dpu->chained?
>
> If userspace modifies this value after the bounds check, it could bypass the
> validation.
>
>>        if (chained >= HSA_MAX_LEVEL1_INDIRECT_ENTRIES) {
>>                XDNA_ERR(xdna, "Invalid DPU data");
> [ ... ]
>
>> @@ -1012,7 +1014,8 @@ static int submit_job_cmds(struct amdxdna_hwctx *hwctx,
>>                XDNA_ERR(xdna, "Invalid cmd payload for chained cmd");
>>                return -EINVAL;
>>        }
>>        ccnt = payload->command_count;
> [Severity: High]
> Since payload points to a GEM object mapped into user memory, can the compiler
> optimize this and re-fetch payload->command_count after the bounds validation?
>
> Without READ_ONCE(), userspace could potentially modify the count concurrently
> to bypass the checks and trigger out-of-bounds reads.
>
>>        /*
>>         * A chain (runlist) must fit within the queue. CERT advances the host-visible
> [ ... ]
>
>> @@ -1032,15 +1034,22 @@ static int submit_job_cmds(struct amdxdna_hwctx *hwctx,
>>        for (i = 0; i < ccnt; i++) {
>>                u32 boh = (u32)(payload->data[i]);
>>                struct amdxdna_gem_obj *abo;
>>
>>                abo = amdxdna_gem_get_obj(hwctx->client, boh, AMDXDNA_BO_SHARE);
>>                if (!abo) {
>>                        XDNA_ERR(xdna, "Failed to find cmd BO %u", boh);
>>                        ret = -ENOENT;
>>                        break;
>>                }
>>
>>                /*
>>                 * submit_one_cmd() blocks in wait_till_connected_hsa_not_full() until the
>>                 * ctx is connected and a slot is free, so a concurrent suspend/disconnect
>>                 * is waited out inline rather than returned here. The first sub-command
>>                 * (i == 0, nothing published yet) waits through a TDR reset and runs on
>>                 * the recreated ctx; a later sub-command returns -ECONNRESET if a reset
>>                 * landed while waiting for a slot, so the published prefix is not split
>>                 * across the reset. It also returns -ERESTARTSYS on a signal, or a
>>                 * validation error. Break on any; a published prefix is then reaped by
>>                 * the job worker's reset drain (see below).
>>                 */
>>                ret = submit_one_cmd(hwctx, abo, i + 1 == ccnt, i == 0, &job->seq);
>>                amdxdna_gem_put_obj(abo);
> [Severity: Critical]
> Does this allow arbitrary DMA to freed memory?
>
> The driver reads the BO handle (boh) directly from the user-shared GEM object.
> Userspace can modify this handle concurrently after the DRM scheduler
> validates job->bos.
>
> Because this substituted BO is not part of job->bos, it lacks a tracking
> fence. By immediately dropping the kernel reference via amdxdna_gem_put_obj(),
> userspace can close the handle and free the memory while the hardware
> asynchronously executes DMA operations on the freed physical memory.
>
>>                if (ret)
>>                        break;
> [ ... ]
>
>> @@ -1054,6 +1056,7 @@ static int submit_job_cmds(struct amdxdna_hwctx *hwctx,
>>         * published prefix and reports the partial chain as failed (ABORT). Only when
>>         * nothing was published (i == 0) is the error returned to the caller.
>>         */
>>        if (i > 0)
>>                return 0;
> [Severity: High]
> If submit_one_cmd() fails for a sub-command mid-chain (e.g., due to an
> invalid payload) after the first command is published, this masks the failure
> and returns success.
>
> Since the final command is never published, CHAIN_FLG_LAST_CMD is never sent
> to the hardware. Will this cause the hardware to wait indefinitely, preventing
> job_done() from ever signaling the dma_fence?
>
>>        return ret;
>>   }
> [ ... ]
>
>> @@ -1182,7 +1184,8 @@ int aie4_cmd_submit(struct amdxdna_hwctx *hwctx, struct amdxdna_sched_job *job,
>>
>>        job->out_fence = dma_fence_get(job->fence);
>>        for (i = 0; i < job->bo_cnt; i++)
>>                dma_resv_add_fence(job->bos[i]->resv, job->out_fence, DMA_RESV_USAGE_WRITE);
> [Severity: High]
> Following up on the mid-chain failure issue in submit_job_cmds(), since the
> fence is attached to shared BOs here, if it never signals due to an abandoned
> hardware queue, won't any future access to those BOs by other processes hang
> forever, creating a denial of service?
>
>>        up_read(&xdna->notifier_lock);
>> diff --git a/drivers/accel/amdxdna/amdxdna_ctx.c b/drivers/accel/amdxdna/amdxdna_ctx.c
>> --- a/drivers/accel/amdxdna/amdxdna_ctx.c
>> +++ b/drivers/accel/amdxdna/amdxdna_ctx.c
> [ ... ]
>
>> @@ -47,7 +49,8 @@ static const char *amdxdna_fence_get_timeline_name(struct dma_fence *fence)
>>         * another process). The device outlives any individual context, so
>>         * dev_name() is safe to call at any point during the fence's lifetime.
>>         */
>>        return dev_name(xdna_fence->dev);
> [Severity: High]
> Can this trigger a use-after-free if the sync_file outlives the DRM device?
>
> Userspace querying the timeline after the device has been torn down or
> hot-unplugged would access the freed struct device pointer here.
>
>>   }
>>
> [ ... ]
>
>> @@ -62,7 +65,8 @@ static struct dma_fence *amdxdna_fence_create(struct amdxdna_hwctx *hwctx)
>>        if (!fence)
>>                return NULL;
>>
>>        fence->dev = hwctx->client->xdna->ddev.dev;
> [Severity: High]
> Does this store the struct device pointer without taking a reference via
> get_device()?
>
> This allows the use-after-free mentioned above when the fence outlives the
> device.


The srcu locking issue will be addressed in next version of the patch, 
and those UAF (use-after-free)

cases will be addressed too.

/David

>
>>        spin_lock_init(&fence->lock);
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=14

^ permalink raw reply	[flat|nested] 33+ messages in thread

end of thread, other threads:[~2026-10-05 21:52 UTC | newest]

Thread overview: 33+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30  3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
2026-09-30  3:32 ` [PATCH V1 01/20] accel/amdxdna: Rename NPU3 firmware files David Zhang
2026-09-30  3:32 ` [PATCH V1 02/20] accel/amdxdna: Remove mmap for doorbell David Zhang
2026-09-30  3:32 ` [PATCH V1 03/20] accel/amdxdna: Add CERT firmware version support David Zhang
2026-09-30  3:53   ` sashiko-bot
2026-09-30  3:32 ` [PATCH V1 04/20] accel/amdxdna: Upgrade firmware version to 6.0 David Zhang
2026-09-30  3:32 ` [PATCH V1 05/20] accel/amdxdna: Add NPU3 classic device support David Zhang
2026-09-30  3:32 ` [PATCH V1 06/20] accel/amdxdna: Add AIE version query to aie4_get_info David Zhang
2026-09-30  3:32 ` [PATCH V1 07/20] accel/amdxdna: Add get and set power_mode for AIE4 David Zhang
2026-09-30  3:56   ` sashiko-bot
2026-09-30  3:32 ` [PATCH V1 08/20] accel/amdxdna: Add clock, DPM frequency, and resource info queries " David Zhang
2026-09-30  3:32 ` [PATCH V1 09/20] accel/amdxdna: Add context switch hysteresis with debugfs control David Zhang
2026-09-30  3:52   ` sashiko-bot
2026-09-30  3:32 ` [PATCH V1 10/20] accel/amdxdna: Refactor AIE4 hardware initialization sequence David Zhang
2026-09-30  3:32 ` [PATCH V1 11/20] accel/amdxdna: Decouple AIE4 doorbell and MSI-X notification transport hooks David Zhang
2026-09-30  3:32 ` [PATCH V1 12/20] accel/amdxdna: Implement AIE4 kernel queue lifecycle and memory layout David Zhang
2026-09-30  4:00   ` sashiko-bot
2026-09-30  3:32 ` [PATCH V1 13/20] accel/amdxdna: Prepare for AIE4 command submission David Zhang
2026-09-30  4:00   ` sashiko-bot
2026-09-30  3:32 ` [PATCH V1 14/20] accel/amdxdna: Implement AIE4 command packet building and submission David Zhang
2026-09-30  4:04   ` sashiko-bot
2026-10-05 21:52     ` Zhang, Yidong (David)
2026-09-30  3:32 ` [PATCH V1 15/20] accel/amdxdna: Finalize runtime PM before acquiring dev_lock on removal David Zhang
2026-09-30  3:32 ` [PATCH V1 16/20] accel/amdxdna: Implement AIE4 suspend and resume David Zhang
2026-09-30  4:07   ` sashiko-bot
2026-10-05 21:49     ` Zhang, Yidong (David)
2026-09-30  3:32 ` [PATCH V1 17/20] accel/amdxdna: Link SR-IOV VFs for power management sequencing David Zhang
2026-09-30  3:59   ` sashiko-bot
2026-09-30  3:32 ` [PATCH V1 18/20] accel/amdxdna: Implement runtime suspend and resume support David Zhang
2026-09-30  4:00   ` sashiko-bot
2026-09-30  3:32 ` [PATCH V1 19/20] accel/amdxdna: Add stub hwctx_config for AIE4 David Zhang
2026-09-30  3:59   ` sashiko-bot
2026-09-30  3:32 ` [PATCH V1 20/20] accel/amdxdna: Enable AIE4 firmware logging to DRAM David Zhang

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox