Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement
@ 2026-08-06  6:34 Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 1/9] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
                   ` (8 more replies)
  0 siblings, 9 replies; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

Based on Igor Paunovic's "[PATCH v2] accel/rocket: request the core
clocks by name", which is waiting on Tomeu and which v5 was carrying a
duplicate of:

  https://lore.kernel.org/linux-rockchip/20260729130743.128876-1-royalnet026@gmail.com/

That duplication was the coordination problem Igor raised on v5. Basing
on it rather than re-adding the same hunk removes it, and it makes the
patch split below fall out cleanly.

Tested on a Radxa ROCK 4D, on next-20260730.

Changes in v6
-------------

  * accel/rocket: the enablement is split in two, as Diederik de Haas
    asked for and Igor seconded with the observation that v5's 6/8 no
    longer applied on a tree carrying the clocks fix. 6/9 is preparation
    only, the soc_data plumbing and the bulk counts, with RK3588 keeping
    four clocks and two resets. 7/9 is the RK3576 enablement.

  * accel/rocket: poll_dying was a one-way latch. rocket_job_fini() set
    it and nothing cleared it, while struct rocket_core survives an
    unbind whenever another core stays bound, so a rebinding core would
    never retire anything through the poll again. rocket_job_init() now
    clears it. Found by Igor, who also explained why the ROCK 4D test
    could not have caught it: with one core enabled every unbind is the
    last one and the core array is reallocated.

  * accel/rocket: rocket_core_reset() still used ARRAY_SIZE(core->resets)
    after acquisition moved to soc->num_resets, so on RK3576 it walked a
    reset that was never acquired. Benign, since the reset core accepts a
    NULL rstc, but patch 5 has just made this path load bearing. Also
    from Igor. It is in 6/9 with the rest of the count changes.

  * accel/rocket: the two register writes in rocket_job_handle_irq() were
    outside job_lock while hw_submit() writes OPERATION_ENABLE inside it,
    so the completion's zero could land after a submit's one and stop a
    task that had just started. They are under the lock now. This is the
    only behaviour change 6/9 makes to RK3588.

  * accel/rocket: the poll work now skips its register writes when no job
    is in flight. Unlike an interrupt it has no hardware condition to
    ack, and with the job already retired by the interrupt path the
    device can have autosuspended underneath it.

  * accel/rocket: where both completion paths are live, a completion
    whose submit the other path has already retired no longer goes on to
    start a further task. v5 checked this only on the poll side, which
    left the mirror case open.

  * pmdomain/rockchip: dev_err_probe() for the reset acquisition, as
    Philipp Zabel asked.

Philipp, on the other half of that: every devm_reset_control_* variant
resolves against dev->of_node, and these resets are on the power-domain
child node rather than the PMU's own node, so the devm form would take
them from the wrong node. The clocks a few lines above use the same
of_*() and manual-put pattern for the same reason. Happy to add a
devm_add_action_or_reset() instead if you would rather the lifetime were
devres managed.

Igor, thank you for the RK3588 characterisation and for running the
per-core unbind and rebind. Everything in v6 that came from your review
is above. The completion path changed again, so it needs another look
rather than a carried tag.

Verified on hardware with every debug knob off: the NPU probes with the
two domain list and no attach failure, a convolution submitted after a
fresh resume is byte exact against the CPU reference, and unbind and
rebind is clean with no warning. Every patch in the series builds on its
own.

I have deliberately stopped quoting the "runs byte exact N times in a
row" figure I used in earlier cover letters. See below: that test could
not distinguish a recomputation from an untouched buffer, and it was
measuring the latter.

What is still wrong
-------------------

The failure is much narrower than I have been describing it, and most of
what I said about it in v3 through v5 was reading an artefact.

I had been reporting that a single convolution is byte exact, that
re-running the same one is byte exact every time, and that what fails is
loading a different configuration after it. The first part is true. The
rest was a stale buffer.

Every one of those "re-runs byte exact" measurements fed the same input
each time, so a correct recomputation and an output buffer that nothing
had touched since the first submit look identical. Feeding the same model
a different input and checksumming the output BO in place separates them:

  A(input X), first submit of the session   correct, crc32 20a556ae
  A(input Y), no reset in between           wrong,   crc32 20a556ae
  A(input Y), after a runtime resume        correct, crc32 dda67317

The third line moves the checksum, so it does see the block's writes. The
second does not move it, with the same configuration loaded and only the
input data different. The second submit does not write its output at all.

So the shape is: only the first submit after a reset computes. Everything
after it is a no-op that leaves the output buffer holding whatever was
there before. "A works, B fails, A works again" needs no configuration
story: A computes, B is a no-op and its freshly zeroed buffer reads back
as the zero point, and A again is a no-op returning A's old result.

That also settles the question Igor raised on v5, and not in favour of
what I claimed there. His alternative was that the block never stops
executing the resident configuration and writes to the previous task's
addresses, which would look the same from the failing job's own buffers.
With the watched buffer latched rather than followed, the resident job's
output is unchanged across the failing submit. It writes nothing anywhere.

Two corrections to the record, both mine:

  * v5 said the failing submit "computes byte exact from a buffer full of
    0xdeadbeef" after its regcmd was corrupted. It does not compute. The
    narrower statement survives, that a repeat submit does not re-read its
    regcmd, because it does not read anything.

  * v3 and v4 said the failing job "writes out a zero point surface" and I
    read that as the MAC producing nothing. Nothing was ever measured
    about the MAC. The buffer is simply never written, and a zeroed shmem
    page plus the +0x80 that teflon applies on readback is 128.

The ping-pong lead from the v3 thread stays retired, and so does the
configuration-load framing that replaced it. The question is now why the
block accepts exactly one task per reset. That is narrower than anything
I have had before, it matches the interrupt behaviour already in this
series, and it means the userspace side was never involved.

One incidental register fact, in case it means something to someone:
PC_BASE_ADDRESS reads back 0x00000000 immediately after being written, on
every submit.

I used Claude Opus 5 to trim this series out of my debugging tree and
generate the diffs.

Jiaxing Hu (9):
  dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core
  dt-bindings: power: rockchip: allow resets in a power domain node
  dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set
  pmdomain/rockchip: add optional per-domain power-on settle delay
  pmdomain/rockchip: cycle optional power-domain resets on power-on
  accel/rocket: select the per-core clock and reset counts from match data
  accel/rocket: add RK3576 NPU (RKNN) support
  arm64: dts: rockchip: rk3576: add NPU (RKNN) nodes
  arm64: dts: rockchip: rk3576-rock-4d: enable NPU

 .../devicetree/bindings/iommu/rockchip,iommu.yaml  |   8 ++
 .../bindings/npu/rockchip,rk3588-rknn-core.yaml    |  47 ++++++-
 .../bindings/power/rockchip,power-controller.yaml  |   8 ++
 arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts    |  10 ++
 arch/arm64/boot/dts/rockchip/rk3576.dtsi           |  80 ++++++++++-
 drivers/accel/rocket/rocket_core.c                 |  26 +++-
 drivers/accel/rocket/rocket_core.h                 |  20 ++-
 drivers/accel/rocket/rocket_device.c               |   4 +
 drivers/accel/rocket/rocket_drv.c                  |  22 ++-
 drivers/accel/rocket/rocket_job.c                  | 149 +++++++++++++++++++--
 drivers/pmdomain/rockchip/pm-domains.c             |  71 +++++++---
 11 files changed, 396 insertions(+), 49 deletions(-)

--
2.43.0


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [RFC PATCH v6 1/9] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core
  2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
@ 2026-08-06  6:34 ` Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 2/9] dt-bindings: power: rockchip: allow resets in a power domain node Jiaxing Hu
                   ` (7 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

The RK3576 NPU has two cores of the same RKNN block the RK3588 binding
already describes, but it wires them up differently: two extra CBUF
clocks, two power domains per core, and a single reset instead of two.
It also has no NPU SRAM supply.

Widen the property ranges to cover both, then pin each SoC back to its
own shape in allOf so nothing loosens for RK3588, and keep sram-supply
required for rockchip,rk3588-rknn-core only.

Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
 .../npu/rockchip,rk3588-rknn-core.yaml        | 47 +++++++++++++++++--
 1 file changed, 44 insertions(+), 3 deletions(-)

diff --git a/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml b/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml
index caca2a490..3b611b64c 100644
--- a/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml
+++ b/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml
@@ -21,6 +21,7 @@ properties:
 
   compatible:
     enum:
+      - rockchip,rk3576-rknn-core
       - rockchip,rk3588-rknn-core
 
   reg:
@@ -33,14 +34,18 @@ properties:
       - const: core # Main NPU core processing unit registers
 
   clocks:
-    maxItems: 4
+    minItems: 4
+    maxItems: 6
 
   clock-names:
+    minItems: 4
     items:
       - const: aclk
       - const: hclk
       - const: npu
       - const: pclk
+      - const: aclk_cbuf
+      - const: hclk_cbuf
 
   interrupts:
     maxItems: 1
@@ -51,12 +56,15 @@ properties:
   npu-supply: true
 
   power-domains:
-    maxItems: 1
+    minItems: 1
+    maxItems: 2
 
   resets:
+    minItems: 1
     maxItems: 2
 
   reset-names:
+    minItems: 1
     items:
       - const: srst_a
       - const: srst_h
@@ -75,7 +83,40 @@ required:
   - resets
   - reset-names
   - npu-supply
-  - sram-supply
+
+allOf:
+  - if:
+      properties:
+        compatible:
+          contains:
+            const: rockchip,rk3588-rknn-core
+    then:
+      properties:
+        clocks:
+          maxItems: 4
+        clock-names:
+          maxItems: 4
+        power-domains:
+          maxItems: 1
+        resets:
+          minItems: 2
+        reset-names:
+          minItems: 2
+      required:
+        - sram-supply
+    else:
+      properties:
+        clocks:
+          minItems: 6
+        clock-names:
+          minItems: 6
+        power-domains:
+          minItems: 2
+        resets:
+          maxItems: 1
+        reset-names:
+          maxItems: 1
+        sram-supply: false
 
 additionalProperties: false
 
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [RFC PATCH v6 2/9] dt-bindings: power: rockchip: allow resets in a power domain node
  2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 1/9] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
@ 2026-08-06  6:34 ` Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 3/9] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set Jiaxing Hu
                   ` (6 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

Some domains do not come up in a usable state on their own and need
their resets cycled once power is on. The RK3576 NPU domains are one
case: without it the first access after power-on takes an async SError.

pd-node has no resets property and every nesting level is
unevaluatedProperties: false, so describing that in DT is rejected
today. Add it alongside clocks.

Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
 .../bindings/power/rockchip,power-controller.yaml         | 8 ++++++++
 1 file changed, 8 insertions(+)

diff --git a/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml b/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml
index b41db576f..f23c1a118 100644
--- a/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml
+++ b/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml
@@ -136,6 +136,14 @@ $defs:
           A number of phandles to clocks that need to be enabled
           while power domain switches state.
 
+      resets:
+        minItems: 1
+        maxItems: 30
+        description: |
+          A number of phandles to resets that need to be cycled once the power
+          domain has been switched on, for domains whose logic does not come up
+          in a usable state by itself.
+
       domain-supply:
         description: domain regulator supply.
 
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [RFC PATCH v6 3/9] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set
  2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 1/9] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 2/9] dt-bindings: power: rockchip: allow resets in a power domain node Jiaxing Hu
@ 2026-08-06  6:34 ` Jiaxing Hu
  2026-08-06  9:23   ` Diederik de Haas
  2026-08-06  6:34 ` [RFC PATCH v6 4/9] pmdomain/rockchip: add optional per-domain power-on settle delay Jiaxing Hu
                   ` (5 subsequent siblings)
  8 siblings, 1 reply; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

The RK3576 NPU MMUs need more than aclk and iface. With only those two
enabled the MMU accepts reads but silently drops register writes: a
DTE_ADDR value written from the power domain, while the domain clocks
are still on, reads back correctly, and the write rk_iommu_resume() does
microseconds later does not land at all. The vendor DT names the CBUF
clocks as that MMU's interface clocks and its driver keeps every NPU
clock on for as long as the device is powered.

The driver side of this is already upstream, commit 841363ebb508
("iommu/rockchip: Take all DT clocks"), which switched rk_iommu to
devm_clk_bulk_get_all(). Widen the schema to match so those nodes can
be described. minItems stays at 2, so every existing devicetree, which
all carry exactly aclk and iface, is unaffected.

Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
 .../devicetree/bindings/iommu/rockchip,iommu.yaml         | 8 ++++++++
 1 file changed, 8 insertions(+)

diff --git a/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml b/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
index 6ce41d11f..a3cedcaaa 100644
--- a/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
+++ b/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
@@ -42,14 +42,22 @@ properties:
     minItems: 1
 
   clocks:
+    minItems: 2
     items:
       - description: Core clock
       - description: Interface clock
+      - description: Compute clock, RK3576 NPU MMUs only
+      - description: Convolution buffer core clock, RK3576 NPU MMUs only
+      - description: Convolution buffer interface clock, RK3576 NPU MMUs only
 
   clock-names:
+    minItems: 2
     items:
       - const: aclk
       - const: iface
+      - const: npu
+      - const: aclk_cbuf
+      - const: hclk_cbuf
 
   "#iommu-cells":
     const: 0
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [RFC PATCH v6 4/9] pmdomain/rockchip: add optional per-domain power-on settle delay
  2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
                   ` (2 preceding siblings ...)
  2026-08-06  6:34 ` [RFC PATCH v6 3/9] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set Jiaxing Hu
@ 2026-08-06  6:34 ` Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 5/9] pmdomain/rockchip: cycle optional power-domain resets on power-on Jiaxing Hu
                   ` (4 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

The RK3576 NPU domains need a short settle time after the idle request
is released before the QoS registers behind the domain answer. Without
it rockchip_pmu_restore_qos() reads back zeroes, and the NPU throws an
async SError on the first cold power-on.

Give rockchip_domain_info an optional delay_us and wait for it between
releasing idle and restoring QoS. Rename DOMAIN_M_O_R_G to
DOMAIN_M_O_R_G_W, since the suffixes name the fields the macro sets and
this one now also carries a wakeup delay; RK3576 is its only user, so
the old spelling is not kept around.

Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
 drivers/pmdomain/rockchip/pm-domains.c | 52 +++++++++++++++-----------
 1 file changed, 30 insertions(+), 22 deletions(-)

diff --git a/drivers/pmdomain/rockchip/pm-domains.c b/drivers/pmdomain/rockchip/pm-domains.c
index ba66ae719..e1857f878 100644
--- a/drivers/pmdomain/rockchip/pm-domains.c
+++ b/drivers/pmdomain/rockchip/pm-domains.c
@@ -18,6 +18,7 @@
 #include <linux/of_address.h>
 #include <linux/of_clk.h>
 #include <linux/clk.h>
+#include <linux/delay.h>
 #include <linux/regmap.h>
 #include <linux/regulator/consumer.h>
 #include <linux/mfd/syscon.h>
@@ -59,6 +60,7 @@ struct rockchip_domain_info {
 	u32 pwr_offset;
 	u32 mem_offset;
 	u32 req_offset;
+	u32 delay_us;
 };
 
 struct rockchip_pmu_info {
@@ -185,7 +187,7 @@ struct rockchip_pmu {
 	.need_regulator = regulator,			\
 }
 
-#define DOMAIN_M_O_R_G(_name, p_offset, pwr, status, m_offset, m_status, r_status, r_offset, req, idle, ack, g_mask, wakeup)	\
+#define DOMAIN_M_O_R_G_W(_name, p_offset, pwr, status, m_offset, m_status, r_status, r_offset, req, idle, ack, g_mask, delay, wakeup)	\
 {							\
 	.name = _name,					\
 	.pwr_offset = p_offset,				\
@@ -200,6 +202,7 @@ struct rockchip_pmu {
 	.req_mask = (req),				\
 	.idle_mask = (idle),				\
 	.clk_ungate_mask = (g_mask),			\
+	.delay_us = (delay),				\
 	.ack_mask = (ack),				\
 	.active_wakeup = wakeup,			\
 }
@@ -258,8 +261,8 @@ struct rockchip_pmu {
 #define DOMAIN_RK3568(name, pwr, req, wakeup, regulator)		\
 	DOMAIN_M_R(name, pwr, pwr, req, req, req, wakeup, regulator)
 
-#define DOMAIN_RK3576(name, p_offset, pwr, status, r_status, r_offset, req, idle, g_mask, wakeup)	\
-	DOMAIN_M_O_R_G(name, p_offset, pwr, status, 0, r_status, r_status, r_offset, req, idle, idle, g_mask, wakeup)
+#define DOMAIN_RK3576(name, p_offset, pwr, status, r_status, r_offset, req, idle, g_mask, delay, wakeup)	\
+	DOMAIN_M_O_R_G_W(name, p_offset, pwr, status, 0, r_status, r_status, r_offset, req, idle, idle, g_mask, delay, wakeup)
 
 /*
  * Dynamic Memory Controller may need to coordinate with us -- see
@@ -681,6 +684,10 @@ static int rockchip_pd_power(struct rockchip_pm_domain *pd, bool power_on)
 		if (ret < 0)
 			goto out;
 
+		/* Some domains need to settle before the QoS registers answer. */
+		if (pd->info->delay_us)
+			udelay(pd->info->delay_us);
+
 		rockchip_pmu_restore_qos(pd);
 	}
 
@@ -1300,25 +1307,26 @@ static const struct rockchip_domain_info rk3568_pm_domains[] = {
 };
 
 static const struct rockchip_domain_info rk3576_pm_domains[] = {
-	[RK3576_PD_NPU]		= DOMAIN_RK3576("npu",    0x0, BIT(0),  BIT(0), 0,       0x0, 0,       0,       0,       false),
-	[RK3576_PD_NVM]		= DOMAIN_RK3576("nvm",    0x0, BIT(6),  0,      BIT(6),  0x4, BIT(2),  BIT(18), BIT(2),  false),
-	[RK3576_PD_SDGMAC]	= DOMAIN_RK3576("sdgmac", 0x0, BIT(7),  0,      BIT(7),  0x4, BIT(1),  BIT(17), 0x6,     false),
-	[RK3576_PD_AUDIO]	= DOMAIN_RK3576("audio",  0x0, BIT(8),  0,      BIT(8),  0x4, BIT(0),  BIT(16), BIT(0),  false),
-	[RK3576_PD_PHP]		= DOMAIN_RK3576("php",    0x0, BIT(9),  0,      BIT(9),  0x0, BIT(15), BIT(15), BIT(15), false),
-	[RK3576_PD_SUBPHP]	= DOMAIN_RK3576("subphp", 0x0, BIT(10), 0,      BIT(10), 0x0, 0,       0,       0,       false),
-	[RK3576_PD_VOP]		= DOMAIN_RK3576("vop",    0x0, BIT(11), 0,      BIT(11), 0x0, 0x6000,  0x6000,  0x6000,  false),
-	[RK3576_PD_VO1]		= DOMAIN_RK3576("vo1",    0x0, BIT(14), 0,      BIT(14), 0x0, BIT(12), BIT(12), 0x7000,  false),
-	[RK3576_PD_VO0]		= DOMAIN_RK3576("vo0",    0x0, BIT(15), 0,      BIT(15), 0x0, BIT(11), BIT(11), 0x6800,  false),
-	[RK3576_PD_USB]		= DOMAIN_RK3576("usb",    0x4, BIT(0),  0,      BIT(16), 0x0, BIT(10), BIT(10), 0x6400,  true),
-	[RK3576_PD_VI]		= DOMAIN_RK3576("vi",     0x4, BIT(1),  0,      BIT(17), 0x0, BIT(9),  BIT(9),  BIT(9),  false),
-	[RK3576_PD_VEPU0]	= DOMAIN_RK3576("vepu0",  0x4, BIT(2),  0,      BIT(18), 0x0, BIT(7),  BIT(7),  0x280,   false),
-	[RK3576_PD_VEPU1]	= DOMAIN_RK3576("vepu1",  0x4, BIT(3),  0,      BIT(19), 0x0, BIT(8),  BIT(8),  BIT(8),  false),
-	[RK3576_PD_VDEC]	= DOMAIN_RK3576("vdec",   0x4, BIT(4),  0,      BIT(20), 0x0, BIT(6),  BIT(6),  BIT(6),  false),
-	[RK3576_PD_VPU]		= DOMAIN_RK3576("vpu",    0x4, BIT(5),  0,      BIT(21), 0x0, BIT(5),  BIT(5),  BIT(5),  false),
-	[RK3576_PD_NPUTOP]	= DOMAIN_RK3576("nputop", 0x4, BIT(6),  0,      BIT(22), 0x0, 0x18,    0x18,    0x18,    false),
-	[RK3576_PD_NPU0]	= DOMAIN_RK3576("npu0",   0x4, BIT(7),  0,      BIT(23), 0x0, BIT(1),  BIT(1),  0x1a,    false),
-	[RK3576_PD_NPU1]	= DOMAIN_RK3576("npu1",   0x4, BIT(8),  0,      BIT(24), 0x0, BIT(2),  BIT(2),  0x1c,    false),
-	[RK3576_PD_GPU]		= DOMAIN_RK3576("gpu",    0x4, BIT(9),  0,      BIT(25), 0x0, BIT(0),  BIT(0),  BIT(0),  false),
+	/*                                            name    p_offset pwr      status  r_status r_offset req      idle     g_mask   delay wakeup */
+	[RK3576_PD_NPU]		= DOMAIN_RK3576("npu",    0x0, BIT(0),  BIT(0), 0,       0x0, 0,       0,       0,       0,    false),
+	[RK3576_PD_NVM]		= DOMAIN_RK3576("nvm",    0x0, BIT(6),  0,      BIT(6),  0x4, BIT(2),  BIT(18), BIT(2),  0,    false),
+	[RK3576_PD_SDGMAC]	= DOMAIN_RK3576("sdgmac", 0x0, BIT(7),  0,      BIT(7),  0x4, BIT(1),  BIT(17), 0x6,     0,    false),
+	[RK3576_PD_AUDIO]	= DOMAIN_RK3576("audio",  0x0, BIT(8),  0,      BIT(8),  0x4, BIT(0),  BIT(16), BIT(0),  0,    false),
+	[RK3576_PD_PHP]		= DOMAIN_RK3576("php",    0x0, BIT(9),  0,      BIT(9),  0x0, BIT(15), BIT(15), BIT(15), 0,    false),
+	[RK3576_PD_SUBPHP]	= DOMAIN_RK3576("subphp", 0x0, BIT(10), 0,      BIT(10), 0x0, 0,       0,       0,       0,    false),
+	[RK3576_PD_VOP]		= DOMAIN_RK3576("vop",    0x0, BIT(11), 0,      BIT(11), 0x0, 0x6000,  0x6000,  0x6000,  0,    false),
+	[RK3576_PD_VO1]		= DOMAIN_RK3576("vo1",    0x0, BIT(14), 0,      BIT(14), 0x0, BIT(12), BIT(12), 0x7000,  0,    false),
+	[RK3576_PD_VO0]		= DOMAIN_RK3576("vo0",    0x0, BIT(15), 0,      BIT(15), 0x0, BIT(11), BIT(11), 0x6800,  0,    false),
+	[RK3576_PD_USB]		= DOMAIN_RK3576("usb",    0x4, BIT(0),  0,      BIT(16), 0x0, BIT(10), BIT(10), 0x6400,  0,    true),
+	[RK3576_PD_VI]		= DOMAIN_RK3576("vi",     0x4, BIT(1),  0,      BIT(17), 0x0, BIT(9),  BIT(9),  BIT(9),  0,    false),
+	[RK3576_PD_VEPU0]	= DOMAIN_RK3576("vepu0",  0x4, BIT(2),  0,      BIT(18), 0x0, BIT(7),  BIT(7),  0x280,   0,    false),
+	[RK3576_PD_VEPU1]	= DOMAIN_RK3576("vepu1",  0x4, BIT(3),  0,      BIT(19), 0x0, BIT(8),  BIT(8),  BIT(8),  0,    false),
+	[RK3576_PD_VDEC]	= DOMAIN_RK3576("vdec",   0x4, BIT(4),  0,      BIT(20), 0x0, BIT(6),  BIT(6),  BIT(6),  0,    false),
+	[RK3576_PD_VPU]		= DOMAIN_RK3576("vpu",    0x4, BIT(5),  0,      BIT(21), 0x0, BIT(5),  BIT(5),  BIT(5),  0,    false),
+	[RK3576_PD_NPUTOP]	= DOMAIN_RK3576("nputop", 0x4, BIT(6),  0,      BIT(22), 0x0, 0x18,    0x18,    0x18,    15,   false),
+	[RK3576_PD_NPU0]	= DOMAIN_RK3576("npu0",   0x4, BIT(7),  0,      BIT(23), 0x0, BIT(1),  BIT(1),  0x1a,    15,   false),
+	[RK3576_PD_NPU1]	= DOMAIN_RK3576("npu1",   0x4, BIT(8),  0,      BIT(24), 0x0, BIT(2),  BIT(2),  0x1c,    15,   false),
+	[RK3576_PD_GPU]		= DOMAIN_RK3576("gpu",    0x4, BIT(9),  0,      BIT(25), 0x0, BIT(0),  BIT(0),  BIT(0),  0,    false),
 };
 
 static const struct rockchip_domain_info rk3588_pm_domains[] = {
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [RFC PATCH v6 5/9] pmdomain/rockchip: cycle optional power-domain resets on power-on
  2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
                   ` (3 preceding siblings ...)
  2026-08-06  6:34 ` [RFC PATCH v6 4/9] pmdomain/rockchip: add optional per-domain power-on settle delay Jiaxing Hu
@ 2026-08-06  6:34 ` Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 6/9] accel/rocket: select the per-core clock and reset counts from match data Jiaxing Hu
                   ` (3 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

Some Rockchip domains come out of power-on with their bus interface in
an undefined state. On the RK3576 NPU this shows up as a hang on the
first register access after the domain is switched on, and pulsing the
domain's resets at this point clears it.

Take the domain node's resets if it has any, and pulse them between
releasing idle and restoring QoS. The resets are optional, so domains
that do not list any are unaffected.

Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
 drivers/pmdomain/rockchip/pm-domains.c | 19 +++++++++++++++++++
 1 file changed, 19 insertions(+)

diff --git a/drivers/pmdomain/rockchip/pm-domains.c b/drivers/pmdomain/rockchip/pm-domains.c
index e1857f878..4eebb5d99 100644
--- a/drivers/pmdomain/rockchip/pm-domains.c
+++ b/drivers/pmdomain/rockchip/pm-domains.c
@@ -19,6 +19,7 @@
 #include <linux/of_clk.h>
 #include <linux/clk.h>
 #include <linux/delay.h>
+#include <linux/reset.h>
 #include <linux/regmap.h>
 #include <linux/regulator/consumer.h>
 #include <linux/mfd/syscon.h>
@@ -103,6 +104,7 @@ struct rockchip_pm_domain {
 	struct clk_bulk_data *clks;
 	struct device_node *node;
 	struct regulator *supply;
+	struct reset_control *resets;
 };
 
 struct rockchip_pmu {
@@ -688,6 +690,13 @@ static int rockchip_pd_power(struct rockchip_pm_domain *pd, bool power_on)
 		if (pd->info->delay_us)
 			udelay(pd->info->delay_us);
 
+		/* Optional: some domains need their resets cycled after power-on. */
+		if (pd->resets) {
+			reset_control_assert(pd->resets);
+			udelay(10);
+			reset_control_deassert(pd->resets);
+		}
+
 		rockchip_pmu_restore_qos(pd);
 	}
 
@@ -857,6 +866,14 @@ static int rockchip_pm_add_one_domain(struct rockchip_pmu *pmu,
 	if (error)
 		goto err_put_clocks;
 
+	pd->resets = of_reset_control_array_get_optional_exclusive(node);
+	if (IS_ERR(pd->resets)) {
+		error = dev_err_probe(pmu->dev, PTR_ERR(pd->resets),
+				      "%pOFn: failed to get resets\n", node);
+		pd->resets = NULL;
+		goto err_unprepare_clocks;
+	}
+
 	pd->num_qos = of_count_phandle_with_args(node, "pm_qos",
 						 NULL);
 
@@ -927,6 +944,7 @@ static int rockchip_pm_add_one_domain(struct rockchip_pmu *pmu,
 	clk_bulk_unprepare(pd->num_clks, pd->clks);
 err_put_clocks:
 	clk_bulk_put(pd->num_clks, pd->clks);
+	reset_control_put(pd->resets);
 	return error;
 }
 
@@ -945,6 +963,7 @@ static void rockchip_pm_remove_one_domain(struct rockchip_pm_domain *pd)
 
 	clk_bulk_unprepare(pd->num_clks, pd->clks);
 	clk_bulk_put(pd->num_clks, pd->clks);
+	reset_control_put(pd->resets);
 
 	/* protect the zeroing of pm->num_clks */
 	mutex_lock(&pd->pmu->mutex);
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [RFC PATCH v6 6/9] accel/rocket: select the per-core clock and reset counts from match data
  2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
                   ` (4 preceding siblings ...)
  2026-08-06  6:34 ` [RFC PATCH v6 5/9] pmdomain/rockchip: cycle optional power-domain resets on power-on Jiaxing Hu
@ 2026-08-06  6:34 ` Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 7/9] accel/rocket: add RK3576 NPU (RKNN) support Jiaxing Hu
                   ` (2 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

The RK3576 carries the same RKNN block with a different set of clocks and
resets, so the counts cannot stay compile-time constants. Add a soc_data
struct to the of_device_id match data and take the bulk counts from it.
RK3588 keeps four clocks and two resets, so nothing changes for it.

rocket_core_reset() is switched over as well. It is the same array, and
leaving it on ARRAY_SIZE() would walk entries that were never acquired
once a SoC asks for fewer.

While moving through this path, take the two register writes in
rocket_job_handle_irq() under job_lock. rocket_job_hw_submit() writes
OPERATION_ENABLE from inside the lock, so a completion handled outside it
can write the zero after that one and stop a task that has just started.

Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
 drivers/accel/rocket/rocket_core.c |  8 ++++----
 drivers/accel/rocket/rocket_core.h |  9 ++++++++-
 drivers/accel/rocket/rocket_drv.c  | 12 +++++++++---
 drivers/accel/rocket/rocket_job.c  | 12 +++++++++---
 4 files changed, 30 insertions(+), 11 deletions(-)

diff --git a/drivers/accel/rocket/rocket_core.c b/drivers/accel/rocket/rocket_core.c
index 5dd260bac..b202d1581 100644
--- a/drivers/accel/rocket/rocket_core.c
+++ b/drivers/accel/rocket/rocket_core.c
@@ -23,7 +23,7 @@ int rocket_core_init(struct rocket_core *core)
 
 	core->resets[0].id = "srst_a";
 	core->resets[1].id = "srst_h";
-	err = devm_reset_control_bulk_get_exclusive(&pdev->dev, ARRAY_SIZE(core->resets),
+	err = devm_reset_control_bulk_get_exclusive(&pdev->dev, core->soc->num_resets,
 						    core->resets);
 	if (err)
 		return dev_err_probe(dev, err, "failed to get resets for core %d\n", core->index);
@@ -32,7 +32,7 @@ int rocket_core_init(struct rocket_core *core)
 	core->clks[1].id = "hclk";
 	core->clks[2].id = "npu";
 	core->clks[3].id = "pclk";
-	err = devm_clk_bulk_get(dev, ARRAY_SIZE(core->clks), core->clks);
+	err = devm_clk_bulk_get(dev, core->soc->num_clks, core->clks);
 	if (err)
 		return dev_err_probe(dev, err, "failed to get clocks for core %d\n", core->index);
 
@@ -109,9 +109,9 @@ void rocket_core_fini(struct rocket_core *core)
 
 void rocket_core_reset(struct rocket_core *core)
 {
-	reset_control_bulk_assert(ARRAY_SIZE(core->resets), core->resets);
+	reset_control_bulk_assert(core->soc->num_resets, core->resets);
 
 	udelay(10);
 
-	reset_control_bulk_deassert(ARRAY_SIZE(core->resets), core->resets);
+	reset_control_bulk_deassert(core->soc->num_resets, core->resets);
 }
diff --git a/drivers/accel/rocket/rocket_core.h b/drivers/accel/rocket/rocket_core.h
index f6d738285..0f424bb86 100644
--- a/drivers/accel/rocket/rocket_core.h
+++ b/drivers/accel/rocket/rocket_core.h
@@ -27,16 +27,23 @@
 #define rocket_core_writel(core, reg, value) \
 	writel(value, (core)->core_iomem + (REG_CORE_##reg) - REG_CORE_S_STATUS)
 
+/* Per-SoC differences, selected by the of_device_id match data. */
+struct rocket_soc_data {
+	unsigned int num_clks;		/* clk_bulk count */
+	unsigned int num_resets;	/* reset_bulk count */
+};
+
 struct rocket_core {
 	struct device *dev;
 	struct rocket_device *rdev;
+	const struct rocket_soc_data *soc;
 	unsigned int index;
 
 	int irq;
 	void __iomem *pc_iomem;
 	void __iomem *cna_iomem;
 	void __iomem *core_iomem;
-	struct clk_bulk_data clks[4];
+	struct clk_bulk_data clks[6];
 	struct reset_control_bulk_data resets[2];
 
 	struct iommu_group *iommu_group;
diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocket_drv.c
index 8bbbce594..6e7dc91c5 100644
--- a/drivers/accel/rocket/rocket_drv.c
+++ b/drivers/accel/rocket/rocket_drv.c
@@ -176,6 +176,7 @@ static int rocket_probe(struct platform_device *pdev)
 
 	rdev->cores[core].rdev = rdev;
 	rdev->cores[core].dev = &pdev->dev;
+	rdev->cores[core].soc = of_device_get_match_data(&pdev->dev);
 	rdev->cores[core].index = core;
 
 	rdev->num_cores++;
@@ -213,8 +214,13 @@ static void rocket_remove(struct platform_device *pdev)
 	}
 }
 
+static const struct rocket_soc_data rk3588_soc_data = {
+	.num_clks = 4,
+	.num_resets = 2,
+};
+
 static const struct of_device_id dt_match[] = {
-	{ .compatible = "rockchip,rk3588-rknn-core" },
+	{ .compatible = "rockchip,rk3588-rknn-core", .data = &rk3588_soc_data },
 	{}
 };
 MODULE_DEVICE_TABLE(of, dt_match);
@@ -240,7 +246,7 @@ static int rocket_device_runtime_resume(struct device *dev)
 	if (core < 0)
 		return -ENODEV;
 
-	err = clk_bulk_prepare_enable(ARRAY_SIZE(rdev->cores[core].clks), rdev->cores[core].clks);
+	err = clk_bulk_prepare_enable(rdev->cores[core].soc->num_clks, rdev->cores[core].clks);
 	if (err) {
 		dev_err(dev, "failed to enable (%d) clocks for core %d\n", err, core);
 		return err;
@@ -260,7 +266,7 @@ static int rocket_device_runtime_suspend(struct device *dev)
 	if (!rocket_job_is_idle(&rdev->cores[core]))
 		return -EBUSY;
 
-	clk_bulk_disable_unprepare(ARRAY_SIZE(rdev->cores[core].clks), rdev->cores[core].clks);
+	clk_bulk_disable_unprepare(rdev->cores[core].soc->num_clks, rdev->cores[core].clks);
 
 	return 0;
 }
diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocket_job.c
index bb77b6bf0..aa26e2977 100644
--- a/drivers/accel/rocket/rocket_job.c
+++ b/drivers/accel/rocket/rocket_job.c
@@ -345,10 +345,15 @@ static void rocket_job_handle_irq(struct rocket_core *core)
 {
 	pm_runtime_mark_last_busy(core->dev);
 
-	rocket_pc_writel(core, OPERATION_ENABLE, 0x0);
-	rocket_pc_writel(core, INTERRUPT_CLEAR, 0x1ffff);
+	scoped_guard(mutex, &core->job_lock) {
+		/*
+		 * Stopping the block belongs under the lock. hw_submit() writes
+		 * OPERATION_ENABLE too, and outside the lock this zero can land
+		 * after that one and kill a task that has only just started.
+		 */
+		rocket_pc_writel(core, OPERATION_ENABLE, 0x0);
+		rocket_pc_writel(core, INTERRUPT_CLEAR, 0x1ffff);
 
-	scoped_guard(mutex, &core->job_lock)
 		if (core->in_flight_job) {
 			if (core->in_flight_job->next_task_idx < core->in_flight_job->task_count) {
 				rocket_job_hw_submit(core, core->in_flight_job);
@@ -360,6 +365,7 @@ static void rocket_job_handle_irq(struct rocket_core *core)
 			pm_runtime_put_autosuspend(core->dev);
 			core->in_flight_job = NULL;
 		}
+	}
 }
 
 static void
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [RFC PATCH v6 7/9] accel/rocket: add RK3576 NPU (RKNN) support
  2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
                   ` (5 preceding siblings ...)
  2026-08-06  6:34 ` [RFC PATCH v6 6/9] accel/rocket: select the per-core clock and reset counts from match data Jiaxing Hu
@ 2026-08-06  6:34 ` Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 8/9] arm64: dts: rockchip: rk3576: add NPU (RKNN) nodes Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 9/9] arm64: dts: rockchip: rk3576-rock-4d: enable NPU Jiaxing Hu
  8 siblings, 0 replies; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

The RK3576 has two cores of the same RKNN block and a few platform
differences:

 - the CBUF (convolution buffer) has its own clock domain, so the core
   needs six clocks rather than four;
 - the BIU reset moved into the power domain, leaving one reset here;
 - the NPU spans two power domains, and a device with more than one is
   skipped by the driver-core single-domain auto-attach, so the list has
   to be attached explicitly;
 - the DPU completion interrupt is armed exactly as on RK3588 but never
   reaches the GIC. The completion is visible in INTERRUPT_RAW_STATUS,
   so sample that from an hrtimer rather than wait for an interrupt that
   does not come. The interrupt stays armed, and if it ever does arrive
   the two paths agree on which submit each completion belongs to.

All of it hangs off the soc_data added earlier, so the RK3588 path keeps
its existing counts and behaviour.

Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
 drivers/accel/rocket/rocket_core.c   |  18 ++++
 drivers/accel/rocket/rocket_core.h   |  15 ++-
 drivers/accel/rocket/rocket_device.c |   4 +
 drivers/accel/rocket/rocket_drv.c    |  10 ++
 drivers/accel/rocket/rocket_job.c    | 142 ++++++++++++++++++++++++---
 5 files changed, 174 insertions(+), 15 deletions(-)

diff --git a/drivers/accel/rocket/rocket_core.c b/drivers/accel/rocket/rocket_core.c
index b202d1581..1c865e247 100644
--- a/drivers/accel/rocket/rocket_core.c
+++ b/drivers/accel/rocket/rocket_core.c
@@ -8,6 +8,7 @@
 #include <linux/err.h>
 #include <linux/iommu.h>
 #include <linux/platform_device.h>
+#include <linux/pm_domain.h>
 #include <linux/pm_runtime.h>
 #include <linux/reset.h>
 
@@ -21,6 +22,7 @@ int rocket_core_init(struct rocket_core *core)
 	u32 version;
 	int err = 0;
 
+	/* RK3576 moves the BIU reset into its power domain and takes only srst_a. */
 	core->resets[0].id = "srst_a";
 	core->resets[1].id = "srst_h";
 	err = devm_reset_control_bulk_get_exclusive(&pdev->dev, core->soc->num_resets,
@@ -32,6 +34,9 @@ int rocket_core_init(struct rocket_core *core)
 	core->clks[1].id = "hclk";
 	core->clks[2].id = "npu";
 	core->clks[3].id = "pclk";
+	/* RK3576 clocks the CBUF separately; the compute path stalls without these. */
+	core->clks[4].id = "aclk_cbuf";
+	core->clks[5].id = "hclk_cbuf";
 	err = devm_clk_bulk_get(dev, core->soc->num_clks, core->clks);
 	if (err)
 		return dev_err_probe(dev, err, "failed to get clocks for core %d\n", core->index);
@@ -69,6 +74,19 @@ int rocket_core_init(struct rocket_core *core)
 		return err;
 	}
 
+	/*
+	 * RK3576 spans two power domains, and a multi-domain device is skipped
+	 * by the driver-core single-domain auto-attach, so attach the list here.
+	 */
+	if (core->soc->multi_power_domain) {
+		struct dev_pm_domain_list *pd_list;
+
+		err = devm_pm_domain_attach_list(dev, NULL, &pd_list);
+		if (err < 0)
+			return dev_err_probe(dev, err,
+					     "failed to attach NPU power domains\n");
+	}
+
 	pm_runtime_use_autosuspend(dev);
 
 	/*
diff --git a/drivers/accel/rocket/rocket_core.h b/drivers/accel/rocket/rocket_core.h
index 0f424bb86..205ff070d 100644
--- a/drivers/accel/rocket/rocket_core.h
+++ b/drivers/accel/rocket/rocket_core.h
@@ -6,6 +6,7 @@
 
 #include <drm/gpu_scheduler.h>
 #include <linux/clk.h>
+#include <linux/hrtimer.h>
 #include <linux/io.h>
 #include <linux/mutex_types.h>
 #include <linux/reset.h>
@@ -29,8 +30,10 @@
 
 /* Per-SoC differences, selected by the of_device_id match data. */
 struct rocket_soc_data {
-	unsigned int num_clks;		/* clk_bulk count */
-	unsigned int num_resets;	/* reset_bulk count */
+	unsigned int num_clks;		/* clk_bulk count: 4 base, 6 with CBUF */
+	unsigned int num_resets;	/* reset_bulk count: 2 base, 1 on RK3576 */
+	bool multi_power_domain;	/* device spans more than one PM domain */
+	bool poll_completion;		/* completion IRQ never reaches the GIC */
 };
 
 struct rocket_core {
@@ -59,6 +62,14 @@ struct rocket_core {
 		atomic_t pending;
 	} reset;
 
+	struct hrtimer poll_timer;
+	struct work_struct poll_work;
+	atomic_t poll_active;
+	unsigned int poll_ticks;
+	unsigned int poll_seq;
+	unsigned int poll_work_seq;
+	bool poll_dying;
+
 	struct drm_gpu_scheduler sched;
 	u64 fence_context;
 	u64 emit_seqno;
diff --git a/drivers/accel/rocket/rocket_device.c b/drivers/accel/rocket/rocket_device.c
index 46e6ee1e7..bfb00f967 100644
--- a/drivers/accel/rocket/rocket_device.c
+++ b/drivers/accel/rocket/rocket_device.c
@@ -31,6 +31,10 @@ struct rocket_device *rocket_device_init(struct platform_device *pdev,
 		if (of_device_is_available(core_node))
 			num_cores++;
 
+	for_each_compatible_node(core_node, NULL, "rockchip,rk3576-rknn-core")
+		if (of_device_is_available(core_node))
+			num_cores++;
+
 	rdev->cores = devm_kcalloc(dev, num_cores, sizeof(*rdev->cores), GFP_KERNEL);
 	if (!rdev->cores)
 		return ERR_PTR(-ENOMEM);
diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocket_drv.c
index 6e7dc91c5..7f7dfa374 100644
--- a/drivers/accel/rocket/rocket_drv.c
+++ b/drivers/accel/rocket/rocket_drv.c
@@ -217,10 +217,20 @@ static void rocket_remove(struct platform_device *pdev)
 static const struct rocket_soc_data rk3588_soc_data = {
 	.num_clks = 4,
 	.num_resets = 2,
+	.multi_power_domain = false,
+	.poll_completion = false,
+};
+
+static const struct rocket_soc_data rk3576_soc_data = {
+	.num_clks = 6,
+	.num_resets = 1,
+	.multi_power_domain = true,
+	.poll_completion = true,
 };
 
 static const struct of_device_id dt_match[] = {
 	{ .compatible = "rockchip,rk3588-rknn-core", .data = &rk3588_soc_data },
+	{ .compatible = "rockchip,rk3576-rknn-core", .data = &rk3576_soc_data },
 	{}
 };
 MODULE_DEVICE_TABLE(of, dt_match);
diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocket_job.c
index aa26e2977..c5f91f4c5 100644
--- a/drivers/accel/rocket/rocket_job.c
+++ b/drivers/accel/rocket/rocket_job.c
@@ -7,6 +7,7 @@
 #include <drm/drm_file.h>
 #include <drm/drm_gem.h>
 #include <drm/rocket_accel.h>
+#include <linux/hrtimer.h>
 #include <linux/interrupt.h>
 #include <linux/overflow.h>
 #include <linux/iommu.h>
@@ -21,6 +22,15 @@
 
 #define JOB_TIMEOUT_MS 500
 
+/*
+ * RK3576 arms the same DPU completion as RK3588, but the interrupt never
+ * reaches the GIC. The completion itself is visible in INTERRUPT_RAW_STATUS,
+ * so sample that instead. The tick cap bounds jobs that never raise it at all,
+ * which is the same open problem as the wrong inference results.
+ */
+#define RK3576_POLL_INTERVAL_NS	1000000LL	/* 1 ms */
+#define RK3576_POLL_MAX_TICKS	8
+
 static struct rocket_job *
 to_rocket_job(struct drm_sched_job *sched_job)
 {
@@ -151,6 +161,14 @@ static void rocket_job_hw_submit(struct rocket_core *core, struct rocket_job *jo
 
 	rocket_pc_writel(core, OPERATION_ENABLE, PC_OPERATION_ENABLE_OP_EN(1));
 
+	if (core->soc->poll_completion) {
+		core->poll_ticks = 0;
+		WRITE_ONCE(core->poll_seq, core->poll_seq + 1);
+		atomic_set(&core->poll_active, 1);
+		hrtimer_start(&core->poll_timer, ns_to_ktime(RK3576_POLL_INTERVAL_NS),
+			      HRTIMER_MODE_REL);
+	}
+
 	dev_dbg(core->dev, "Submitted regcmd at 0x%llx to core %d", task->regcmd, core->index);
 }
 
@@ -341,30 +359,108 @@ static struct dma_fence *rocket_job_run(struct drm_sched_job *sched_job)
 	return ERR_PTR(ret);
 }
 
+static enum hrtimer_restart rocket_poll_timer_fn(struct hrtimer *timer)
+{
+	struct rocket_core *core = container_of(timer, struct rocket_core, poll_timer);
+	u32 raw;
+
+	if (!atomic_read(&core->poll_active))
+		return HRTIMER_NORESTART;
+
+	WRITE_ONCE(core->poll_work_seq, READ_ONCE(core->poll_seq));
+
+	raw = rocket_pc_readl(core, INTERRUPT_RAW_STATUS);
+	if ((raw & (PC_INTERRUPT_RAW_STATUS_DPU_0 | PC_INTERRUPT_RAW_STATUS_DPU_1)) ||
+	    ++core->poll_ticks >= RK3576_POLL_MAX_TICKS) {
+		atomic_set(&core->poll_active, 0);
+		schedule_work(&core->poll_work);
+		return HRTIMER_NORESTART;
+	}
+
+	hrtimer_forward_now(timer, ns_to_ktime(RK3576_POLL_INTERVAL_NS));
+	return HRTIMER_RESTART;
+}
+
+/* Start the job's next task, or retire it. Caller holds job_lock. */
+static void rocket_job_next_locked(struct rocket_core *core)
+{
+	lockdep_assert_held(&core->job_lock);
+
+	if (!core->in_flight_job)
+		return;
+
+	if (core->in_flight_job->next_task_idx < core->in_flight_job->task_count) {
+		rocket_job_hw_submit(core, core->in_flight_job);
+		return;
+	}
+
+	iommu_detach_group(NULL, iommu_group_get(core->dev));
+	dma_fence_signal(core->in_flight_job->done_fence);
+	pm_runtime_put_autosuspend(core->dev);
+	core->in_flight_job = NULL;
+}
+
+static void rocket_poll_work_fn(struct work_struct *work)
+{
+	struct rocket_core *core = container_of(work, struct rocket_core, poll_work);
+
+	pm_runtime_mark_last_busy(core->dev);
+
+	scoped_guard(mutex, &core->job_lock) {
+		/*
+		 * The interrupt can land while this work is queued, retire the job
+		 * and start the next task. poll_seq only advances under job_lock,
+		 * in hw_submit, so comparing it here says whether that happened.
+		 */
+		if (READ_ONCE(core->poll_dying) ||
+		    READ_ONCE(core->poll_work_seq) != core->poll_seq)
+			return;
+
+		/*
+		 * Unlike an interrupt there is no hardware condition to ack here,
+		 * so with no job in flight there is nothing to write, and the
+		 * device may have autosuspended underneath this work already.
+		 */
+		if (!core->in_flight_job)
+			return;
+
+		rocket_pc_writel(core, OPERATION_ENABLE, 0x0);
+		rocket_pc_writel(core, INTERRUPT_CLEAR, 0x1ffff);
+
+		rocket_job_next_locked(core);
+	}
+}
+
 static void rocket_job_handle_irq(struct rocket_core *core)
 {
+	unsigned int seq = 0;
+
+	if (core->soc->poll_completion) {
+		seq = READ_ONCE(core->poll_seq);
+		atomic_set(&core->poll_active, 0);
+		hrtimer_cancel(&core->poll_timer);
+	}
+
 	pm_runtime_mark_last_busy(core->dev);
 
 	scoped_guard(mutex, &core->job_lock) {
 		/*
-		 * Stopping the block belongs under the lock. hw_submit() writes
-		 * OPERATION_ENABLE too, and outside the lock this zero can land
-		 * after that one and kill a task that has only just started.
+		 * Stopping the block belongs under the lock. A submit from the
+		 * other completion path writes OPERATION_ENABLE too, and outside
+		 * the lock this zero can land after that one and kill a task that
+		 * has only just started.
 		 */
 		rocket_pc_writel(core, OPERATION_ENABLE, 0x0);
 		rocket_pc_writel(core, INTERRUPT_CLEAR, 0x1ffff);
 
-		if (core->in_flight_job) {
-			if (core->in_flight_job->next_task_idx < core->in_flight_job->task_count) {
-				rocket_job_hw_submit(core, core->in_flight_job);
-				return;
-			}
+		/*
+		 * Where both completion paths are live, one whose submit the other
+		 * has already retired must not go on to start a further task.
+		 */
+		if (core->soc->poll_completion && seq != core->poll_seq)
+			return;
 
-			iommu_detach_group(NULL, iommu_group_get(core->dev));
-			dma_fence_signal(core->in_flight_job->done_fence);
-			pm_runtime_put_autosuspend(core->dev);
-			core->in_flight_job = NULL;
-		}
+		rocket_job_next_locked(core);
 	}
 }
 
@@ -466,6 +562,11 @@ int rocket_job_init(struct rocket_core *core)
 	int ret;
 
 	INIT_WORK(&core->reset.work, rocket_reset_work);
+	INIT_WORK(&core->poll_work, rocket_poll_work_fn);
+	hrtimer_setup(&core->poll_timer, rocket_poll_timer_fn, CLOCK_MONOTONIC,
+		      HRTIMER_MODE_REL);
+	atomic_set(&core->poll_active, 0);
+	core->poll_dying = false;
 	spin_lock_init(&core->fence_lock);
 	mutex_init(&core->job_lock);
 
@@ -507,8 +608,23 @@ int rocket_job_init(struct rocket_core *core)
 
 void rocket_job_fini(struct rocket_core *core)
 {
+	/*
+	 * Stop the poll from starting hardware work before tearing anything
+	 * down: it submits the next task, and drm_sched_fini() does not wait
+	 * for work already queued. Cancel after the scheduler is gone, so a
+	 * job running now cannot re-arm the timer behind the cancel.
+	 */
+	if (core->soc->poll_completion)
+		WRITE_ONCE(core->poll_dying, true);
+
 	drm_sched_fini(&core->sched);
 
+	if (core->soc->poll_completion) {
+		atomic_set(&core->poll_active, 0);
+		hrtimer_cancel(&core->poll_timer);
+		cancel_work_sync(&core->poll_work);
+	}
+
 	cancel_work_sync(&core->reset.work);
 	destroy_workqueue(core->reset.wq);
 }
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [RFC PATCH v6 8/9] arm64: dts: rockchip: rk3576: add NPU (RKNN) nodes
  2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
                   ` (6 preceding siblings ...)
  2026-08-06  6:34 ` [RFC PATCH v6 7/9] accel/rocket: add RK3576 NPU (RKNN) support Jiaxing Hu
@ 2026-08-06  6:34 ` Jiaxing Hu
  2026-08-06  6:34 ` [RFC PATCH v6 9/9] arm64: dts: rockchip: rk3576-rock-4d: enable NPU Jiaxing Hu
  8 siblings, 0 replies; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

Add the two RKNN cores and their IOMMUs, plus the NPU power-domain
resets the pmdomain driver now cycles on power-on. Both cores are
disabled by default; boards enable what they wire up.

Each core lists both NPU power domains, its own first. The compute path
needs NPU1 powered even when only core 0 runs, and a node with a single
domain would be auto-attached by the driver core before the driver can
attach the list itself. The IOMMUs keep one domain each, since they rely
on that same auto-attach.

Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
 arch/arm64/boot/dts/rockchip/rk3576.dtsi | 80 +++++++++++++++++++++++-
 1 file changed, 78 insertions(+), 2 deletions(-)

diff --git a/arch/arm64/boot/dts/rockchip/rk3576.dtsi b/arch/arm64/boot/dts/rockchip/rk3576.dtsi
index b0c0d3c8b..1e6dd039f 100644
--- a/arch/arm64/boot/dts/rockchip/rk3576.dtsi
+++ b/arch/arm64/boot/dts/rockchip/rk3576.dtsi
@@ -1070,14 +1070,22 @@ power-domain@RK3576_PD_NPUTOP {
 						power-domain@RK3576_PD_NPU0 {
 							reg = <RK3576_PD_NPU0>;
 							clocks = <&cru HCLK_RKNN_ROOT>,
-								 <&cru ACLK_RKNN0>;
+								 <&cru ACLK_RKNN0>,
+								 <&cru CLK_RKNN_DSU0>,
+								 <&cru ACLK_RKNN_CBUF>,
+								 <&cru HCLK_RKNN_CBUF>;
+							resets = <&cru SRST_A_RKNN0_BIU>;
 							pm_qos = <&qos_npu_m0>;
 							#power-domain-cells = <0>;
 						};
 						power-domain@RK3576_PD_NPU1 {
 							reg = <RK3576_PD_NPU1>;
 							clocks = <&cru HCLK_RKNN_ROOT>,
-								 <&cru ACLK_RKNN1>;
+								 <&cru ACLK_RKNN1>,
+								 <&cru CLK_RKNN_DSU0>,
+								 <&cru ACLK_RKNN_CBUF>,
+								 <&cru HCLK_RKNN_CBUF>;
+							resets = <&cru SRST_A_RKNN1_BIU>;
 							pm_qos = <&qos_npu_m1>;
 							#power-domain-cells = <0>;
 						};
@@ -1832,6 +1840,74 @@ qos_npu_m1ro: qos@27f22100 {
 			reg = <0x0 0x27f22100 0x0 0x20>;
 		};
 
+		rknn_core_0: npu@27700000 {
+			compatible = "rockchip,rk3576-rknn-core";
+			reg = <0x0 0x27700000 0x0 0x1000>,
+			      <0x0 0x27701000 0x0 0x1000>,
+			      <0x0 0x27703000 0x0 0x1000>;
+			reg-names = "pc", "cna", "core";
+			interrupts = <GIC_SPI 247 IRQ_TYPE_LEVEL_HIGH>;
+			clocks = <&cru ACLK_RKNN0>, <&cru HCLK_RKNN_ROOT>,
+				 <&cru CLK_RKNN_DSU0>, <&cru PCLK_NPUTOP_ROOT>,
+				 <&cru ACLK_RKNN_CBUF>, <&cru HCLK_RKNN_CBUF>;
+			clock-names = "aclk", "hclk", "npu", "pclk",
+				      "aclk_cbuf", "hclk_cbuf";
+			resets = <&cru SRST_A_RKNN0>;
+			reset-names = "srst_a";
+			power-domains = <&power RK3576_PD_NPU0>, <&power RK3576_PD_NPU1>;
+			iommus = <&rknn_mmu_0>;
+			status = "disabled";
+		};
+
+		rknn_mmu_0: iommu@27702000 {
+			compatible = "rockchip,rk3576-iommu", "rockchip,rk3568-iommu";
+			reg = <0x0 0x27702000 0x0 0x100>,
+			      <0x0 0x27702100 0x0 0x100>;
+			interrupts = <GIC_SPI 247 IRQ_TYPE_LEVEL_HIGH>;
+			clocks = <&cru ACLK_RKNN0>, <&cru HCLK_RKNN_ROOT>,
+				 <&cru CLK_RKNN_DSU0>, <&cru ACLK_RKNN_CBUF>,
+				 <&cru HCLK_RKNN_CBUF>;
+			clock-names = "aclk", "iface", "npu",
+				      "aclk_cbuf", "hclk_cbuf";
+			#iommu-cells = <0>;
+			power-domains = <&power RK3576_PD_NPU0>;
+			status = "disabled";
+		};
+
+		rknn_core_1: npu@27708000 {
+			compatible = "rockchip,rk3576-rknn-core";
+			reg = <0x0 0x27708000 0x0 0x1000>,
+			      <0x0 0x27709000 0x0 0x1000>,
+			      <0x0 0x2770b000 0x0 0x1000>;
+			reg-names = "pc", "cna", "core";
+			interrupts = <GIC_SPI 248 IRQ_TYPE_LEVEL_HIGH>;
+			clocks = <&cru ACLK_RKNN1>, <&cru HCLK_RKNN_ROOT>,
+				 <&cru CLK_RKNN_DSU0>, <&cru PCLK_NPUTOP_ROOT>,
+				 <&cru ACLK_RKNN_CBUF>, <&cru HCLK_RKNN_CBUF>;
+			clock-names = "aclk", "hclk", "npu", "pclk",
+				      "aclk_cbuf", "hclk_cbuf";
+			resets = <&cru SRST_A_RKNN1>;
+			reset-names = "srst_a";
+			power-domains = <&power RK3576_PD_NPU1>, <&power RK3576_PD_NPU0>;
+			iommus = <&rknn_mmu_1>;
+			status = "disabled";
+		};
+
+		rknn_mmu_1: iommu@2770a000 {
+			compatible = "rockchip,rk3576-iommu", "rockchip,rk3568-iommu";
+			reg = <0x0 0x2770a000 0x0 0x100>,
+			      <0x0 0x2770a100 0x0 0x100>;
+			interrupts = <GIC_SPI 248 IRQ_TYPE_LEVEL_HIGH>;
+			clocks = <&cru ACLK_RKNN1>, <&cru HCLK_RKNN_ROOT>,
+				 <&cru CLK_RKNN_DSU0>, <&cru ACLK_RKNN_CBUF>,
+				 <&cru HCLK_RKNN_CBUF>;
+			clock-names = "aclk", "iface", "npu",
+				      "aclk_cbuf", "hclk_cbuf";
+			#iommu-cells = <0>;
+			power-domains = <&power RK3576_PD_NPU1>;
+			status = "disabled";
+		};
+
 		gmac0: ethernet@2a220000 {
 			compatible = "rockchip,rk3576-gmac", "snps,dwmac-4.20a";
 			reg = <0x0 0x2a220000 0x0 0x10000>;
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [RFC PATCH v6 9/9] arm64: dts: rockchip: rk3576-rock-4d: enable NPU
  2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
                   ` (7 preceding siblings ...)
  2026-08-06  6:34 ` [RFC PATCH v6 8/9] arm64: dts: rockchip: rk3576: add NPU (RKNN) nodes Jiaxing Hu
@ 2026-08-06  6:34 ` Jiaxing Hu
  8 siblings, 0 replies; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  6:34 UTC (permalink / raw)
  To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
	ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel, Jiaxing Hu

Enable rknn_core_0 and its IOMMU on the Radxa ROCK 4D and supply the
core from vdd_npu_s0.

The supply is marked always-on because the NPU power domains are what
gate the block here, and dropping the rail underneath them takes an
async SError on the next power-on rather than a clean retry. Only
rknn_core_0 is enabled: the driver binds one core per node and the
second core is left to whoever can test it.

Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
 arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts b/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts
index 272af1012..965e0906b 100644
--- a/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts
+++ b/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts
@@ -442,6 +442,7 @@ regulator-state-mem {
 			};
 
 			vdd_npu_s0: dcdc-reg2 {
+				regulator-always-on;
 				regulator-boot-on;
 				regulator-enable-ramp-delay = <400>;
 				regulator-min-microvolt = <550000>;
@@ -869,3 +870,12 @@ vp0_out_hdmi: endpoint@ROCKCHIP_VOP2_EP_HDMI0 {
 		remote-endpoint = <&hdmi_in_vp0>;
 	};
 };
+
+&rknn_core_0 {
+	npu-supply = <&vdd_npu_s0>;
+	status = "okay";
+};
+
+&rknn_mmu_0 {
+	status = "okay";
+};
-- 
2.43.0



^ permalink raw reply related	[flat|nested] 13+ messages in thread

* Re: [RFC PATCH v6 3/9] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set
  2026-08-06  6:34 ` [RFC PATCH v6 3/9] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set Jiaxing Hu
@ 2026-08-06  9:23   ` Diederik de Haas
  2026-08-06  9:55     ` Jiaxing Hu
  0 siblings, 1 reply; 13+ messages in thread
From: Diederik de Haas @ 2026-08-06  9:23 UTC (permalink / raw)
  To: Jiaxing Hu, tomeu, heiko, robh, krzk+dt, conor+dt, joro, will,
	robin.murphy, ulfh, p.zabel, ogabbay, zhangqing
  Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
	linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
	linux-kernel

On Thu Aug 6, 2026 at 8:34 AM CEST, Jiaxing Hu wrote:
> The RK3576 NPU MMUs need more than aclk and iface. With only those two
> enabled the MMU accepts reads but silently drops register writes: a
> DTE_ADDR value written from the power domain, while the domain clocks
> are still on, reads back correctly, and the write rk_iommu_resume() does
> microseconds later does not land at all. The vendor DT names the CBUF
> clocks as that MMU's interface clocks and its driver keeps every NPU
> clock on for as long as the device is powered.
>
> The driver side of this is already upstream, commit 841363ebb508
> ("iommu/rockchip: Take all DT clocks"), which switched rk_iommu to
> devm_clk_bulk_get_all(). Widen the schema to match so those nodes can
> be described. minItems stays at 2, so every existing devicetree, which
> all carry exactly aclk and iface, is unaffected.

I agree with all remarks Sashiko made wrt this patch:
https://sashiko.dev/#/patchset/20260805063826.95682-1-gahing%40gahingwoo.com?part=3

Now every MMU with "rockchip,rk3568-iommu", "rockchip,rk3588-iommu" or
"rockchip,rk3576-iommu" is allowed to have a minimum of 2 clocks, instead
of having exactly 2 clocks. That does not sound desirable.

Cheers,
  Diederik
>
> Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
> ---
>  .../devicetree/bindings/iommu/rockchip,iommu.yaml         | 8 ++++++++
>  1 file changed, 8 insertions(+)
>
> diff --git a/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml b/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
> index 6ce41d11f..a3cedcaaa 100644
> --- a/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
> +++ b/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
> @@ -42,14 +42,22 @@ properties:
>      minItems: 1
>  
>    clocks:
> +    minItems: 2
>      items:
>        - description: Core clock
>        - description: Interface clock
> +      - description: Compute clock, RK3576 NPU MMUs only
> +      - description: Convolution buffer core clock, RK3576 NPU MMUs only
> +      - description: Convolution buffer interface clock, RK3576 NPU MMUs only
>  
>    clock-names:
> +    minItems: 2
>      items:
>        - const: aclk
>        - const: iface
> +      - const: npu
> +      - const: aclk_cbuf
> +      - const: hclk_cbuf
>  
>    "#iommu-cells":
>      const: 0




^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [RFC PATCH v6 3/9] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set
  2026-08-06  9:23   ` Diederik de Haas
@ 2026-08-06  9:55     ` Jiaxing Hu
  2026-08-06 11:29       ` Diederik de Haas
  0 siblings, 1 reply; 13+ messages in thread
From: Jiaxing Hu @ 2026-08-06  9:55 UTC (permalink / raw)
  To: diederik, heiko, robh, krzk+dt, conor+dt, joro, will,
	robin.murphy, tomeu
  Cc: royalnet026, alchark, chaoyi.chen, iommu, linux-rockchip,
	devicetree, dri-devel, linux-arm-kernel, linux-kernel, Jiaxing Hu

Hi Diederik,

> Now every MMU with "rockchip,rk3568-iommu", "rockchip,rk3588-iommu" or
> "rockchip,rk3576-iommu" is allowed to have a minimum of 2 clocks, instead
> of having exactly 2 clocks. That does not sound desirable.

You are right, and it is worse than sounding undesirable, it actually
happens. I gave an RK3588 NPU MMU a bogus third clock and v6's schema
accepted it without a word. That is a real loss of coverage for every
existing Rockchip IOMMU and I should not have sent it that way.

Fixed for v7 the way you and the bot suggest, with a compatible of its
own:

  compatible = "rockchip,rk3576-npu-iommu", "rockchip,rk3568-iommu";

and an allOf that pins each side:

  if compatible contains rockchip,rk3576-npu-iommu
    then clocks/clock-names minItems: 5
    else clocks/clock-names maxItems: 2

so the NPU MMUs are required to carry all five and everything else is
back to exactly two. Checked in both directions: the three clock RK3588
node is rejected again, an NPU MMU with only aclk and iface is rejected,
and every rockchip dtb in the tree validates clean.

No driver change goes with it. rk_iommu matches only "rockchip,iommu"
and "rockchip,rk3568-iommu", and the fallback stays, so the new string
is documentation only.

The name is the part I am least sure of. Everything else in that binding
ends in -iommu, which is why I did not use -mmu to pair with the
rknn-core node it belongs to. Happy to change it if Heiko or Krzysztof
prefer something else.

Thanks for catching it.

Cheers,
Jiaxing


^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [RFC PATCH v6 3/9] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set
  2026-08-06  9:55     ` Jiaxing Hu
@ 2026-08-06 11:29       ` Diederik de Haas
  0 siblings, 0 replies; 13+ messages in thread
From: Diederik de Haas @ 2026-08-06 11:29 UTC (permalink / raw)
  To: Jiaxing Hu, diederik, heiko, robh, krzk+dt, conor+dt, joro, will,
	robin.murphy, tomeu
  Cc: royalnet026, alchark, chaoyi.chen, iommu, linux-rockchip,
	devicetree, dri-devel, linux-arm-kernel, linux-kernel

Hi Jiaxing,

On Thu Aug 6, 2026 at 11:55 AM CEST, Jiaxing Hu wrote:
> Hi Diederik,
>
>> Now every MMU with "rockchip,rk3568-iommu", "rockchip,rk3588-iommu" or
>> "rockchip,rk3576-iommu" is allowed to have a minimum of 2 clocks, instead
>> of having exactly 2 clocks. That does not sound desirable.
>
> You are right, and it is worse than sounding undesirable, it actually

Yeah, it was a 'bit' of an understatement ;-)

> happens. I gave an RK3588 NPU MMU a bogus third clock and v6's schema
> accepted it without a word. That is a real loss of coverage for every
> existing Rockchip IOMMU and I should not have sent it that way.
>
> Fixed for v7 the way you and the bot suggest, with a compatible of its
> own:

If you haven't already, it's probably worth checking whether Sashiko made
other useful remarks. I don't feel qualified to judge those, so I didn't
reference those. But they made be valid as well. Or hallucinations ;-)

>   compatible = "rockchip,rk3576-npu-iommu", "rockchip,rk3568-iommu";
>
> and an allOf that pins each side:
>
>   if compatible contains rockchip,rk3576-npu-iommu
>     then clocks/clock-names minItems: 5
>     else clocks/clock-names maxItems: 2
>
> so the NPU MMUs are required to carry all five and everything else is
> back to exactly two. Checked in both directions: the three clock RK3588
> node is rejected again, an NPU MMU with only aclk and iface is rejected,
> and every rockchip dtb in the tree validates clean.
>
> No driver change goes with it. rk_iommu matches only "rockchip,iommu"
> and "rockchip,rk3568-iommu", and the fallback stays, so the new string
> is documentation only.

I'll leave it up to others to comment whether that's correct or not.

Another thing you could consider is splitting this NPU iommu 'stuff' into
a separate patch set and drop the RFC 'prefix' for that series.
IIUC the RFC is (only) related to the working of the NPU on RK3576.

Cheers,
  Diederik

> The name is the part I am least sure of. Everything else in that binding
> ends in -iommu, which is why I did not use -mmu to pair with the
> rknn-core node it belongs to. Happy to change it if Heiko or Krzysztof
> prefer something else.
>
> Thanks for catching it.
>
> Cheers,
> Jiaxing
>
> _______________________________________________
> Linux-rockchip mailing list
> Linux-rockchip@lists.infradead.org
> http://lists.infradead.org/mailman/listinfo/linux-rockchip




^ permalink raw reply	[flat|nested] 13+ messages in thread

end of thread, other threads:[~2026-08-06 11:29 UTC | newest]

Thread overview: 13+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-06  6:34 [RFC PATCH v6 0/9] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
2026-08-06  6:34 ` [RFC PATCH v6 1/9] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
2026-08-06  6:34 ` [RFC PATCH v6 2/9] dt-bindings: power: rockchip: allow resets in a power domain node Jiaxing Hu
2026-08-06  6:34 ` [RFC PATCH v6 3/9] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set Jiaxing Hu
2026-08-06  9:23   ` Diederik de Haas
2026-08-06  9:55     ` Jiaxing Hu
2026-08-06 11:29       ` Diederik de Haas
2026-08-06  6:34 ` [RFC PATCH v6 4/9] pmdomain/rockchip: add optional per-domain power-on settle delay Jiaxing Hu
2026-08-06  6:34 ` [RFC PATCH v6 5/9] pmdomain/rockchip: cycle optional power-domain resets on power-on Jiaxing Hu
2026-08-06  6:34 ` [RFC PATCH v6 6/9] accel/rocket: select the per-core clock and reset counts from match data Jiaxing Hu
2026-08-06  6:34 ` [RFC PATCH v6 7/9] accel/rocket: add RK3576 NPU (RKNN) support Jiaxing Hu
2026-08-06  6:34 ` [RFC PATCH v6 8/9] arm64: dts: rockchip: rk3576: add NPU (RKNN) nodes Jiaxing Hu
2026-08-06  6:34 ` [RFC PATCH v6 9/9] arm64: dts: rockchip: rk3576-rock-4d: enable NPU Jiaxing Hu

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox