* [PATCH v7 01/10] accel/rocket: take the completion register writes under job_lock
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
@ 2026-08-12 9:40 ` Jiaxing Hu
2026-08-12 12:47 ` Igor Paunovic
2026-08-12 9:40 ` [PATCH v7 02/10] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
` (8 subsequent siblings)
9 siblings, 1 reply; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:40 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
rocket_job_handle_irq() writes OPERATION_ENABLE and INTERRUPT_CLEAR before
taking job_lock, while rocket_job_hw_submit() writes OPERATION_ENABLE from
inside it. The two can therefore race: a completion being handled on one core
can write its zero after a submit on the same core has written its one, and
stop a task that has only just started.
Nothing in tree hits this often, because the interrupt is the only completion
path and it does not overlap its own submit, but the ordering is wrong on its
own terms.
Move both writes inside the existing scoped_guard() rather than adding a second
critical section, so stopping the block and deciding what to start next are one
atomic step.
Fixes: 0810d5ad88a1 ("accel/rocket: Add job submission IOCTL")
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
drivers/accel/rocket/rocket_job.c | 12 +++++++++---
1 file changed, 9 insertions(+), 3 deletions(-)
diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocket_job.c
index bb77b6bf0..4c01b703e 100644
--- a/drivers/accel/rocket/rocket_job.c
+++ b/drivers/accel/rocket/rocket_job.c
@@ -345,10 +345,15 @@ static void rocket_job_handle_irq(struct rocket_core *core)
{
pm_runtime_mark_last_busy(core->dev);
- rocket_pc_writel(core, OPERATION_ENABLE, 0x0);
- rocket_pc_writel(core, INTERRUPT_CLEAR, 0x1ffff);
+ scoped_guard(mutex, &core->job_lock) {
+ /*
+ * Stopping the block belongs under the lock. hw_submit() writes
+ * OPERATION_ENABLE too, and outside the lock this zero can land
+ * after that one and stop a task that has only just started.
+ */
+ rocket_pc_writel(core, OPERATION_ENABLE, 0x0);
+ rocket_pc_writel(core, INTERRUPT_CLEAR, 0x1ffff);
- scoped_guard(mutex, &core->job_lock)
if (core->in_flight_job) {
if (core->in_flight_job->next_task_idx < core->in_flight_job->task_count) {
rocket_job_hw_submit(core, core->in_flight_job);
@@ -360,6 +365,7 @@ static void rocket_job_handle_irq(struct rocket_core *core)
pm_runtime_put_autosuspend(core->dev);
core->in_flight_job = NULL;
}
+ }
}
static void
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v7 01/10] accel/rocket: take the completion register writes under job_lock
2026-08-12 9:40 ` [PATCH v7 01/10] accel/rocket: take the completion register writes under job_lock Jiaxing Hu
@ 2026-08-12 12:47 ` Igor Paunovic
0 siblings, 0 replies; 23+ messages in thread
From: Igor Paunovic @ 2026-08-12 12:47 UTC (permalink / raw)
To: Jiaxing Hu, tomeu, heiko, robh, krzk+dt, conor+dt, joro, will,
robin.murphy, ulfh, p.zabel, ogabbay, zhangqing
Cc: Igor Paunovic, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel
Tested-by: Igor Paunovic <royalnet026@gmail.com> # RK3588, three cores
I ran this on an Orange Pi 5 Plus across all three NPU cores, against a
base without the series. Both modules were built the same way and
neither carried any local DVFS work.
base: v7.2 rocket
+ Guangshuo Li's "clear rdev on device init failure"
+ my "request the core clocks by name" v2
+ my lifecycle v2 1/2 and 2/2
test: the same, plus 1/10, 7/10 and 8/10 from this series
Six phases per module: all three cores bound; core 2 unbound and
rebound; core 0 unbound and rebound; all three unbound and all three
rebound. One MobileNet V1 run per phase through the Teflon delegate.
The oracle is the sha256 of the tensors that both change between
different inputs and stay stable across repeats, so a stale output
buffer cannot pass as a recomputation.
base this series
three cores 89.3 90.0 inf/s
core 2 unbound 87.6 88.8
core 2 rebound 88.9 88.2
core 0 unbound 75.3 75.0
core 0 rebound 88.6 88.4
all three cycled 88.7 88.3
All twelve runs produce identical oracle hashes and the same
classification. Interrupts per inference are 42.75 in both, and the
distribution matches phase for phase: with core 0 bound it takes 41.7
of them and core 1 takes 1.02; with core 0 unbound the same work moves
to core 1. Neither round logged anything beyond the probe messages.
That comes to 2596 inferences and 111048 completion interrupts through
rocket_job_handle_irq() with the two writes moved under job_lock, with
no difference in result from the same count without them.
On the change itself: I could not construct the race on the normal
path. The scheduler runs one job at a time and the fence is signalled
under the same lock after the writes, so a submit cannot overlap the
completion it follows.
Where I think it is reachable is the reset path. rocket_reset() calls
drm_sched_stop() and then says "Remaining interrupts have been
handled", but drm_sched_stop() stops the scheduler, not the threaded
IRQ handler. A handler already in flight can therefore run alongside
rocket_reset(), and after drm_sched_start() alongside a fresh job.
Making the write and the decision one step is the right shape for
that. It does not stop a late handler from writing the zero into a job
that is not the one whose interrupt it is handling, though - would a
synchronize_irq(core->irq) before the guard in rocket_reset() be worth
having as well?
One note on the base, since it matters to anyone repeating this. The
core-0 rebind step needs my lifecycle series underneath. Without it,
that rebind hands the returning core the index of a core that is still
live: the driver prints "core 2" for fdab0000.npu, inference starts
returning a different answer, and the teardown that follows dies in
destroy_workqueue() under drm_sched_fini() with a poisoned list
pointer, leaving an unkillable D state. None of that is your series
doing - it reproduces with 1/10, 7/10 and 8/10 absent - but it does
mean the three-core test cannot run to completion on a tree without it.
Igor
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v7 02/10] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
2026-08-12 9:40 ` [PATCH v7 01/10] accel/rocket: take the completion register writes under job_lock Jiaxing Hu
@ 2026-08-12 9:40 ` Jiaxing Hu
2026-08-13 7:04 ` Krzysztof Kozlowski
2026-08-12 9:40 ` [PATCH v7 03/10] dt-bindings: power: rockchip: allow resets in a power domain node Jiaxing Hu
` (7 subsequent siblings)
9 siblings, 1 reply; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:40 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
The RK3576 NPU has two cores of the same RKNN block the RK3588 binding
already describes, but it wires them up differently: two extra CBUF
clocks, two power domains per core, and a single reset instead of two.
It also has no NPU SRAM supply.
Widen the property ranges to cover both, then pin each SoC back to its
own shape in allOf so nothing loosens for RK3588, and keep sram-supply
required for rockchip,rk3588-rknn-core only.
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
.../npu/rockchip,rk3588-rknn-core.yaml | 47 +++++++++++++++++--
1 file changed, 44 insertions(+), 3 deletions(-)
diff --git a/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml b/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml
index caca2a490..3b611b64c 100644
--- a/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml
+++ b/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml
@@ -21,6 +21,7 @@ properties:
compatible:
enum:
+ - rockchip,rk3576-rknn-core
- rockchip,rk3588-rknn-core
reg:
@@ -33,14 +34,18 @@ properties:
- const: core # Main NPU core processing unit registers
clocks:
- maxItems: 4
+ minItems: 4
+ maxItems: 6
clock-names:
+ minItems: 4
items:
- const: aclk
- const: hclk
- const: npu
- const: pclk
+ - const: aclk_cbuf
+ - const: hclk_cbuf
interrupts:
maxItems: 1
@@ -51,12 +56,15 @@ properties:
npu-supply: true
power-domains:
- maxItems: 1
+ minItems: 1
+ maxItems: 2
resets:
+ minItems: 1
maxItems: 2
reset-names:
+ minItems: 1
items:
- const: srst_a
- const: srst_h
@@ -75,7 +83,40 @@ required:
- resets
- reset-names
- npu-supply
- - sram-supply
+
+allOf:
+ - if:
+ properties:
+ compatible:
+ contains:
+ const: rockchip,rk3588-rknn-core
+ then:
+ properties:
+ clocks:
+ maxItems: 4
+ clock-names:
+ maxItems: 4
+ power-domains:
+ maxItems: 1
+ resets:
+ minItems: 2
+ reset-names:
+ minItems: 2
+ required:
+ - sram-supply
+ else:
+ properties:
+ clocks:
+ minItems: 6
+ clock-names:
+ minItems: 6
+ power-domains:
+ minItems: 2
+ resets:
+ maxItems: 1
+ reset-names:
+ maxItems: 1
+ sram-supply: false
additionalProperties: false
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v7 02/10] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core
2026-08-12 9:40 ` [PATCH v7 02/10] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
@ 2026-08-13 7:04 ` Krzysztof Kozlowski
0 siblings, 0 replies; 23+ messages in thread
From: Krzysztof Kozlowski @ 2026-08-13 7:04 UTC (permalink / raw)
To: Jiaxing Hu
Cc: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing, royalnet026, alchark,
chaoyi.chen, diederik, dri-devel, linux-rockchip, iommu, linux-pm,
devicetree, linux-arm-kernel, linux-kernel
On Wed, Aug 12, 2026 at 09:40:57PM +1200, Jiaxing Hu wrote:
> The RK3576 NPU has two cores of the same RKNN block the RK3588 binding
> already describes, but it wires them up differently: two extra CBUF
> clocks, two power domains per core, and a single reset instead of two.
> It also has no NPU SRAM supply.
>
> Widen the property ranges to cover both, then pin each SoC back to its
> own shape in allOf so nothing loosens for RK3588, and keep sram-supply
> required for rockchip,rk3588-rknn-core only.
>
> Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
> ---
> .../npu/rockchip,rk3588-rknn-core.yaml | 47 +++++++++++++++++--
> 1 file changed, 44 insertions(+), 3 deletions(-)
Reviewed-by: Krzysztof Kozlowski <krzysztof.kozlowski@oss.qualcomm.com>
Best regards,
Krzysztof
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v7 03/10] dt-bindings: power: rockchip: allow resets in a power domain node
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
2026-08-12 9:40 ` [PATCH v7 01/10] accel/rocket: take the completion register writes under job_lock Jiaxing Hu
2026-08-12 9:40 ` [PATCH v7 02/10] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
@ 2026-08-12 9:40 ` Jiaxing Hu
2026-08-13 7:06 ` Krzysztof Kozlowski
2026-08-12 9:40 ` [PATCH v7 04/10] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set Jiaxing Hu
` (6 subsequent siblings)
9 siblings, 1 reply; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:40 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
Some domains do not come up in a usable state on their own and need
their resets cycled once power is on. The RK3576 NPU domains are one
case: without it the first access after power-on takes an async SError.
pd-node has no resets property and every nesting level is
unevaluatedProperties: false, so describing that in DT is rejected
today. Add it alongside clocks.
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
.../bindings/power/rockchip,power-controller.yaml | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml b/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml
index b41db576f..f23c1a118 100644
--- a/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml
+++ b/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml
@@ -136,6 +136,14 @@ $defs:
A number of phandles to clocks that need to be enabled
while power domain switches state.
+ resets:
+ minItems: 1
+ maxItems: 30
+ description: |
+ A number of phandles to resets that need to be cycled once the power
+ domain has been switched on, for domains whose logic does not come up
+ in a usable state by itself.
+
domain-supply:
description: domain regulator supply.
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v7 03/10] dt-bindings: power: rockchip: allow resets in a power domain node
2026-08-12 9:40 ` [PATCH v7 03/10] dt-bindings: power: rockchip: allow resets in a power domain node Jiaxing Hu
@ 2026-08-13 7:06 ` Krzysztof Kozlowski
2026-08-14 8:21 ` Jiaxing Hu
0 siblings, 1 reply; 23+ messages in thread
From: Krzysztof Kozlowski @ 2026-08-13 7:06 UTC (permalink / raw)
To: Jiaxing Hu
Cc: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing, royalnet026, alchark,
chaoyi.chen, diederik, dri-devel, linux-rockchip, iommu, linux-pm,
devicetree, linux-arm-kernel, linux-kernel
On Wed, Aug 12, 2026 at 09:40:58PM +1200, Jiaxing Hu wrote:
> Some domains do not come up in a usable state on their own and need
> their resets cycled once power is on. The RK3576 NPU domains are one
> case: without it the first access after power-on takes an async SError.
>
> pd-node has no resets property and every nesting level is
> unevaluatedProperties: false, so describing that in DT is rejected
> today. Add it alongside clocks.
This paragraph is redundant. Why are you explaining correct syntax?
>
> Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
> ---
> .../bindings/power/rockchip,power-controller.yaml | 8 ++++++++
> 1 file changed, 8 insertions(+)
>
> diff --git a/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml b/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml
> index b41db576f..f23c1a118 100644
> --- a/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml
> +++ b/Documentation/devicetree/bindings/power/rockchip,power-controller.yaml
> @@ -136,6 +136,14 @@ $defs:
> A number of phandles to clocks that need to be enabled
> while power domain switches state.
>
> + resets:
> + minItems: 1
> + maxItems: 30
30 resets per one power domain? and none got to the example in this
file?
> + description: |
Do not need '|' unless you need to preserve formatting.
> + A number of phandles to resets that need to be cycled once the power
> + domain has been switched on, for domains whose logic does not come up
> + in a usable state by itself.
> +
> domain-supply:
> description: domain regulator supply.
>
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v7 03/10] dt-bindings: power: rockchip: allow resets in a power domain node
2026-08-13 7:06 ` Krzysztof Kozlowski
@ 2026-08-14 8:21 ` Jiaxing Hu
0 siblings, 0 replies; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-14 8:21 UTC (permalink / raw)
To: krzk
Cc: heiko, robh, krzk+dt, conor+dt, ulf.hansson, tomeu, royalnet026,
diederik, chaoyi.chen, devicetree, linux-pm, linux-rockchip,
linux-arm-kernel, linux-kernel, Jiaxing Hu
Hi Krzysztof,
> This paragraph is redundant. Why are you explaining correct syntax?
No good reason. v8 drops it.
> 30 resets per one power domain? and none got to the example in this
> file?
One. Both RK3576 NPU domains carry exactly one, SRST_A_RKNN0_BIU and
SRST_A_RKNN1_BIU, and 30 came from the clocks property above it, copied
without asking what it would mean here. v8 has maxItems: 1, which is
what this series actually needs, and it can be widened by whoever turns
up with a domain that needs more.
And no, nothing got to the example, which it should have. v8 adds one.
> Do not need '|' unless you need to preserve formatting.
v8 drops it.
Nothing above is fixed yet, only decided. I told another reviewer on v6
that something was fixed for v7 and then sent v7 without it, so I would
rather say what v8 will contain than describe it as done.
Thanks for the review.
Jiaxing
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v7 04/10] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
` (2 preceding siblings ...)
2026-08-12 9:40 ` [PATCH v7 03/10] dt-bindings: power: rockchip: allow resets in a power domain node Jiaxing Hu
@ 2026-08-12 9:40 ` Jiaxing Hu
2026-08-12 10:45 ` Diederik de Haas
2026-08-12 9:41 ` [PATCH v7 05/10] pmdomain/rockchip: add optional per-domain power-on settle delay Jiaxing Hu
` (5 subsequent siblings)
9 siblings, 1 reply; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:40 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
The RK3576 NPU MMUs need more than aclk and iface. With only those two
enabled the MMU accepts reads but silently drops register writes: a
DTE_ADDR value written from the power domain, while the domain clocks
are still on, reads back correctly, and the write rk_iommu_resume() does
microseconds later does not land at all. The vendor DT names the CBUF
clocks as that MMU's interface clocks and its driver keeps every NPU
clock on for as long as the device is powered.
The driver side of this is already upstream, commit 841363ebb508
("iommu/rockchip: Take all DT clocks"), which switched rk_iommu to
devm_clk_bulk_get_all(). Widen the schema to match so those nodes can
be described. minItems stays at 2, so every existing devicetree, which
all carry exactly aclk and iface, is unaffected.
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
.../devicetree/bindings/iommu/rockchip,iommu.yaml | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml b/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
index 6ce41d11f..a3cedcaaa 100644
--- a/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
+++ b/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
@@ -42,14 +42,22 @@ properties:
minItems: 1
clocks:
+ minItems: 2
items:
- description: Core clock
- description: Interface clock
+ - description: Compute clock, RK3576 NPU MMUs only
+ - description: Convolution buffer core clock, RK3576 NPU MMUs only
+ - description: Convolution buffer interface clock, RK3576 NPU MMUs only
clock-names:
+ minItems: 2
items:
- const: aclk
- const: iface
+ - const: npu
+ - const: aclk_cbuf
+ - const: hclk_cbuf
"#iommu-cells":
const: 0
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v7 04/10] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set
2026-08-12 9:40 ` [PATCH v7 04/10] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set Jiaxing Hu
@ 2026-08-12 10:45 ` Diederik de Haas
2026-08-13 9:27 ` Jiaxing Hu
0 siblings, 1 reply; 23+ messages in thread
From: Diederik de Haas @ 2026-08-12 10:45 UTC (permalink / raw)
To: Jiaxing Hu, tomeu, heiko, robh, krzk+dt, conor+dt, joro, will,
robin.murphy, ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel
Hi Jiaxing,
On Wed Aug 12, 2026 at 11:40 AM CEST, Jiaxing Hu wrote:
> The RK3576 NPU MMUs need more than aclk and iface. With only those two
> enabled the MMU accepts reads but silently drops register writes: a
> DTE_ADDR value written from the power domain, while the domain clocks
> are still on, reads back correctly, and the write rk_iommu_resume() does
> microseconds later does not land at all. The vendor DT names the CBUF
> clocks as that MMU's interface clocks and its driver keeps every NPU
> clock on for as long as the device is powered.
>
> The driver side of this is already upstream, commit 841363ebb508
> ("iommu/rockchip: Take all DT clocks"), which switched rk_iommu to
> devm_clk_bulk_get_all(). Widen the schema to match so those nodes can
> be described. minItems stays at 2, so every existing devicetree, which
> all carry exactly aclk and iface, is unaffected.
>
> Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
> ---
> .../devicetree/bindings/iommu/rockchip,iommu.yaml | 8 ++++++++
> 1 file changed, 8 insertions(+)
>
> diff --git a/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml b/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
> index 6ce41d11f..a3cedcaaa 100644
> --- a/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
> +++ b/Documentation/devicetree/bindings/iommu/rockchip,iommu.yaml
> @@ -42,14 +42,22 @@ properties:
> minItems: 1
>
> clocks:
> + minItems: 2
> items:
> - description: Core clock
> - description: Interface clock
> + - description: Compute clock, RK3576 NPU MMUs only
> + - description: Convolution buffer core clock, RK3576 NPU MMUs only
> + - description: Convolution buffer interface clock, RK3576 NPU MMUs only
Drop the ", RK3576 NPU MMUs only" part as it is not future proof, not
needed, not enforceable and not enforced.
IIUC, only a RK3576 NPU MMU can and should have 5 clocks, but a non-NPU
RK3576 MMU should only have 2 clocks, just like any MMU for RK3568 and
RK3588.
So you'd need a new compatible for RK3576 NPU MMU and enforce that only
that one has exactly 5 clocks, while all other compatibles are only
allowed to have 2 clocks.
Right now, it is allowed that a ``rockchip,rk3568-iommu`` compatible has
5 clocks while a RK3576 NPU MMU only has 2. Both are incorrect.
Cheers,
Diederik
> clock-names:
> + minItems: 2
> items:
> - const: aclk
> - const: iface
> + - const: npu
> + - const: aclk_cbuf
> + - const: hclk_cbuf
>
> "#iommu-cells":
> const: 0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v7 04/10] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set
2026-08-12 10:45 ` Diederik de Haas
@ 2026-08-13 9:27 ` Jiaxing Hu
0 siblings, 0 replies; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-13 9:27 UTC (permalink / raw)
To: diederik
Cc: heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy, tomeu,
royalnet026, iommu, devicetree, linux-rockchip, linux-arm-kernel,
linux-kernel, Jiaxing Hu
Hi Diederik,
You are right, and worse than that, this is the same thing you asked
for on v6 and I told you it was fixed.
My reply then said v7 would carry a compatible of its own and pin each
side with an allOf. It does not. What v7 actually contains is a
minItems of 2 and three descriptions ending in "RK3576 NPU MMUs only",
which is a comment, not a schema, and it leaves both of the cases you
name allowed: an rk3568-iommu with five clocks, and an RK3576 NPU MMU
with two. I do not have an explanation for the gap between what I said
and what I sent, only the fix.
v8 does what you and the bot asked for the first time. The MMU nodes
get their own compatible,
compatible = "rockchip,rk3576-npu-iommu", "rockchip,rk3568-iommu";
the enum gains it, and an allOf pins both sides so neither can borrow
the other's clock set:
allOf:
- if:
properties:
compatible:
contains:
const: rockchip,rk3576-npu-iommu
then:
properties:
clocks:
minItems: 5
maxItems: 5
clock-names:
items:
- const: aclk
- const: iface
- const: npu
- const: aclk_cbuf
- const: hclk_cbuf
else:
properties:
clocks:
maxItems: 2
clock-names:
maxItems: 2
so every existing devicetree keeps exactly two clocks and only the new
compatible may have five. The ", RK3576 NPU MMUs only" wording goes
away with it, since the schema then says it.
The DTS patch changes with it, since v7's NPU MMU nodes use the plain
rockchip,rk3576-iommu string.
I will not claim it is fixed this time until the patch is in front of
you.
Thanks for catching it twice,
Jiaxing
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v7 05/10] pmdomain/rockchip: add optional per-domain power-on settle delay
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
` (3 preceding siblings ...)
2026-08-12 9:40 ` [PATCH v7 04/10] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set Jiaxing Hu
@ 2026-08-12 9:41 ` Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 06/10] pmdomain/rockchip: cycle optional power-domain resets on power-on Jiaxing Hu
` (4 subsequent siblings)
9 siblings, 0 replies; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:41 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
The RK3576 NPU domains need a short settle time after the idle request
is released before the QoS registers behind the domain answer. Without
it rockchip_pmu_restore_qos() reads back zeroes, and the NPU throws an
async SError on the first cold power-on.
Give rockchip_domain_info an optional delay_us and wait for it between
releasing idle and restoring QoS. Rename DOMAIN_M_O_R_G to
DOMAIN_M_O_R_G_W, since the suffixes name the fields the macro sets and
this one now also carries a wakeup delay; RK3576 is its only user, so
the old spelling is not kept around.
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
drivers/pmdomain/rockchip/pm-domains.c | 52 +++++++++++++++-----------
1 file changed, 30 insertions(+), 22 deletions(-)
diff --git a/drivers/pmdomain/rockchip/pm-domains.c b/drivers/pmdomain/rockchip/pm-domains.c
index ba66ae719..e1857f878 100644
--- a/drivers/pmdomain/rockchip/pm-domains.c
+++ b/drivers/pmdomain/rockchip/pm-domains.c
@@ -18,6 +18,7 @@
#include <linux/of_address.h>
#include <linux/of_clk.h>
#include <linux/clk.h>
+#include <linux/delay.h>
#include <linux/regmap.h>
#include <linux/regulator/consumer.h>
#include <linux/mfd/syscon.h>
@@ -59,6 +60,7 @@ struct rockchip_domain_info {
u32 pwr_offset;
u32 mem_offset;
u32 req_offset;
+ u32 delay_us;
};
struct rockchip_pmu_info {
@@ -185,7 +187,7 @@ struct rockchip_pmu {
.need_regulator = regulator, \
}
-#define DOMAIN_M_O_R_G(_name, p_offset, pwr, status, m_offset, m_status, r_status, r_offset, req, idle, ack, g_mask, wakeup) \
+#define DOMAIN_M_O_R_G_W(_name, p_offset, pwr, status, m_offset, m_status, r_status, r_offset, req, idle, ack, g_mask, delay, wakeup) \
{ \
.name = _name, \
.pwr_offset = p_offset, \
@@ -200,6 +202,7 @@ struct rockchip_pmu {
.req_mask = (req), \
.idle_mask = (idle), \
.clk_ungate_mask = (g_mask), \
+ .delay_us = (delay), \
.ack_mask = (ack), \
.active_wakeup = wakeup, \
}
@@ -258,8 +261,8 @@ struct rockchip_pmu {
#define DOMAIN_RK3568(name, pwr, req, wakeup, regulator) \
DOMAIN_M_R(name, pwr, pwr, req, req, req, wakeup, regulator)
-#define DOMAIN_RK3576(name, p_offset, pwr, status, r_status, r_offset, req, idle, g_mask, wakeup) \
- DOMAIN_M_O_R_G(name, p_offset, pwr, status, 0, r_status, r_status, r_offset, req, idle, idle, g_mask, wakeup)
+#define DOMAIN_RK3576(name, p_offset, pwr, status, r_status, r_offset, req, idle, g_mask, delay, wakeup) \
+ DOMAIN_M_O_R_G_W(name, p_offset, pwr, status, 0, r_status, r_status, r_offset, req, idle, idle, g_mask, delay, wakeup)
/*
* Dynamic Memory Controller may need to coordinate with us -- see
@@ -681,6 +684,10 @@ static int rockchip_pd_power(struct rockchip_pm_domain *pd, bool power_on)
if (ret < 0)
goto out;
+ /* Some domains need to settle before the QoS registers answer. */
+ if (pd->info->delay_us)
+ udelay(pd->info->delay_us);
+
rockchip_pmu_restore_qos(pd);
}
@@ -1300,25 +1307,26 @@ static const struct rockchip_domain_info rk3568_pm_domains[] = {
};
static const struct rockchip_domain_info rk3576_pm_domains[] = {
- [RK3576_PD_NPU] = DOMAIN_RK3576("npu", 0x0, BIT(0), BIT(0), 0, 0x0, 0, 0, 0, false),
- [RK3576_PD_NVM] = DOMAIN_RK3576("nvm", 0x0, BIT(6), 0, BIT(6), 0x4, BIT(2), BIT(18), BIT(2), false),
- [RK3576_PD_SDGMAC] = DOMAIN_RK3576("sdgmac", 0x0, BIT(7), 0, BIT(7), 0x4, BIT(1), BIT(17), 0x6, false),
- [RK3576_PD_AUDIO] = DOMAIN_RK3576("audio", 0x0, BIT(8), 0, BIT(8), 0x4, BIT(0), BIT(16), BIT(0), false),
- [RK3576_PD_PHP] = DOMAIN_RK3576("php", 0x0, BIT(9), 0, BIT(9), 0x0, BIT(15), BIT(15), BIT(15), false),
- [RK3576_PD_SUBPHP] = DOMAIN_RK3576("subphp", 0x0, BIT(10), 0, BIT(10), 0x0, 0, 0, 0, false),
- [RK3576_PD_VOP] = DOMAIN_RK3576("vop", 0x0, BIT(11), 0, BIT(11), 0x0, 0x6000, 0x6000, 0x6000, false),
- [RK3576_PD_VO1] = DOMAIN_RK3576("vo1", 0x0, BIT(14), 0, BIT(14), 0x0, BIT(12), BIT(12), 0x7000, false),
- [RK3576_PD_VO0] = DOMAIN_RK3576("vo0", 0x0, BIT(15), 0, BIT(15), 0x0, BIT(11), BIT(11), 0x6800, false),
- [RK3576_PD_USB] = DOMAIN_RK3576("usb", 0x4, BIT(0), 0, BIT(16), 0x0, BIT(10), BIT(10), 0x6400, true),
- [RK3576_PD_VI] = DOMAIN_RK3576("vi", 0x4, BIT(1), 0, BIT(17), 0x0, BIT(9), BIT(9), BIT(9), false),
- [RK3576_PD_VEPU0] = DOMAIN_RK3576("vepu0", 0x4, BIT(2), 0, BIT(18), 0x0, BIT(7), BIT(7), 0x280, false),
- [RK3576_PD_VEPU1] = DOMAIN_RK3576("vepu1", 0x4, BIT(3), 0, BIT(19), 0x0, BIT(8), BIT(8), BIT(8), false),
- [RK3576_PD_VDEC] = DOMAIN_RK3576("vdec", 0x4, BIT(4), 0, BIT(20), 0x0, BIT(6), BIT(6), BIT(6), false),
- [RK3576_PD_VPU] = DOMAIN_RK3576("vpu", 0x4, BIT(5), 0, BIT(21), 0x0, BIT(5), BIT(5), BIT(5), false),
- [RK3576_PD_NPUTOP] = DOMAIN_RK3576("nputop", 0x4, BIT(6), 0, BIT(22), 0x0, 0x18, 0x18, 0x18, false),
- [RK3576_PD_NPU0] = DOMAIN_RK3576("npu0", 0x4, BIT(7), 0, BIT(23), 0x0, BIT(1), BIT(1), 0x1a, false),
- [RK3576_PD_NPU1] = DOMAIN_RK3576("npu1", 0x4, BIT(8), 0, BIT(24), 0x0, BIT(2), BIT(2), 0x1c, false),
- [RK3576_PD_GPU] = DOMAIN_RK3576("gpu", 0x4, BIT(9), 0, BIT(25), 0x0, BIT(0), BIT(0), BIT(0), false),
+ /* name p_offset pwr status r_status r_offset req idle g_mask delay wakeup */
+ [RK3576_PD_NPU] = DOMAIN_RK3576("npu", 0x0, BIT(0), BIT(0), 0, 0x0, 0, 0, 0, 0, false),
+ [RK3576_PD_NVM] = DOMAIN_RK3576("nvm", 0x0, BIT(6), 0, BIT(6), 0x4, BIT(2), BIT(18), BIT(2), 0, false),
+ [RK3576_PD_SDGMAC] = DOMAIN_RK3576("sdgmac", 0x0, BIT(7), 0, BIT(7), 0x4, BIT(1), BIT(17), 0x6, 0, false),
+ [RK3576_PD_AUDIO] = DOMAIN_RK3576("audio", 0x0, BIT(8), 0, BIT(8), 0x4, BIT(0), BIT(16), BIT(0), 0, false),
+ [RK3576_PD_PHP] = DOMAIN_RK3576("php", 0x0, BIT(9), 0, BIT(9), 0x0, BIT(15), BIT(15), BIT(15), 0, false),
+ [RK3576_PD_SUBPHP] = DOMAIN_RK3576("subphp", 0x0, BIT(10), 0, BIT(10), 0x0, 0, 0, 0, 0, false),
+ [RK3576_PD_VOP] = DOMAIN_RK3576("vop", 0x0, BIT(11), 0, BIT(11), 0x0, 0x6000, 0x6000, 0x6000, 0, false),
+ [RK3576_PD_VO1] = DOMAIN_RK3576("vo1", 0x0, BIT(14), 0, BIT(14), 0x0, BIT(12), BIT(12), 0x7000, 0, false),
+ [RK3576_PD_VO0] = DOMAIN_RK3576("vo0", 0x0, BIT(15), 0, BIT(15), 0x0, BIT(11), BIT(11), 0x6800, 0, false),
+ [RK3576_PD_USB] = DOMAIN_RK3576("usb", 0x4, BIT(0), 0, BIT(16), 0x0, BIT(10), BIT(10), 0x6400, 0, true),
+ [RK3576_PD_VI] = DOMAIN_RK3576("vi", 0x4, BIT(1), 0, BIT(17), 0x0, BIT(9), BIT(9), BIT(9), 0, false),
+ [RK3576_PD_VEPU0] = DOMAIN_RK3576("vepu0", 0x4, BIT(2), 0, BIT(18), 0x0, BIT(7), BIT(7), 0x280, 0, false),
+ [RK3576_PD_VEPU1] = DOMAIN_RK3576("vepu1", 0x4, BIT(3), 0, BIT(19), 0x0, BIT(8), BIT(8), BIT(8), 0, false),
+ [RK3576_PD_VDEC] = DOMAIN_RK3576("vdec", 0x4, BIT(4), 0, BIT(20), 0x0, BIT(6), BIT(6), BIT(6), 0, false),
+ [RK3576_PD_VPU] = DOMAIN_RK3576("vpu", 0x4, BIT(5), 0, BIT(21), 0x0, BIT(5), BIT(5), BIT(5), 0, false),
+ [RK3576_PD_NPUTOP] = DOMAIN_RK3576("nputop", 0x4, BIT(6), 0, BIT(22), 0x0, 0x18, 0x18, 0x18, 15, false),
+ [RK3576_PD_NPU0] = DOMAIN_RK3576("npu0", 0x4, BIT(7), 0, BIT(23), 0x0, BIT(1), BIT(1), 0x1a, 15, false),
+ [RK3576_PD_NPU1] = DOMAIN_RK3576("npu1", 0x4, BIT(8), 0, BIT(24), 0x0, BIT(2), BIT(2), 0x1c, 15, false),
+ [RK3576_PD_GPU] = DOMAIN_RK3576("gpu", 0x4, BIT(9), 0, BIT(25), 0x0, BIT(0), BIT(0), BIT(0), 0, false),
};
static const struct rockchip_domain_info rk3588_pm_domains[] = {
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* [PATCH v7 06/10] pmdomain/rockchip: cycle optional power-domain resets on power-on
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
` (4 preceding siblings ...)
2026-08-12 9:41 ` [PATCH v7 05/10] pmdomain/rockchip: add optional per-domain power-on settle delay Jiaxing Hu
@ 2026-08-12 9:41 ` Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 07/10] accel/rocket: select the per-core clock and reset counts from match data Jiaxing Hu
` (3 subsequent siblings)
9 siblings, 0 replies; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:41 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
Some Rockchip domains come out of power-on with their bus interface in
an undefined state. On the RK3576 NPU this shows up as a hang on the
first register access after the domain is switched on, and pulsing the
domain's resets at this point clears it.
Take the domain node's resets if it has any, and pulse them between
releasing idle and restoring QoS. The resets are optional, so domains
that do not list any are unaffected.
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
drivers/pmdomain/rockchip/pm-domains.c | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
diff --git a/drivers/pmdomain/rockchip/pm-domains.c b/drivers/pmdomain/rockchip/pm-domains.c
index e1857f878..4eebb5d99 100644
--- a/drivers/pmdomain/rockchip/pm-domains.c
+++ b/drivers/pmdomain/rockchip/pm-domains.c
@@ -19,6 +19,7 @@
#include <linux/of_clk.h>
#include <linux/clk.h>
#include <linux/delay.h>
+#include <linux/reset.h>
#include <linux/regmap.h>
#include <linux/regulator/consumer.h>
#include <linux/mfd/syscon.h>
@@ -103,6 +104,7 @@ struct rockchip_pm_domain {
struct clk_bulk_data *clks;
struct device_node *node;
struct regulator *supply;
+ struct reset_control *resets;
};
struct rockchip_pmu {
@@ -688,6 +690,13 @@ static int rockchip_pd_power(struct rockchip_pm_domain *pd, bool power_on)
if (pd->info->delay_us)
udelay(pd->info->delay_us);
+ /* Optional: some domains need their resets cycled after power-on. */
+ if (pd->resets) {
+ reset_control_assert(pd->resets);
+ udelay(10);
+ reset_control_deassert(pd->resets);
+ }
+
rockchip_pmu_restore_qos(pd);
}
@@ -857,6 +866,14 @@ static int rockchip_pm_add_one_domain(struct rockchip_pmu *pmu,
if (error)
goto err_put_clocks;
+ pd->resets = of_reset_control_array_get_optional_exclusive(node);
+ if (IS_ERR(pd->resets)) {
+ error = dev_err_probe(pmu->dev, PTR_ERR(pd->resets),
+ "%pOFn: failed to get resets\n", node);
+ pd->resets = NULL;
+ goto err_unprepare_clocks;
+ }
+
pd->num_qos = of_count_phandle_with_args(node, "pm_qos",
NULL);
@@ -927,6 +944,7 @@ static int rockchip_pm_add_one_domain(struct rockchip_pmu *pmu,
clk_bulk_unprepare(pd->num_clks, pd->clks);
err_put_clocks:
clk_bulk_put(pd->num_clks, pd->clks);
+ reset_control_put(pd->resets);
return error;
}
@@ -945,6 +963,7 @@ static void rockchip_pm_remove_one_domain(struct rockchip_pm_domain *pd)
clk_bulk_unprepare(pd->num_clks, pd->clks);
clk_bulk_put(pd->num_clks, pd->clks);
+ reset_control_put(pd->resets);
/* protect the zeroing of pm->num_clks */
mutex_lock(&pd->pmu->mutex);
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* [PATCH v7 07/10] accel/rocket: select the per-core clock and reset counts from match data
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
` (5 preceding siblings ...)
2026-08-12 9:41 ` [PATCH v7 06/10] pmdomain/rockchip: cycle optional power-domain resets on power-on Jiaxing Hu
@ 2026-08-12 9:41 ` Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support Jiaxing Hu
` (2 subsequent siblings)
9 siblings, 0 replies; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:41 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
The RK3576 carries the same RKNN block with a different set of clocks and
resets, so the counts cannot stay compile-time constants. Add a soc_data
struct to the of_device_id match data and take the bulk counts from it.
RK3588 keeps four clocks and two resets, so nothing changes for it, and
the arrays keep their present sizes: the SoC that needs a longer one
grows it in the patch that adds the names.
rocket_core_reset() is switched over as well. It is the same array, and
leaving it on ARRAY_SIZE() would walk entries that were never acquired
once a SoC asks for fewer.
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
drivers/accel/rocket/rocket_core.c | 8 ++++----
drivers/accel/rocket/rocket_core.h | 7 +++++++
drivers/accel/rocket/rocket_drv.c | 12 +++++++++---
3 files changed, 20 insertions(+), 7 deletions(-)
diff --git a/drivers/accel/rocket/rocket_core.c b/drivers/accel/rocket/rocket_core.c
index 5dd260bac..b202d1581 100644
--- a/drivers/accel/rocket/rocket_core.c
+++ b/drivers/accel/rocket/rocket_core.c
@@ -23,7 +23,7 @@ int rocket_core_init(struct rocket_core *core)
core->resets[0].id = "srst_a";
core->resets[1].id = "srst_h";
- err = devm_reset_control_bulk_get_exclusive(&pdev->dev, ARRAY_SIZE(core->resets),
+ err = devm_reset_control_bulk_get_exclusive(&pdev->dev, core->soc->num_resets,
core->resets);
if (err)
return dev_err_probe(dev, err, "failed to get resets for core %d\n", core->index);
@@ -32,7 +32,7 @@ int rocket_core_init(struct rocket_core *core)
core->clks[1].id = "hclk";
core->clks[2].id = "npu";
core->clks[3].id = "pclk";
- err = devm_clk_bulk_get(dev, ARRAY_SIZE(core->clks), core->clks);
+ err = devm_clk_bulk_get(dev, core->soc->num_clks, core->clks);
if (err)
return dev_err_probe(dev, err, "failed to get clocks for core %d\n", core->index);
@@ -109,9 +109,9 @@ void rocket_core_fini(struct rocket_core *core)
void rocket_core_reset(struct rocket_core *core)
{
- reset_control_bulk_assert(ARRAY_SIZE(core->resets), core->resets);
+ reset_control_bulk_assert(core->soc->num_resets, core->resets);
udelay(10);
- reset_control_bulk_deassert(ARRAY_SIZE(core->resets), core->resets);
+ reset_control_bulk_deassert(core->soc->num_resets, core->resets);
}
diff --git a/drivers/accel/rocket/rocket_core.h b/drivers/accel/rocket/rocket_core.h
index f6d738285..ba74c5339 100644
--- a/drivers/accel/rocket/rocket_core.h
+++ b/drivers/accel/rocket/rocket_core.h
@@ -27,9 +27,16 @@
#define rocket_core_writel(core, reg, value) \
writel(value, (core)->core_iomem + (REG_CORE_##reg) - REG_CORE_S_STATUS)
+/* Per-SoC differences, selected by the of_device_id match data. */
+struct rocket_soc_data {
+ unsigned int num_clks; /* clk_bulk count */
+ unsigned int num_resets; /* reset_bulk count */
+};
+
struct rocket_core {
struct device *dev;
struct rocket_device *rdev;
+ const struct rocket_soc_data *soc;
unsigned int index;
int irq;
diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocket_drv.c
index 8bbbce594..6e7dc91c5 100644
--- a/drivers/accel/rocket/rocket_drv.c
+++ b/drivers/accel/rocket/rocket_drv.c
@@ -176,6 +176,7 @@ static int rocket_probe(struct platform_device *pdev)
rdev->cores[core].rdev = rdev;
rdev->cores[core].dev = &pdev->dev;
+ rdev->cores[core].soc = of_device_get_match_data(&pdev->dev);
rdev->cores[core].index = core;
rdev->num_cores++;
@@ -213,8 +214,13 @@ static void rocket_remove(struct platform_device *pdev)
}
}
+static const struct rocket_soc_data rk3588_soc_data = {
+ .num_clks = 4,
+ .num_resets = 2,
+};
+
static const struct of_device_id dt_match[] = {
- { .compatible = "rockchip,rk3588-rknn-core" },
+ { .compatible = "rockchip,rk3588-rknn-core", .data = &rk3588_soc_data },
{}
};
MODULE_DEVICE_TABLE(of, dt_match);
@@ -240,7 +246,7 @@ static int rocket_device_runtime_resume(struct device *dev)
if (core < 0)
return -ENODEV;
- err = clk_bulk_prepare_enable(ARRAY_SIZE(rdev->cores[core].clks), rdev->cores[core].clks);
+ err = clk_bulk_prepare_enable(rdev->cores[core].soc->num_clks, rdev->cores[core].clks);
if (err) {
dev_err(dev, "failed to enable (%d) clocks for core %d\n", err, core);
return err;
@@ -260,7 +266,7 @@ static int rocket_device_runtime_suspend(struct device *dev)
if (!rocket_job_is_idle(&rdev->cores[core]))
return -EBUSY;
- clk_bulk_disable_unprepare(ARRAY_SIZE(rdev->cores[core].clks), rdev->cores[core].clks);
+ clk_bulk_disable_unprepare(rdev->cores[core].soc->num_clks, rdev->cores[core].clks);
return 0;
}
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
` (6 preceding siblings ...)
2026-08-12 9:41 ` [PATCH v7 07/10] accel/rocket: select the per-core clock and reset counts from match data Jiaxing Hu
@ 2026-08-12 9:41 ` Jiaxing Hu
2026-08-12 12:48 ` Igor Paunovic
2026-08-12 9:41 ` [PATCH v7 09/10] arm64: dts: rockchip: rk3576: add NPU (RKNN) nodes Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 10/10] arm64: dts: rockchip: rk3576-rock-4d: enable NPU Jiaxing Hu
9 siblings, 1 reply; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:41 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
The RK3576 has two cores of the same RKNN block and a few platform
differences:
- the CBUF (convolution buffer) has its own clock domain, so the core
needs six clocks rather than four;
- the BIU reset moved into the power domain, leaving one reset here;
- the NPU spans two power domains, and a device with more than one is
skipped by the driver-core single-domain auto-attach, so the list has
to be attached explicitly;
- PC_TASK_CON packs the task number with sixteen bits rather than
twelve, moving the three controls above it up by four.
That last one is the reason this series has been reporting, since v3,
that the block accepts exactly one task per reset. rocket_registers.h is
generated from the RK3588 description, so writing it unchanged to an
RK3576 asks for task_number 0x7001, which is 28673 tasks, and puts
TASK_COUNT_CLEAR on a bit that does nothing. The counter is then only
ever cleared by a reset.
The layout was confirmed by Chaoyi Chen of Rockchip, including a fourth
control at BIT(18), task_last_layer_clear, which belongs on every submit
alongside the count clear:
https://lore.kernel.org/all/4f300b78-d96d-4d98-8819-dc292b0c9b97@rock-chips.com/
With that written correctly a job of several tasks runs to completion,
the completion interrupt arrives, and /proc/interrupts counts up. A
convolution submitted three times with three different inputs is byte
exact against the CPU reference each time, with no reset in between and
with nothing retiring the job but the interrupt.
All of it hangs off the soc_data added earlier, so the RK3588 path keeps
its existing counts and behaviour.
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
drivers/accel/rocket/rocket_core.c | 20 ++++++++
drivers/accel/rocket/rocket_core.h | 8 +--
drivers/accel/rocket/rocket_device.c | 4 ++
drivers/accel/rocket/rocket_drv.c | 10 ++++
drivers/accel/rocket/rocket_job.c | 76 ++++++++++++++++++++++------
5 files changed, 99 insertions(+), 19 deletions(-)
diff --git a/drivers/accel/rocket/rocket_core.c b/drivers/accel/rocket/rocket_core.c
index b202d1581..5f3155135 100644
--- a/drivers/accel/rocket/rocket_core.c
+++ b/drivers/accel/rocket/rocket_core.c
@@ -8,6 +8,7 @@
#include <linux/err.h>
#include <linux/iommu.h>
#include <linux/platform_device.h>
+#include <linux/pm_domain.h>
#include <linux/pm_runtime.h>
#include <linux/reset.h>
@@ -21,6 +22,7 @@ int rocket_core_init(struct rocket_core *core)
u32 version;
int err = 0;
+ /* RK3576 moves the BIU reset into its power domain and takes only srst_a. */
core->resets[0].id = "srst_a";
core->resets[1].id = "srst_h";
err = devm_reset_control_bulk_get_exclusive(&pdev->dev, core->soc->num_resets,
@@ -32,6 +34,9 @@ int rocket_core_init(struct rocket_core *core)
core->clks[1].id = "hclk";
core->clks[2].id = "npu";
core->clks[3].id = "pclk";
+ /* RK3576 clocks the CBUF separately; the compute path stalls without these. */
+ core->clks[4].id = "aclk_cbuf";
+ core->clks[5].id = "hclk_cbuf";
err = devm_clk_bulk_get(dev, core->soc->num_clks, core->clks);
if (err)
return dev_err_probe(dev, err, "failed to get clocks for core %d\n", core->index);
@@ -60,6 +65,21 @@ int rocket_core_init(struct rocket_core *core)
if (err)
return err;
+ /*
+ * RK3576 spans two power domains, and a multi-domain device is skipped
+ * by the driver-core single-domain auto-attach, so attach the list here.
+ * This goes before the first thing that would have to be unwound, so a
+ * failure can simply return.
+ */
+ if (core->soc->multi_power_domain) {
+ struct dev_pm_domain_list *pd_list;
+
+ err = devm_pm_domain_attach_list(dev, NULL, &pd_list);
+ if (err < 0)
+ return dev_err_probe(dev, err,
+ "failed to attach NPU power domains\n");
+ }
+
core->iommu_group = iommu_group_get(dev);
err = rocket_job_init(core);
diff --git a/drivers/accel/rocket/rocket_core.h b/drivers/accel/rocket/rocket_core.h
index ba74c5339..8c8d1f453 100644
--- a/drivers/accel/rocket/rocket_core.h
+++ b/drivers/accel/rocket/rocket_core.h
@@ -29,8 +29,10 @@
/* Per-SoC differences, selected by the of_device_id match data. */
struct rocket_soc_data {
- unsigned int num_clks; /* clk_bulk count */
- unsigned int num_resets; /* reset_bulk count */
+ unsigned int num_clks; /* clk_bulk count: 4 base, 6 with CBUF */
+ unsigned int num_resets; /* reset_bulk count: 2 base, 1 on RK3576 */
+ bool multi_power_domain; /* device spans more than one PM domain */
+ bool task_con_16bit; /* PC_TASK_CON uses the 16-bit task number */
};
struct rocket_core {
@@ -43,7 +45,7 @@ struct rocket_core {
void __iomem *pc_iomem;
void __iomem *cna_iomem;
void __iomem *core_iomem;
- struct clk_bulk_data clks[4];
+ struct clk_bulk_data clks[6];
struct reset_control_bulk_data resets[2];
struct iommu_group *iommu_group;
diff --git a/drivers/accel/rocket/rocket_device.c b/drivers/accel/rocket/rocket_device.c
index 46e6ee1e7..bfb00f967 100644
--- a/drivers/accel/rocket/rocket_device.c
+++ b/drivers/accel/rocket/rocket_device.c
@@ -31,6 +31,10 @@ struct rocket_device *rocket_device_init(struct platform_device *pdev,
if (of_device_is_available(core_node))
num_cores++;
+ for_each_compatible_node(core_node, NULL, "rockchip,rk3576-rknn-core")
+ if (of_device_is_available(core_node))
+ num_cores++;
+
rdev->cores = devm_kcalloc(dev, num_cores, sizeof(*rdev->cores), GFP_KERNEL);
if (!rdev->cores)
return ERR_PTR(-ENOMEM);
diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocket_drv.c
index 6e7dc91c5..f333fe466 100644
--- a/drivers/accel/rocket/rocket_drv.c
+++ b/drivers/accel/rocket/rocket_drv.c
@@ -217,10 +217,20 @@ static void rocket_remove(struct platform_device *pdev)
static const struct rocket_soc_data rk3588_soc_data = {
.num_clks = 4,
.num_resets = 2,
+ .multi_power_domain = false,
+ .task_con_16bit = false,
+};
+
+static const struct rocket_soc_data rk3576_soc_data = {
+ .num_clks = 6,
+ .num_resets = 1,
+ .multi_power_domain = true,
+ .task_con_16bit = true,
};
static const struct of_device_id dt_match[] = {
{ .compatible = "rockchip,rk3588-rknn-core", .data = &rk3588_soc_data },
+ { .compatible = "rockchip,rk3576-rknn-core", .data = &rk3576_soc_data },
{}
};
MODULE_DEVICE_TABLE(of, dt_match);
diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocket_job.c
index 4c01b703e..493b3bf97 100644
--- a/drivers/accel/rocket/rocket_job.c
+++ b/drivers/accel/rocket/rocket_job.c
@@ -21,6 +21,35 @@
#define JOB_TIMEOUT_MS 500
+/*
+ * RK3576 arms the same DPU completion as RK3588, but the interrupt never
+ * reaches the GIC. The completion itself is visible in INTERRUPT_RAW_STATUS,
+ * so sample that instead. The tick cap bounds jobs that never raise it at all,
+ * which is the same open problem as the wrong inference results.
+ */
+/*
+ * PC_TASK_CON packs the task number with three controls, and the field widths
+ * are not the same on every SoC. rocket_registers.h is generated from the
+ * RK3588 description, where the task number is twelve bits:
+ *
+ * RK3588 BIT[11:0] task_number, BIT[12] pp_en, BIT[13] count_clear
+ * RK3576 BIT[15:0] task_number, BIT[16] pp_en, BIT[17] count_clear,
+ * BIT[18] last_layer_clear
+ *
+ * The RK3576 layout was confirmed by Chaoyi Chen of Rockchip:
+ * https://lore.kernel.org/all/4f300b78-d96d-4d98-8819-dc292b0c9b97@rock-chips.com/
+ *
+ * Writing the RK3588 layout to an RK3576 therefore asks for task_number
+ * 0x7001, that is 28673 tasks, and lands the count clear on a bit that does
+ * nothing. The task counter is then only ever cleared by a reset, which is
+ * exactly the "one task per reset" behaviour this series has been reporting
+ * since v3.
+ */
+#define RK3576_PC_TASK_CON_TASK_NUMBER(n) ((n) & 0xffff)
+#define RK3576_PC_TASK_CON_PP_EN BIT(16)
+#define RK3576_PC_TASK_CON_COUNT_CLEAR BIT(17)
+#define RK3576_PC_TASK_CON_LAST_LAYER_CLEAR BIT(18)
+
static struct rocket_job *
to_rocket_job(struct drm_sched_job *sched_job)
{
@@ -142,10 +171,17 @@ static void rocket_job_hw_submit(struct rocket_core *core, struct rocket_job *jo
rocket_pc_writel(core, INTERRUPT_MASK, PC_INTERRUPT_MASK_DPU_0 | PC_INTERRUPT_MASK_DPU_1);
rocket_pc_writel(core, INTERRUPT_CLEAR, PC_INTERRUPT_CLEAR_DPU_0 | PC_INTERRUPT_CLEAR_DPU_1);
- rocket_pc_writel(core, TASK_CON, PC_TASK_CON_RESERVED_0(1) |
- PC_TASK_CON_TASK_COUNT_CLEAR(1) |
- PC_TASK_CON_TASK_NUMBER(1) |
- PC_TASK_CON_TASK_PP_EN(1));
+ if (core->soc->task_con_16bit)
+ rocket_pc_writel(core, TASK_CON,
+ RK3576_PC_TASK_CON_LAST_LAYER_CLEAR |
+ RK3576_PC_TASK_CON_COUNT_CLEAR |
+ RK3576_PC_TASK_CON_PP_EN |
+ RK3576_PC_TASK_CON_TASK_NUMBER(1));
+ else
+ rocket_pc_writel(core, TASK_CON, PC_TASK_CON_RESERVED_0(1) |
+ PC_TASK_CON_TASK_COUNT_CLEAR(1) |
+ PC_TASK_CON_TASK_NUMBER(1) |
+ PC_TASK_CON_TASK_PP_EN(1));
rocket_pc_writel(core, TASK_DMA_BASE_ADDR, PC_TASK_DMA_BASE_ADDR_DMA_BASE_ADDR(0x0));
@@ -341,6 +377,25 @@ static struct dma_fence *rocket_job_run(struct drm_sched_job *sched_job)
return ERR_PTR(ret);
}
+/* Start the job's next task, or retire it. Caller holds job_lock. */
+static void rocket_job_next_locked(struct rocket_core *core)
+{
+ lockdep_assert_held(&core->job_lock);
+
+ if (!core->in_flight_job)
+ return;
+
+ if (core->in_flight_job->next_task_idx < core->in_flight_job->task_count) {
+ rocket_job_hw_submit(core, core->in_flight_job);
+ return;
+ }
+
+ iommu_detach_group(NULL, iommu_group_get(core->dev));
+ dma_fence_signal(core->in_flight_job->done_fence);
+ pm_runtime_put_autosuspend(core->dev);
+ core->in_flight_job = NULL;
+}
+
static void rocket_job_handle_irq(struct rocket_core *core)
{
pm_runtime_mark_last_busy(core->dev);
@@ -354,17 +409,7 @@ static void rocket_job_handle_irq(struct rocket_core *core)
rocket_pc_writel(core, OPERATION_ENABLE, 0x0);
rocket_pc_writel(core, INTERRUPT_CLEAR, 0x1ffff);
- if (core->in_flight_job) {
- if (core->in_flight_job->next_task_idx < core->in_flight_job->task_count) {
- rocket_job_hw_submit(core, core->in_flight_job);
- return;
- }
-
- iommu_detach_group(NULL, iommu_group_get(core->dev));
- dma_fence_signal(core->in_flight_job->done_fence);
- pm_runtime_put_autosuspend(core->dev);
- core->in_flight_job = NULL;
- }
+ rocket_job_next_locked(core);
}
}
@@ -644,7 +689,6 @@ int rocket_ioctl_submit(struct drm_device *dev, void *data, struct drm_file *fil
}
}
-
for (i = 0; i < args->job_count; i++)
rocket_ioctl_submit_job(dev, file, &jobs[i]);
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support
2026-08-12 9:41 ` [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support Jiaxing Hu
@ 2026-08-12 12:48 ` Igor Paunovic
2026-08-13 9:26 ` Jiaxing Hu
0 siblings, 1 reply; 23+ messages in thread
From: Igor Paunovic @ 2026-08-12 12:48 UTC (permalink / raw)
To: Jiaxing Hu, tomeu, heiko, robh, krzk+dt, conor+dt, joro, will,
robin.murphy, ulfh, p.zabel, ogabbay, zhangqing
Cc: Igor Paunovic, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel
Two things here, one of which I think has to be fixed before this
lands.
The first is a comment that outlived its subject. This patch adds the
following just above the PC_TASK_CON block:
/*
* RK3576 arms the same DPU completion as RK3588, but the interrupt
* never reaches the GIC. The completion itself is visible in
* INTERRUPT_RAW_STATUS, so sample that instead. The tick cap bounds
* jobs that never raise it at all, which is the same open problem as
* the wrong inference results.
*/
That is the v6 comment for the poll. It states the premise your cover
letter withdraws, it describes machinery this version deletes, and it
has no code under it - the next line opens the second comment block.
Left in, the driver would carry a claim that contradicts both the
commit introducing it and the register description two paragraphs
below it.
The second is placement rather than correctness. This patch also
factors the completion tail out of rocket_job_handle_irq() into
rocket_job_next_locked(). I read that as behaviour-neutral on RK3588 -
the return that used to leave handle_irq() now leaves the helper, and
scoped_guard drops the lock either way - and the numbers I posted on
1/10 bear it out. But it restructures the shared completion path in a
patch whose subject is adding RK3576, which puts a bisect in the wrong
place if it ever turns out not to be neutral. It would sit more
naturally in 1/10, which already touches that function, or in a patch
of its own.
Both of the things I raised on v6 are right in this version. The power
domain list is attached before anything that would have to be unwound,
and the comment saying why a plain return is correct there is a good
addition. clks[] grows in the same patch that adds the two names.
I also went looking for an ARRAY_SIZE(core->clks) or
ARRAY_SIZE(core->resets) left behind, since that would walk six entries
on a four-clock RK3588. All six are converted in 7/10, including the
two in rocket_drv.c's runtime PM callbacks, which are the easiest pair
to miss.
Igor
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support
2026-08-12 12:48 ` Igor Paunovic
@ 2026-08-13 9:26 ` Jiaxing Hu
2026-08-13 9:56 ` Igor Paunovic
0 siblings, 1 reply; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-13 9:26 UTC (permalink / raw)
To: royalnet026
Cc: tomeu, heiko, chaoyi.chen, alchark, dri-devel, linux-rockchip,
linux-arm-kernel, linux-kernel, Jiaxing Hu
Hi Igor,
Thank you for the run, and for reading the patch rather than only
testing it. Both of your points are right and both are fixed for v8.
I am answering your 1/10 question here as well so it stays in one
place.
> That is the v6 comment for the poll.
Yes, and it should not have shipped. The cover letter withdraws the
premise that comment states, the patch under it deletes the machinery
it describes, and it has no code beneath it at all. It is gone in v8.
For the record on how it survived: it was fixed in my tree the day the
correction went to the list and the fix never made it into the series I
formatted. That is the second time a fix has existed here and not
reached what I sent, so I now diff the posted patches against the tree
before sending rather than trusting that they match.
> It would sit more naturally in 1/10, which already touches that
> function, or in a patch of its own.
A patch of its own, placed after 1/10 rather than before it. 1/10 is a
fix with a Fixes tag that someone may want to backport, and it should
stay the smallest thing that fixes the bug. Refactoring the function
first would put the backport on top of a restructure it does not need.
So v8 is 1/10 unchanged, then the extraction on its own, then the
RK3576 patch with no shared-path changes left in it.
> would a synchronize_irq(core->irq) before the guard in
> rocket_reset() be worth having as well?
I think yes, and before the guard is the only place it can go. The
handler takes job_lock, so calling it inside the scoped_guard would
wait for a handler that is waiting for the lock we hold. Before the
guard nothing is held, and both callers, the timedout_job callback and
reset_work, are process context, so it is safe there.
It also closes exactly the window you describe rather than a different
one. drm_sched_stop() stops the scheduler and returns; a threaded
handler already running is untouched by it, and the comment sitting
above that code says "Remaining interrupts have been handled", which is
the assumption your reading breaks. synchronize_irq() makes that
sentence true instead of hopeful.
What it does not do is stop a handler that has already read
in_flight_job from finishing its work on a job the reset is about to
drop. That one wants the check and the write to be one step under the
lock, which is what 1/10 does. The two changes are complementary and I
will send them as such, with the comment reworded to say what is
actually guaranteed.
Your RK3588 numbers are also the only evidence anyone has that 1/10
costs nothing on the path it protects, since I cannot run three cores
here. Carrying the tag to v8:
Tested-by: Igor Paunovic <royalnet026@gmail.com> # RK3588, three cores
One piece of news from the userspace side, since you run MobileNet
through Teflon yourself. As of today the whole of MobileNet V1 runs on
the RK3576 with the open stack: 995 of its 1001 outputs land within one
count of the CPU reference, against 1001 channels of zero in every run
before this. The kernel side of that is this series unchanged; what
moved was four Mesa faults, the last of which was a coefficient buffer
whose second operand has to be 16 byte aligned, which is why every
layer whose output channel count was not a multiple of eight came back
empty.
Thanks again,
Jiaxing
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support
2026-08-13 9:26 ` Jiaxing Hu
@ 2026-08-13 9:56 ` Igor Paunovic
2026-08-14 8:26 ` Jiaxing Hu
0 siblings, 1 reply; 23+ messages in thread
From: Igor Paunovic @ 2026-08-13 9:56 UTC (permalink / raw)
To: Jiaxing Hu
Cc: Igor Paunovic, tomeu, heiko, chaoyi.chen, alchark, dri-devel,
linux-rockchip, linux-arm-kernel, linux-kernel
Hi Jiaxing,
Two short things.
On synchronize_irq(): your placement is right and my question was not
thought through that far. Inside the guard it would wait for a handler
that is itself waiting for job_lock, which is a deadlock rather than a
fence. Before it, with nothing held and both callers in process
context, is the only place it works. I had the window right and the
location wrong.
I will run v8 on the three cores here when it is out, and I will drive
the reset path deliberately rather than wait for a timeout to happen on
its own, since that is the path the two changes are there for.
On MobileNet: 995 of 1001 within one count is a different kind of
number from what this series has been reporting, and it took four Mesa
faults to get there. That is worth saying out loud.
If it would help to know whether the remaining six are RK3576 specific
or common to the stack, I can run the same comparison on RK3588. I
already run MobileNet V1 through the Teflon delegate here, but my
oracle is bit-exactness across repeated runs rather than a per-output
comparison against the CPU, so it would not have noticed six outputs
being off by more than a count. Say the word and I will point it at the
CPU reference the way you did.
Igor
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support
2026-08-13 9:56 ` Igor Paunovic
@ 2026-08-14 8:26 ` Jiaxing Hu
2026-08-14 11:08 ` Igor Paunovic
0 siblings, 1 reply; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-14 8:26 UTC (permalink / raw)
To: royalnet026
Cc: tomeu, heiko, chaoyi.chen, alchark, dri-devel, linux-rockchip,
linux-arm-kernel, linux-kernel, Jiaxing Hu
Hi Igor,
> I had the window right and the location wrong.
That is the useful half. v8 carries synchronize_irq(core->irq) before the
guard in rocket_reset(), with the comment above it saying what is
actually guaranteed rather than "Remaining interrupts have been handled".
Driving the reset path deliberately rather than waiting for a timeout is
worth more than the rest of the run put together, since that is the only
path either change is for.
> Say the word and I will point it at the CPU reference the way you did.
Please do, and thank you. It is the one comparison I cannot produce, and
it separates two things that look identical from here: a defect specific
to this SoC, and the reference's own rounding compounding through a
chain.
Two things before you spend time on it.
The number moved. When I wrote 995 of 1001 there was still one fault
left in Mesa, an output channel count that is not a multiple of two,
which the CNA reads in pairs. With that fixed it is 1000 of 1001, so it
is one output rather than six, and whether one output is even worth
chasing is a fair question. The comparison is still worth having for the
LAYERS rather than the final vector.
And the oracle matters more than the run. A per output comparison
against the CPU is not enough on its own past the first layer or two,
because tflite's requant and the hardware's disagree by design and that
disagreement compounds: at operator 6 a flawless accelerator scores 4 of
128 channels against the CPU. vendor-capture/chainmodel.py in
https://github.com/gahingwoo/linux-rk3576-npu
runs the graph twice from the model file, once with tflite's
SaturatingRoundingDoublingHighMul and RoundingDivideByPOT and once with
the hardware's single half up shift, and prints what a perfect
accelerator would score at every operator. Read your numbers against
that column rather than against 128 of 128, or every deep layer will
look broken on both SoCs.
If it is easier, mn_L00 through mn_L26 in that repository are MobileNet
with its graph output moved to each operator's output, which is a four
byte patch of the flatbuffer and needs no converter. Those are what the
per layer table came from.
Jiaxing
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support
2026-08-14 8:26 ` Jiaxing Hu
@ 2026-08-14 11:08 ` Igor Paunovic
0 siblings, 0 replies; 23+ messages in thread
From: Igor Paunovic @ 2026-08-14 11:08 UTC (permalink / raw)
To: Jiaxing Hu
Cc: Igor Paunovic, Tomeu Vizoso, Heiko Stübner, Chaoyi Chen,
alchark, dri-devel, linux-rockchip, linux-arm-kernel,
linux-kernel
Hi Jiaxing,
Here is the RK3588 column, all 27 operators, ROCKET_SEED=7, scored the
way perch.py scores: a channel is good when its maxdiff (md below)
against max(cpu, output zero point) is at most 1.
Setup: Orange Pi 5 Plus (RK3588), kernel 7.2.0-rc6, the rocket driver
from this kernel's tree rebuilt with my clocks-by-name and devfreq
patches on top, Mesa at bf70ab68a21, teflon delegate, model
mobilenet_v1_1_224_quant.tflite from the Mesa test suite
(md5 4f348b87dca3315d2b3646cf5a3b31cf), per-operator models generated
with the four byte output patch you described. The "correct hw" column
is your chainmodel.py against the same model file. One difference to
flag up front: against this model file chainmodel prints 36/256 for
operator 8 where your table has 34/256, so our model files are not
byte-identical, and the columns below should be read against each
other rather than against your FINDINGS numbers.
op kind correct hw RK3588
0 conv 32/32 md 1 32/32 md 1
1 depthwise 28/32 md 3 28/32 md 3
2 1x1 22/64 md 6 22/64 md 6
3 depthwise 21/64 md 13 21/64 md 13
4 1x1 18/128 md 14 8/128 md 15
5 depthwise 9/128 md 13 3/128 md 15
6 1x1 4/128 md 23 1/128 md 24
7 depthwise 7/128 md 10 3/128 md 12
8 1x1 36/256 md 7 20/256 md 17
9 depthwise 32/256 md 11 22/256 md 13
10 1x1 29/256 md 9 28/256 md 10
11 depthwise 53/256 md 8 48/256 md 11
12 1x1 166/512 md 11 121/512 md 13
13 depthwise 142/512 md 9 137/512 md 12
14 1x1 82/512 md 7 72/512 md 10
15 depthwise 154/512 md 10 115/512 md 15
16 1x1 82/512 md 7 71/512 md 10
17 depthwise 156/512 md 14 141/512 md 21
18 1x1 92/512 md 7 77/512 md 10
19 depthwise 170/512 md 7 152/512 md 12
20 1x1 102/512 md 8 84/512 md 10
21 depthwise 166/512 md 13 150/512 md 12
22 1x1 174/512 md 6 161/512 md 7
23 depthwise 293/512 md 5 261/512 md 7
24 1x1 671/1024 md 6 636/1024 md 5
25 depthwise 718/1024 md 9 684/1024 md 8
26 1x1 572/1024 md 25 574/1024 md 23
Three things stand out from here.
Operators 0 through 3 score identically to your chain simulation --
same good-channel counts, same maxdiff -- including operator 3, the
stride 2 depthwise with the asymmetric padding you suspect for the
first RK3576 divergence. They are not byte-identical to the simulated
hardware: diffing the raw tensors against the requant_hw chain shows a
few hundred elements per surface already off by 1-4 at operators 0-3.
That looks like the same small extra rounding difference that pushes
the scores below your column from operator 4 on; through operator 3 it
just stays under the maxdiff <= 1 scoring threshold.
There is no md 255 anywhere. From operator 4 on, RK3588 sits somewhat
below the simulation (8 vs 18 at op 4, 1 vs 4 at op 6), but the
maxdiff never exceeds 24 across all 27 operators and the deep layers
track the simulation closely (574 vs 572 at op 26).
A control run with ROCKET_SEED=11 keeps the same character: operator 0
still 32/32, no saturated maxdiff anywhere, worst case md 35 at
operator 6.
So from the RK3588 side your read looks right: the deep-layer
compounding is the reference artifact, and the RK3576 collapse from
operator 4 with maxdiff 255 has no counterpart here.
Igor
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v7 09/10] arm64: dts: rockchip: rk3576: add NPU (RKNN) nodes
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
` (7 preceding siblings ...)
2026-08-12 9:41 ` [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support Jiaxing Hu
@ 2026-08-12 9:41 ` Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 10/10] arm64: dts: rockchip: rk3576-rock-4d: enable NPU Jiaxing Hu
9 siblings, 0 replies; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:41 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
Add the two RKNN cores and their IOMMUs, plus the NPU power-domain
resets the pmdomain driver now cycles on power-on. Both cores are
disabled by default; boards enable what they wire up.
Each core lists both NPU power domains, its own first. The compute path
needs NPU1 powered even when only core 0 runs, and a node with a single
domain would be auto-attached by the driver core before the driver can
attach the list itself. The IOMMUs keep one domain each, since they rely
on that same auto-attach.
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
arch/arm64/boot/dts/rockchip/rk3576.dtsi | 80 +++++++++++++++++++++++-
1 file changed, 78 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/boot/dts/rockchip/rk3576.dtsi b/arch/arm64/boot/dts/rockchip/rk3576.dtsi
index b0c0d3c8b..1e6dd039f 100644
--- a/arch/arm64/boot/dts/rockchip/rk3576.dtsi
+++ b/arch/arm64/boot/dts/rockchip/rk3576.dtsi
@@ -1070,14 +1070,22 @@ power-domain@RK3576_PD_NPUTOP {
power-domain@RK3576_PD_NPU0 {
reg = <RK3576_PD_NPU0>;
clocks = <&cru HCLK_RKNN_ROOT>,
- <&cru ACLK_RKNN0>;
+ <&cru ACLK_RKNN0>,
+ <&cru CLK_RKNN_DSU0>,
+ <&cru ACLK_RKNN_CBUF>,
+ <&cru HCLK_RKNN_CBUF>;
+ resets = <&cru SRST_A_RKNN0_BIU>;
pm_qos = <&qos_npu_m0>;
#power-domain-cells = <0>;
};
power-domain@RK3576_PD_NPU1 {
reg = <RK3576_PD_NPU1>;
clocks = <&cru HCLK_RKNN_ROOT>,
- <&cru ACLK_RKNN1>;
+ <&cru ACLK_RKNN1>,
+ <&cru CLK_RKNN_DSU0>,
+ <&cru ACLK_RKNN_CBUF>,
+ <&cru HCLK_RKNN_CBUF>;
+ resets = <&cru SRST_A_RKNN1_BIU>;
pm_qos = <&qos_npu_m1>;
#power-domain-cells = <0>;
};
@@ -1832,6 +1840,74 @@ qos_npu_m1ro: qos@27f22100 {
reg = <0x0 0x27f22100 0x0 0x20>;
};
+ rknn_core_0: npu@27700000 {
+ compatible = "rockchip,rk3576-rknn-core";
+ reg = <0x0 0x27700000 0x0 0x1000>,
+ <0x0 0x27701000 0x0 0x1000>,
+ <0x0 0x27703000 0x0 0x1000>;
+ reg-names = "pc", "cna", "core";
+ interrupts = <GIC_SPI 247 IRQ_TYPE_LEVEL_HIGH>;
+ clocks = <&cru ACLK_RKNN0>, <&cru HCLK_RKNN_ROOT>,
+ <&cru CLK_RKNN_DSU0>, <&cru PCLK_NPUTOP_ROOT>,
+ <&cru ACLK_RKNN_CBUF>, <&cru HCLK_RKNN_CBUF>;
+ clock-names = "aclk", "hclk", "npu", "pclk",
+ "aclk_cbuf", "hclk_cbuf";
+ resets = <&cru SRST_A_RKNN0>;
+ reset-names = "srst_a";
+ power-domains = <&power RK3576_PD_NPU0>, <&power RK3576_PD_NPU1>;
+ iommus = <&rknn_mmu_0>;
+ status = "disabled";
+ };
+
+ rknn_mmu_0: iommu@27702000 {
+ compatible = "rockchip,rk3576-iommu", "rockchip,rk3568-iommu";
+ reg = <0x0 0x27702000 0x0 0x100>,
+ <0x0 0x27702100 0x0 0x100>;
+ interrupts = <GIC_SPI 247 IRQ_TYPE_LEVEL_HIGH>;
+ clocks = <&cru ACLK_RKNN0>, <&cru HCLK_RKNN_ROOT>,
+ <&cru CLK_RKNN_DSU0>, <&cru ACLK_RKNN_CBUF>,
+ <&cru HCLK_RKNN_CBUF>;
+ clock-names = "aclk", "iface", "npu",
+ "aclk_cbuf", "hclk_cbuf";
+ #iommu-cells = <0>;
+ power-domains = <&power RK3576_PD_NPU0>;
+ status = "disabled";
+ };
+
+ rknn_core_1: npu@27708000 {
+ compatible = "rockchip,rk3576-rknn-core";
+ reg = <0x0 0x27708000 0x0 0x1000>,
+ <0x0 0x27709000 0x0 0x1000>,
+ <0x0 0x2770b000 0x0 0x1000>;
+ reg-names = "pc", "cna", "core";
+ interrupts = <GIC_SPI 248 IRQ_TYPE_LEVEL_HIGH>;
+ clocks = <&cru ACLK_RKNN1>, <&cru HCLK_RKNN_ROOT>,
+ <&cru CLK_RKNN_DSU0>, <&cru PCLK_NPUTOP_ROOT>,
+ <&cru ACLK_RKNN_CBUF>, <&cru HCLK_RKNN_CBUF>;
+ clock-names = "aclk", "hclk", "npu", "pclk",
+ "aclk_cbuf", "hclk_cbuf";
+ resets = <&cru SRST_A_RKNN1>;
+ reset-names = "srst_a";
+ power-domains = <&power RK3576_PD_NPU1>, <&power RK3576_PD_NPU0>;
+ iommus = <&rknn_mmu_1>;
+ status = "disabled";
+ };
+
+ rknn_mmu_1: iommu@2770a000 {
+ compatible = "rockchip,rk3576-iommu", "rockchip,rk3568-iommu";
+ reg = <0x0 0x2770a000 0x0 0x100>,
+ <0x0 0x2770a100 0x0 0x100>;
+ interrupts = <GIC_SPI 248 IRQ_TYPE_LEVEL_HIGH>;
+ clocks = <&cru ACLK_RKNN1>, <&cru HCLK_RKNN_ROOT>,
+ <&cru CLK_RKNN_DSU0>, <&cru ACLK_RKNN_CBUF>,
+ <&cru HCLK_RKNN_CBUF>;
+ clock-names = "aclk", "iface", "npu",
+ "aclk_cbuf", "hclk_cbuf";
+ #iommu-cells = <0>;
+ power-domains = <&power RK3576_PD_NPU1>;
+ status = "disabled";
+ };
+
gmac0: ethernet@2a220000 {
compatible = "rockchip,rk3576-gmac", "snps,dwmac-4.20a";
reg = <0x0 0x2a220000 0x0 0x10000>;
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* [PATCH v7 10/10] arm64: dts: rockchip: rk3576-rock-4d: enable NPU
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
` (8 preceding siblings ...)
2026-08-12 9:41 ` [PATCH v7 09/10] arm64: dts: rockchip: rk3576: add NPU (RKNN) nodes Jiaxing Hu
@ 2026-08-12 9:41 ` Jiaxing Hu
2026-08-12 10:20 ` Chaoyi Chen
9 siblings, 1 reply; 23+ messages in thread
From: Jiaxing Hu @ 2026-08-12 9:41 UTC (permalink / raw)
To: tomeu, heiko, robh, krzk+dt, conor+dt, joro, will, robin.murphy,
ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, chaoyi.chen, diederik, dri-devel,
linux-rockchip, iommu, linux-pm, devicetree, linux-arm-kernel,
linux-kernel, Jiaxing Hu
Enable rknn_core_0 and its IOMMU on the Radxa ROCK 4D and supply the
core from vdd_npu_s0.
The supply is marked always-on because the NPU power domains are what
gate the block here, and dropping the rail underneath them takes an
async SError on the next power-on rather than a clean retry. Only
rknn_core_0 is enabled: the driver binds one core per node and the
second core is left to whoever can test it.
Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
---
arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts | 10 ++++++++++
1 file changed, 10 insertions(+)
diff --git a/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts b/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts
index 272af1012..965e0906b 100644
--- a/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts
+++ b/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts
@@ -442,6 +442,7 @@ regulator-state-mem {
};
vdd_npu_s0: dcdc-reg2 {
+ regulator-always-on;
regulator-boot-on;
regulator-enable-ramp-delay = <400>;
regulator-min-microvolt = <550000>;
@@ -869,3 +870,12 @@ vp0_out_hdmi: endpoint@ROCKCHIP_VOP2_EP_HDMI0 {
remote-endpoint = <&hdmi_in_vp0>;
};
};
+
+&rknn_core_0 {
+ npu-supply = <&vdd_npu_s0>;
+ status = "okay";
+};
+
+&rknn_mmu_0 {
+ status = "okay";
+};
--
2.43.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v7 10/10] arm64: dts: rockchip: rk3576-rock-4d: enable NPU
2026-08-12 9:41 ` [PATCH v7 10/10] arm64: dts: rockchip: rk3576-rock-4d: enable NPU Jiaxing Hu
@ 2026-08-12 10:20 ` Chaoyi Chen
0 siblings, 0 replies; 23+ messages in thread
From: Chaoyi Chen @ 2026-08-12 10:20 UTC (permalink / raw)
To: Jiaxing Hu, tomeu, heiko, robh, krzk+dt, conor+dt, joro, will,
robin.murphy, ulfh, p.zabel, ogabbay, zhangqing
Cc: royalnet026, alchark, diederik, dri-devel, linux-rockchip, iommu,
linux-pm, devicetree, linux-arm-kernel, linux-kernel
Hi Jiaxing,
On 8/12/2026 5:41 PM, Jiaxing Hu wrote:
> Enable rknn_core_0 and its IOMMU on the Radxa ROCK 4D and supply the
> core from vdd_npu_s0.
>
> The supply is marked always-on because the NPU power domains are what
> gate the block here, and dropping the rail underneath them takes an
> async SError on the next power-on rather than a clean retry. Only
> rknn_core_0 is enabled: the driver binds one core per node and the
> second core is left to whoever can test it.
>
> Signed-off-by: Jiaxing Hu <gahing@gahingwoo.com>
> ---
> arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts | 10 ++++++++++
> 1 file changed, 10 insertions(+)
>
> diff --git a/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts b/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts
> index 272af1012..965e0906b 100644
> --- a/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts
> +++ b/arch/arm64/boot/dts/rockchip/rk3576-rock-4d.dts
> @@ -442,6 +442,7 @@ regulator-state-mem {
> };
>
> vdd_npu_s0: dcdc-reg2 {
> + regulator-always-on;
> regulator-boot-on;
> regulator-enable-ramp-delay = <400>;
> regulator-min-microvolt = <550000>;
> @@ -869,3 +870,12 @@ vp0_out_hdmi: endpoint@ROCKCHIP_VOP2_EP_HDMI0 {
> remote-endpoint = <&hdmi_in_vp0>;
> };
> };
> +
> +&rknn_core_0 {
> + npu-supply = <&vdd_npu_s0>;
Out of curiosity, I searched for code about this supply in the rocket
driver and found nothing.
Then what is the consumer of this regulator? I have reason to suspect
they were automatically disabled.
> + status = "okay";
> +};
> +
> +&rknn_mmu_0 {
> + status = "okay";
> +};
--
Best,
Chaoyi
^ permalink raw reply [flat|nested] 23+ messages in thread